Difference between revisions of "AI benchmarks"

From GISAXS
Jump to: navigation, search
(Methods)
(Assistant/Agentic)
(3 intermediate revisions by the same user not shown)
Line 33: Line 33:
 
==Visual==
 
==Visual==
 
* 2025-03: [https://arxiv.org/abs/2503.14607 Can Large Vision Language Models Read Maps Like a Human?] MapBench
 
* 2025-03: [https://arxiv.org/abs/2503.14607 Can Large Vision Language Models Read Maps Like a Human?] MapBench
 +
 +
==Conversation==
 +
* 2025-01: [https://arxiv.org/abs/2501.17399 MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs] ([https://scale.com/research/multichallenge project], [https://github.com/ekwinox117/multi-challenge code], [https://scale.com/leaderboard/multichallenge leaderboard])
  
 
==Creativity==
 
==Creativity==
Line 42: Line 45:
 
==Reasoning==
 
==Reasoning==
 
* [https://scale.com/leaderboard/enigma_eval ENIGMAEVAL]: "reasoning" leaderboard ([https://static.scale.com/uploads/654197dc94d34f66c0f5184e/EnigmaEval%20v4.pdf paper])
 
* [https://scale.com/leaderboard/enigma_eval ENIGMAEVAL]: "reasoning" leaderboard ([https://static.scale.com/uploads/654197dc94d34f66c0f5184e/EnigmaEval%20v4.pdf paper])
 +
* [https://bethgelab.github.io/sober-reasoning/ Sober Reasoning Leaderboard]
 +
** 2025-04: [https://arxiv.org/abs/2504.07086 A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility]
  
 
==Assistant/Agentic==
 
==Assistant/Agentic==
 +
See: [[AI_Agents#Optimization|AI Agents: Optimization]]
 
* [https://arxiv.org/abs/2311.12983 GAIA: a benchmark for General AI Assistants]
 
* [https://arxiv.org/abs/2311.12983 GAIA: a benchmark for General AI Assistants]
 
* [https://www.galileo.ai/blog/agent-leaderboard Galileo AI] [https://huggingface.co/spaces/galileo-ai/agent-leaderboard Agent Leaderboard]
 
* [https://www.galileo.ai/blog/agent-leaderboard Galileo AI] [https://huggingface.co/spaces/galileo-ai/agent-leaderboard Agent Leaderboard]

Revision as of 16:28, 14 April 2025

General

Methods

Task Length

GmZHL8xWQAAtFlF.jpeg

Assess Specific Attributes

Various

Hallucination

Software/Coding

Visual

Conversation

Creativity

Reasoning

Assistant/Agentic

See: AI Agents: Optimization

Science

See: Science Benchmarks