Difference between revisions of "AI benchmarks"

From GISAXS
Jump to: navigation, search
(Assess Specific Attributes)
(Assistant/Agentic)
(One intermediate revision by the same user not shown)
Line 35: Line 35:
  
 
==Conversation==
 
==Conversation==
* 2025-01: [https://arxiv.org/abs/2501.17399 MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs] ([https://scale.com/research/multichallenge project], [https://github.com/ekwinox117/multi-challenge code])
+
* 2025-01: [https://arxiv.org/abs/2501.17399 MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs] ([https://scale.com/research/multichallenge project], [https://github.com/ekwinox117/multi-challenge code], [https://scale.com/leaderboard/multichallenge leaderboard])
  
 
==Creativity==
 
==Creativity==
Line 49: Line 49:
  
 
==Assistant/Agentic==
 
==Assistant/Agentic==
 +
See: [[AI_Agents#Optimization|AI Agents: Optimization]]
 
* [https://arxiv.org/abs/2311.12983 GAIA: a benchmark for General AI Assistants]
 
* [https://arxiv.org/abs/2311.12983 GAIA: a benchmark for General AI Assistants]
 
* [https://www.galileo.ai/blog/agent-leaderboard Galileo AI] [https://huggingface.co/spaces/galileo-ai/agent-leaderboard Agent Leaderboard]
 
* [https://www.galileo.ai/blog/agent-leaderboard Galileo AI] [https://huggingface.co/spaces/galileo-ai/agent-leaderboard Agent Leaderboard]

Revision as of 16:28, 14 April 2025

General

Methods

Task Length

GmZHL8xWQAAtFlF.jpeg

Assess Specific Attributes

Various

Hallucination

Software/Coding

Visual

Conversation

Creativity

Reasoning

Assistant/Agentic

See: AI Agents: Optimization

Science

See: Science Benchmarks