Difference between revisions of "AI benchmarks"

From GISAXS
Jump to: navigation, search
(Reasoning)
(Assess Specific Attributes)
Line 33: Line 33:
 
==Visual==
 
==Visual==
 
* 2025-03: [https://arxiv.org/abs/2503.14607 Can Large Vision Language Models Read Maps Like a Human?] MapBench
 
* 2025-03: [https://arxiv.org/abs/2503.14607 Can Large Vision Language Models Read Maps Like a Human?] MapBench
 +
 +
==Conversation==
 +
* 2025-01: [https://arxiv.org/abs/2501.17399 MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs] ([https://scale.com/research/multichallenge project], [https://github.com/ekwinox117/multi-challenge code])
  
 
==Creativity==
 
==Creativity==

Revision as of 16:26, 14 April 2025

General

Methods

Task Length

GmZHL8xWQAAtFlF.jpeg

Assess Specific Attributes

Various

Hallucination

Software/Coding

Visual

Conversation

Creativity

Reasoning

Assistant/Agentic

Science

See: Science Benchmarks