Difference between revisions of "AI benchmarks"

From GISAXS
Jump to: navigation, search
(Visual)
(Assess Specific Attributes)
 
Line 30: Line 30:
 
==Software/Coding==
 
==Software/Coding==
 
* 2025-02: [https://arxiv.org/abs/2502.12115SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?] ([https://github.com/openai/SWELancer-Benchmark code])
 
* 2025-02: [https://arxiv.org/abs/2502.12115SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?] ([https://github.com/openai/SWELancer-Benchmark code])
 +
 +
==Math==
 +
* [https://www.vals.ai/benchmarks/aime-2025-03-13 AIME Benchmark]
  
 
==Visual==
 
==Visual==

Latest revision as of 13:39, 16 April 2025

General

Methods

Task Length

GmZHL8xWQAAtFlF.jpeg

Assess Specific Attributes

Various

Hallucination

Software/Coding

Math

Visual

Conversation

Creativity

Reasoning

Assistant/Agentic

See: AI Agents: Optimization

Science

See: Science Benchmarks