Benchmarks.

Understand what each score measures, how to read it, and where comparison stops. All 166 reported benchmarks remain visible—even when only one lab publishes them.

01
knowledge workAgents' Last Exam

Can the model complete long-running professional workflows across many occupations?

Measures
End-to-end agent work across 55 professional fields.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
02
knowledge workManagement Consulting Tasks (Internal)

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
03
knowledge workBig Finance Bench

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
04
codingSWE-Bench Pro

Can an agent resolve difficult software issues from real repositories?

Measures
Repository-level issue resolution on the SWE-Bench Pro task set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
05
codingDeepSWE v1.1

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
06
agentsTerminal-Bench 2.1

Can a model-driven agent complete realistic work inside a terminal?

Measures
Multi-step command-line, coding, configuration, debugging, and systems tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
07
scienceGeneBench Pro

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
08
scienceLifeSciBench

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
09
scienceMedChemBench (Internal)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
10
knowledge workHealthBench Professional

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
11
agentsOSWorld 2.0

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
12
agentsBrowseComp

Can the model find hard-to-locate facts through multi-step web research?

Measures
Persistent browsing, source discovery, and synthesis for intentionally difficult questions.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
13
visionBenchCAD

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
14
visionBenchCAD (python tool)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
15
cybersecurityCapture-the-Flag Challenges

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
16
cybersecuritySEC-Bench Pro

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
17
cybersecurityExploitBench

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
18
cybersecurityExploitGym

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
19
codingInternal Research Debugging Evaluation

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
20
codingKernelGen 1P

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
21
codingNanoGPT

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
22
codingPostTrainBench Lite

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
23
codingRSI Index

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
24
visionMMMU Pro (no tools)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
25
visionMMMU Pro (with tools)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
26
visiongdp.pdf

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
27
scienceGPQA Diamond

Can the model reason through specialist biology, physics, and chemistry questions?

Measures
Multiple-choice science questions written and validated by domain experts.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
28
reasoningFrontierMath Tier 1-3 (v2)

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
29
reasoningFrontierMath Tier 4 (v2)

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
30
agentsAutomationBench

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
31
agentsToolathlon

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
32
long contextOpenAI MRCR v2 · 8-needle · 256K-512K

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
33
long contextOpenAI MRCR v2 · 8-needle · 512K-1M

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
34
long contextGraphWalks BFS · 256K F1

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
35
long contextGraphWalks BFS · 1M F1

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
36
reasoningARC-AGI-3

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
37
codingDeepSWE

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
38
codingProgram Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
39
codingFrontierSWE

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
40
codingSWE Marathon

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
41
codingPostTrain Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
42
codingMLS Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
43
codingKimi Code Bench 2.0 (Internal)

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
44
agentsDeepSearchQA · F1

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
45
agentsToolathlon-Verified

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
46
agentsMCP Atlas

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
47
knowledge workJob Bench

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
48
knowledge workOffice QA Pro

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
49
knowledge workSpreadsheetBench 2

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
50
knowledge workDECK-Bench (Internal)

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
51
reasoningHLE-Full

Can the model answer deliberately difficult expert questions across academic fields?

Measures
Broad expert-level reasoning and knowledge without external tools.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
52
reasoningHLE-Full with tools

How much do search or code tools help on Humanity's Last Exam?

Measures
The full expert-question set with the reporting lab's stated tool setup.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
53
visionMMMU-Pro with python

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
54
visionCharXiv · RQ

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
55
visionCharXiv · RQ with python

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
56
visionMathVision

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
57
visionMathVision with python

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
58
visionBabyVision with python

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
59
visionZeroBench main · pass@5

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
60
visionZeroBench main with python · pass@5

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
61
visionWorldVQA ForceAnswer

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
62
visionOmniDocBench

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
63
visionPerceptionBench

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
64
codingFrontierCode · Diamond

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
65
visionBlueprint-Bench 2

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
66
agentsOSWorld-Verified

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
67
knowledge workLegal Agent Benchmark

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
68
scienceBioMysteryBench · hard

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
69
scienceBioMysteryBench · human solved

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
70
reasoningHumanity's Last Exam · search + code

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
71
reasoningARC-AGI-2 · verified

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
72
agentsTerminal-Bench 2.0

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
73
agentsTerminal-Bench 2.0 · Codex harness

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
74
codingSWE-Bench Verified

Can an agent fix a real GitHub issue and pass the repository's tests?

Measures
The share of engineer-verified software issues resolved by the full model-plus-agent system.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
75
codingLiveCodeBench Pro · Elo

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. This benchmark reports points rather than percent correct.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
76
scienceSciCode

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
77
agentsτ²-bench · Retail

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
78
agentsτ²-bench · Telecom

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
79
reasoningMMMLU

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
80
long contextMRCR v2 · 8-needle · 128K average

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
81
long contextMRCR v2 · 8-needle · 1M pointwise

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
82
knowledge workFinance Agent v2

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
83
visionCharXiv Reasoning (no tools)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
84
codingMLE-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
85
visionCharXiv Reasoning (with tools)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
86
long contextMRCR Long Context (1M context window)

Can the model retain and reason over evidence spread across a very long input?

Measures
Long-context retrieval and reasoning, not context-window size alone.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
87
agentsOSWorld 2.0 · strict binary completion

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
88
agentsOSWorld 2.0 · partial score

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
89
agentsWebArena-Verified

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
90
knowledge workHealthBench Pro

Can the model produce useful work in a professional knowledge-work setting?

Measures
Professional analysis, document, finance, office, or domain-specific tasks.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
91
visionBabyVision (with tools)

Can the model extract and reason over information in images or documents?

Measures
Visual perception and multimodal reasoning under the reported protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
92
scienceMBCT

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
93
scienceVCT

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
94
scienceHPCT

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
95
scienceWMDP-Bio

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
96
scienceWMDP-Chem

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
97
scienceProtocolQA

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
98
scienceSeqQA (agentic)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
99
scienceABC Bench (Fragment Design)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
100
scienceABC Bench (Liquid Handling)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
101
scienceABC Bench (Screening Evasion)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
102
scienceBioDesign Tools (avg)

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
103
cybersecurityCybench (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
104
cybersecurityCurated CTFs (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
105
cybersecurityCyberGym (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
106
cybersecurityExploitGym (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
107
cybersecurityCyScenarioBench (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
108
cybersecuritySocial Engineering

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
109
agentsAIRS-Bench

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
110
agentsSHADE-Arena

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
111
agentsGDM-Stealth (of 4)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
112
agentsGDM Situational Awareness

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
113
scienceRefusals: BioTIER

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
The source reports a refusal percentage; interpret with the named safety objective.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
114
scienceRefusals: Chemical Agents

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
The source reports a refusal percentage; interpret with the named safety objective.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
115
cybersecurityCyber Misuse Chat (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
116
cybersecurityCatastrophic Cyber Misuse (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
117
cybersecurityPoly-Guard Bench (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
118
agentsMASK

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
119
agentsAgentic Misalignment

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
120
cybersecurityJailbreak StrongREJECT v2 (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
121
cybersecurityFORTRESS (ARS)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports an attack-risk score as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
122
agentsAgentHarm (ASR)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
123
agentsAgentDojo (pass@1 ASR)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
124
agentsGraySwan ART (pass@1 ASR)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Lower is better. The source reports an attack success rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
125
agentsOR-Bench (FRR)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Lower is better. The source reports a false-refusal rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
126
cybersecurityCyber Misuse Chat (FRR)

Can the model diagnose or complete difficult cybersecurity tasks?

Measures
Security analysis or exploitation under a controlled evaluation environment.
Read it
Lower is better. The source reports a false-refusal rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
127
agentsAgentHarm Verified (benign) (FRR)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Lower is better. The source reports a false-refusal rate as a percentage.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
128
agentsSAVE-Bench

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
129
reasoningInternal Sycophancy

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
130
reasoningDeceptionBench

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
131
reasoningHLE Calibration

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
132
codingCursorBench-3

Can the coding agent solve realistic, underspecified tasks drawn from Cursor's own engineering work?

Measures
Correctness on private, long-horizon software tasks run in Cursor's production-like harness.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
133
codingSWE-bench Multilingual

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
134
agentsTerminal-Bench · Cursor report

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
135
scienceCritPt

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
136
reasoningAIME 2026

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
137
reasoningHMMT · November 2025

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
138
reasoningHMMT · February 2026

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
139
reasoningIMOAnswerBench

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
140
codingNL2Repo

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
141
agentsTerminal-Bench 2.1 · Terminus-2

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
142
agentsTerminal-Bench 2.1 · best reported harness

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
143
agentsMCP-Atlas · public set

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
144
agentsTool-Decathlon

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
145
codingSWE-fficiency

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
146
codingKernelBench Hard

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
147
codingPostTrainBench · normalized score

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. This benchmark reports points rather than percent correct.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
148
codingMulti-SWE-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
149
codingVIBE-Pro · average

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
150
agentsBrowseComp · context management

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
151
agentsWide Search

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
152
agentsRISE

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
153
agentsBFCL multi-turn

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
154
reasoningAIME 2025

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
155
reasoningIFBench

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
156
reasoningHumanity's Last Exam · full set

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
157
codingKimi Code Bench v2 (Internal)

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
158
codingMLS Bench Lite

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
159
agentsKimi Claw 24/7 Bench (Internal)

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source ↗
160
agentsMCP Mark Verified

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
161
agentsTerminal-Bench Hard

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
162
codingLiveCodeBench v6

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
163
scienceAstaBench · overall

Can the model reason accurately about difficult scientific material?

Measures
Scientific knowledge and reasoning under the benchmark's published question set.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
164
reasoningHMMT 2025 · pass@1

Can the model solve difficult problems that require more than factual recall?

Measures
Multi-step reasoning on academic, mathematical, or abstract problems.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
165
codingCodeforces · rating

Can the model complete substantial programming work under this benchmark's agent setup?

Measures
Software implementation, debugging, or repository work under the published evaluation protocol.
Read it
Higher is better. This benchmark reports points rather than percent correct.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗
166
agentsτ²-bench · pass@1

Can the model plan, use tools, and finish a multi-step task?

Measures
Agentic execution across tools, environments, or long-running workflows.
Read it
Higher is better. The value is the percentage reported in this lab's table.
Boundary
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source ↗