Benchmarks.
Understand what each score measures, how to read it, and where comparison stops. All 166 reported benchmarks remain visible—even when only one lab publishes them.
01knowledge workAgents' Last Exam+
“Can the model complete long-running professional workflows across many occupations?”
- Measures
- End-to-end agent work across 55 professional fields.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
02knowledge workManagement Consulting Tasks (Internal)+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
03knowledge workBig Finance Bench+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
04codingSWE-Bench Pro+
“Can an agent resolve difficult software issues from real repositories?”
- Measures
- Repository-level issue resolution on the SWE-Bench Pro task set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
05codingDeepSWE v1.1+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
06agentsTerminal-Bench 2.1+
“Can a model-driven agent complete realistic work inside a terminal?”
- Measures
- Multi-step command-line, coding, configuration, debugging, and systems tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
07scienceGeneBench Pro+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
08scienceLifeSciBench+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
09scienceMedChemBench (Internal)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
10knowledge workHealthBench Professional+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
11agentsOSWorld 2.0+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
12agentsBrowseComp+
“Can the model find hard-to-locate facts through multi-step web research?”
- Measures
- Persistent browsing, source discovery, and synthesis for intentionally difficult questions.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
13visionBenchCAD+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
14visionBenchCAD (python tool)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
15cybersecurityCapture-the-Flag Challenges+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
16cybersecuritySEC-Bench Pro+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
17cybersecurityExploitBench+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
18cybersecurityExploitGym+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
19codingInternal Research Debugging Evaluation+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
20codingKernelGen 1P+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
21codingNanoGPT+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
22codingPostTrainBench Lite+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
23codingRSI Index+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
24visionMMMU Pro (no tools)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
25visionMMMU Pro (with tools)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
26visiongdp.pdf+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
27scienceGPQA Diamond+
“Can the model reason through specialist biology, physics, and chemistry questions?”
- Measures
- Multiple-choice science questions written and validated by domain experts.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
28reasoningFrontierMath Tier 1-3 (v2)+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
29reasoningFrontierMath Tier 4 (v2)+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
30agentsAutomationBench+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
31agentsToolathlon+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
32long contextOpenAI MRCR v2 · 8-needle · 256K-512K+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
33long contextOpenAI MRCR v2 · 8-needle · 512K-1M+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
34long contextGraphWalks BFS · 256K F1+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
35long contextGraphWalks BFS · 1M F1+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
36reasoningARC-AGI-3+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
37codingDeepSWE+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
38codingProgram Bench+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
39codingFrontierSWE+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
40codingSWE Marathon+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
41codingPostTrain Bench+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
42codingMLS Bench+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
43codingKimi Code Bench 2.0 (Internal)+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
44agentsDeepSearchQA · F1+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
45agentsToolathlon-Verified+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
46agentsMCP Atlas+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
47knowledge workJob Bench+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
48knowledge workOffice QA Pro+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
49knowledge workSpreadsheetBench 2+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
50knowledge workDECK-Bench (Internal)+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
51reasoningHLE-Full+
“Can the model answer deliberately difficult expert questions across academic fields?”
- Measures
- Broad expert-level reasoning and knowledge without external tools.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
52reasoningHLE-Full with tools+
“How much do search or code tools help on Humanity's Last Exam?”
- Measures
- The full expert-question set with the reporting lab's stated tool setup.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
53visionMMMU-Pro with python+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
54visionCharXiv · RQ+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
55visionCharXiv · RQ with python+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
56visionMathVision+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
57visionMathVision with python+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
58visionBabyVision with python+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
59visionZeroBench main · pass@5+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
60visionZeroBench main with python · pass@5+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
61visionWorldVQA ForceAnswer+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
62visionOmniDocBench+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
63visionPerceptionBench+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
64codingFrontierCode · Diamond+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
65visionBlueprint-Bench 2+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
66agentsOSWorld-Verified+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
67knowledge workLegal Agent Benchmark+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
68scienceBioMysteryBench · hard+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
69scienceBioMysteryBench · human solved+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
70reasoningHumanity's Last Exam · search + code+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
71reasoningARC-AGI-2 · verified+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
72agentsTerminal-Bench 2.0+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
73agentsTerminal-Bench 2.0 · Codex harness+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
74codingSWE-Bench Verified+
“Can an agent fix a real GitHub issue and pass the repository's tests?”
- Measures
- The share of engineer-verified software issues resolved by the full model-plus-agent system.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
75codingLiveCodeBench Pro · Elo+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. This benchmark reports points rather than percent correct.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
76scienceSciCode+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
77agentsτ²-bench · Retail+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
78agentsτ²-bench · Telecom+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
79reasoningMMMLU+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
80long contextMRCR v2 · 8-needle · 128K average+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
81long contextMRCR v2 · 8-needle · 1M pointwise+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
82knowledge workFinance Agent v2+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
83visionCharXiv Reasoning (no tools)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
84codingMLE-Bench+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
85visionCharXiv Reasoning (with tools)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
86long contextMRCR Long Context (1M context window)+
“Can the model retain and reason over evidence spread across a very long input?”
- Measures
- Long-context retrieval and reasoning, not context-window size alone.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
87agentsOSWorld 2.0 · strict binary completion+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
88agentsOSWorld 2.0 · partial score+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
89agentsWebArena-Verified+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
90knowledge workHealthBench Pro+
“Can the model produce useful work in a professional knowledge-work setting?”
- Measures
- Professional analysis, document, finance, office, or domain-specific tasks.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
91visionBabyVision (with tools)+
“Can the model extract and reason over information in images or documents?”
- Measures
- Visual perception and multimodal reasoning under the reported protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
92scienceMBCT+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
93scienceVCT+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
94scienceHPCT+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
95scienceWMDP-Bio+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
96scienceWMDP-Chem+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
97scienceProtocolQA+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
98scienceSeqQA (agentic)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
99scienceABC Bench (Fragment Design)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
100scienceABC Bench (Liquid Handling)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
101scienceABC Bench (Screening Evasion)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
102scienceBioDesign Tools (avg)+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
103cybersecurityCybench (pass@1)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
104cybersecurityCurated CTFs (pass@1)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
105cybersecurityCyberGym (pass@1)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
106cybersecurityExploitGym (pass@1)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
107cybersecurityCyScenarioBench (pass@1)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
108cybersecuritySocial Engineering+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
109agentsAIRS-Bench+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
110agentsSHADE-Arena+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
111agentsGDM-Stealth (of 4)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
112agentsGDM Situational Awareness+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
113scienceRefusals: BioTIER+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- The source reports a refusal percentage; interpret with the named safety objective.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
114scienceRefusals: Chemical Agents+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- The source reports a refusal percentage; interpret with the named safety objective.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
115cybersecurityCyber Misuse Chat (ASR)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
116cybersecurityCatastrophic Cyber Misuse (ASR)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
117cybersecurityPoly-Guard Bench (ASR)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
118agentsMASK+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
119agentsAgentic Misalignment+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
120cybersecurityJailbreak StrongREJECT v2 (ASR)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
121cybersecurityFORTRESS (ARS)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports an attack-risk score as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
122agentsAgentHarm (ASR)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
123agentsAgentDojo (pass@1 ASR)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
124agentsGraySwan ART (pass@1 ASR)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Lower is better. The source reports an attack success rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
125agentsOR-Bench (FRR)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Lower is better. The source reports a false-refusal rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
126cybersecurityCyber Misuse Chat (FRR)+
“Can the model diagnose or complete difficult cybersecurity tasks?”
- Measures
- Security analysis or exploitation under a controlled evaluation environment.
- Read it
- Lower is better. The source reports a false-refusal rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
127agentsAgentHarm Verified (benign) (FRR)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Lower is better. The source reports a false-refusal rate as a percentage.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
128agentsSAVE-Bench+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
129reasoningInternal Sycophancy+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
130reasoningDeceptionBench+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
131reasoningHLE Calibration+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
132codingCursorBench-3+
“Can the coding agent solve realistic, underspecified tasks drawn from Cursor's own engineering work?”
- Measures
- Correctness on private, long-horizon software tasks run in Cursor's production-like harness.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
133codingSWE-bench Multilingual+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
134agentsTerminal-Bench · Cursor report+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
135scienceCritPt+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
136reasoningAIME 2026+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
137reasoningHMMT · November 2025+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
138reasoningHMMT · February 2026+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
139reasoningIMOAnswerBench+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
140codingNL2Repo+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
141agentsTerminal-Bench 2.1 · Terminus-2+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
142agentsTerminal-Bench 2.1 · best reported harness+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
143agentsMCP-Atlas · public set+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
144agentsTool-Decathlon+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
145codingSWE-fficiency+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
146codingKernelBench Hard+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
147codingPostTrainBench · normalized score+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. This benchmark reports points rather than percent correct.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
148codingMulti-SWE-Bench+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
149codingVIBE-Pro · average+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
150agentsBrowseComp · context management+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
151agentsWide Search+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
152agentsRISE+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
153agentsBFCL multi-turn+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
154reasoningAIME 2025+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
155reasoningIFBench+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
156reasoningHumanity's Last Exam · full set+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
157codingKimi Code Bench v2 (Internal)+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
158codingMLS Bench Lite+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
159agentsKimi Claw 24/7 Bench (Internal)+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
160agentsMCP Mark Verified+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
161agentsTerminal-Bench Hard+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
162codingLiveCodeBench v6+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
163scienceAstaBench · overall+
“Can the model reason accurately about difficult scientific material?”
- Measures
- Scientific knowledge and reasoning under the benchmark's published question set.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
164reasoningHMMT 2025 · pass@1+
“Can the model solve difficult problems that require more than factual recall?”
- Measures
- Multi-step reasoning on academic, mathematical, or abstract problems.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
165codingCodeforces · rating+
“Can the model complete substantial programming work under this benchmark's agent setup?”
- Measures
- Software implementation, debugging, or repository work under the published evaluation protocol.
- Read it
- Higher is better. This benchmark reports points rather than percent correct.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
166agentsτ²-bench · pass@1+
“Can the model plan, use tools, and finish a multi-step task?”
- Measures
- Agentic execution across tools, environments, or long-running workflows.
- Read it
- Higher is better. The value is the percentage reported in this lab's table.
- Boundary
- Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.