Definitions by family
Each family keeps its versions and reported measures together without treating them as the same test.
scienceAAV Capsid Packaging Prediction1 version / measure
AAV Capsid Packaging Prediction · Spearman rank correlation
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Spearman rank correlation reported by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceABC Bench (Fragment Design)1 version / measure
ABC Bench (Fragment Design)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceABC Bench (Liquid Handling)1 version / measure
ABC Bench (Liquid Handling)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceABC Bench (Screening Evasion)1 version / measure
ABC Bench (Screening Evasion)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceAdvanced Screening Evasion1 version / measure
Advanced Screening Evasion
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score reported by OpenAI for the helpful-only checkpoint.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsAgentDojo (pass@1 ASR)1 version / measure
AgentDojo (pass@1 ASR)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAgentHarm (ASR)1 version / measure
AgentHarm (ASR)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAgentHarm Verified (benign) (FRR)1 version / measure
AgentHarm Verified (benign) (FRR)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAgentic Misalignment1 version / measure
Agentic Misalignment
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workAgents’ Last Exam4 versions / measures
Agents' Last Exam · average partial-credit score
Can the model complete long-running professional workflows across many occupations?
Average partial-credit performance on end-to-end agent work across 55 professional subdomains in reproducible desktop sandboxes. Higher is better. Average partial-credit score expressed as a percentage.
ALE-V1 is a living benchmark. Model, harness, effort variant, run count, and task coverage all matter; do not compare this average score with ALE's perfect-run pass rate as if they were the same measurement.
Method / sourceAgents' Last Exam · perfect-run pass rate
How often does the model-and-agent configuration complete an entire professional workflow perfectly?
The share of ALE runs that earn a perfect score across long-running professional workflows in reproducible desktop sandboxes. Higher is better. Percentage of runs that earned a perfect score.
ALE-V1 is a living benchmark. Model, harness, effort variant, run count, and task coverage all matter; do not compare this pass rate with ALE's average-score metric as if they were the same measurement.
Method / sourceAgents' Last Exam · OpenAI reported score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAgents' Last Exam
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAgentWorldBench1 version / measure
AgentWorldBench · overall rubric mean
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workAGIEval1 version / measure
AGIEval · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionAI2D_TEST1 version / measure
AI2D_TEST
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningAIME2 versions / measures
AIME 2026
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAIME 2025
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningAIME261 version / measure
AIME26
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAIRS-Bench1 version / measure
AIRS-Bench
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAISE-Bench10 versions / measures
AISE-Bench · Answer Content · Completeness
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · Answer Content · Correctness
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · Answer Content · F1-LM
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · Answer Content · Faithful
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · API-based Judge · Paragraph Accuracy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · API-based Judge · Success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · References and Formatting · Edit Dist.
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. Raw edit distance printed by the repository.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · References and Formatting · Format
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · References and Formatting · Precision
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAISE-Bench · References and Formatting · Recall
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionAndroidWorld1 version / measure
AndroidWorld
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningApex1 version / measure
Apex · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningApex Shortlist1 version / measure
Apex Shortlist · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningARC-AGI3 versions / measures
ARC-AGI-3
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.
ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.
Method / sourceARC-AGI-2
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.
ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.
Method / sourceARC-AGI-1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.
ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.
Method / sourcescienceAstaBench1 version / measure
AstaBench · overall
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAutomation-Bench1 version / measure
Automation-Bench
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsAutomationBench14 versions / measures
AutomationBench · September 2026 launch table
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAutomationBench · public 600-task set
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's public table reports strict task-completion pass rate as a percentage.
This is the public 600-task table, not the private held-out page leaderboard. Exact source label, highest available reasoning effort, repository revision, and the public task-set definition remain material; no cost is inferred.
Method / sourceAutomationBench · official private leaderboard
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official strict task-completion rate is expressed as a percentage.
This is the private held-out leaderboard, not the public 600-task README table. Exact model label, reasoning effort, page version, row rank, cost markers, fallback behavior, and the private task-set revision remain material; no rank is treated as a performance metric.
Method / sourceAutomationBench · Finance domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Finance domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · HR domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official HR domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · Marketing domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Marketing domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · Operations domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Operations domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · Sales domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Sales domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · Support domain summary
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Support domain summary score is a strict task-completion percentage.
This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.
Method / sourceAutomationBench · Anthropic release comparison
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAutomationBench · OpenAI launch comparison
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAutomationBench · private set
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAutomationBench Public
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceAutomationBench · business workflows
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionBabyVision1 version / measure
BabyVision · with tools
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionBabyVision with python1 version / measure
BabyVision with python
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningBBH1 version / measure
BBH · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionBenchCAD1 version / measure
BenchCAD
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionBenchCAD (python tool)1 version / measure
BenchCAD (python tool)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsBFCL1 version / measure
BFCL v4 · success after LWM RL
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsBFCL multi-turn1 version / measure
BFCL multi-turn
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workBig Finance Bench1 version / measure
Big Finance Bench
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingBigCodeBench1 version / measure
BigCodeBench · Pass@1
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceBioDesign Tools (avg)1 version / measure
BioDesign Tools (avg)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceBioMysteryBench4 versions / measures
BioMysteryBench · human difficult
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBioMysteryBench · human solvable
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBioMysteryBench · hard
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBioMysteryBench · human solved
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionBlueprint-Bench1 version / measure
Blueprint-Bench 2
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsBrowseComp4 versions / measures
BrowseComp · agentic search
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBrowseComp
Can the model find hard-to-locate facts through multi-step web research?
Persistent browsing, source discovery, and synthesis for intentionally difficult questions. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBrowseComp · Pass@1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceBrowseComp · context management
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workC-Eval2 versions / measures
C-Eval · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceC-Eval
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCapture-the-Flag Challenges1 version / measure
Capture-the-Flag Challenges
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCatastrophic Cyber Misuse (ASR)1 version / measure
Catastrophic Cyber Misuse (ASR)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionCC-OCR1 version / measure
CC-OCR
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionChartography1 version / measure
Chartography · with tools
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionCharXiv2 versions / measures
CharXiv · RQ
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCharXiv · RQ with python
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionCharXiv Reasoning (no tools)1 version / measure
CharXiv Reasoning (no tools)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionCharXiv Reasoning (with tools)1 version / measure
CharXiv Reasoning (with tools)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningChinese-SimpleQA1 version / measure
Chinese-SimpleQA · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsClaw-Eval4 versions / measures
Claw-Eval · Orchard-Claw
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceClaw-Eval · Orchard-Claw with ZeroClaw
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceClaw-Eval · success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceClaw-Eval · success after LWM RL
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingClaw-Eval Avg1 version / measure
Claw-Eval Avg
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingClaw-Eval Pass^31 version / measure
Claw-Eval Pass^3
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningCLUEWSC1 version / measure
CLUEWSC · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningCMath1 version / measure
CMath · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workCMMLU1 version / measure
CMMLU · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingCode Arena1 version / measure
Code Arena · Web development · Elo
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingCodeforces2 versions / measures
Codeforces · rating
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCodeforces · rating
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsCodex harness2 versions / measures
Codex harness · Orchard-trained Orchard-Claw
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCodex harness · untrained Orchard-Claw
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsContinuous progress classification1 version / measure
Continuous progress classification · accuracy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceCoronavirus–ACE2 Cell-Entry Screen2 versions / measures
Coronavirus–ACE2 Cell-Entry Screen · composite score
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Composite score reported by OpenAI for held-out cell-entry measurements.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceCoronavirus–ACE2 Cell-Entry Screen · helpful-only composite score
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Composite score for OpenAI's helpful-only Astra checkpoint.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcelong contextCorpusQA 1M1 version / measure
CorpusQA 1M · ACC
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionCountBench1 version / measure
CountBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceCritPt1 version / measure
CritPt
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsCross-embodiment skill transfer1 version / measure
Cross-embodiment skill transfer · average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCurated CTFs (pass@1)1 version / measure
Curated CTFs (pass@1)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingCursorBench4 versions / measures
CursorBench 4.0
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cursor's official CursorBench 4.0 score is a percentage.
Effort and harness are model-specific. The page warns that results have variance; small score differences may not be statistically meaningful. Costs are the publisher's pricing-based average, not observed billed amounts.
Method / sourceCursorBench 4.0 · Medium Opus 5.5 comparison
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCursorBench 3.2.0
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCursorBench-3
Can the coding agent solve realistic, underspecified tasks drawn from Cursor's own engineering work?
Correctness on private, long-horizon software tasks run in Cursor's production-like harness. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecybersecurityCybench (pass@1)1 version / measure
Cybench (pass@1)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyber Misuse Chat (ASR)1 version / measure
Cyber Misuse Chat (ASR)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyber Misuse Chat (FRR)1 version / measure
Cyber Misuse Chat (FRR)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports a false-refusal rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyber range exercises1 version / measure
Cyber range exercises · combined pass rate
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyberGym2 versions / measures
CyberGym
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceCyberGym · success rate
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyberGym (pass@1)1 version / measure
CyberGym (pass@1)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityCyScenarioBench (pass@1)1 version / measure
CyScenarioBench (pass@1)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningDeceptionBench1 version / measure
DeceptionBench
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workDECK-Bench (Internal)1 version / measure
DECK-Bench (Internal)
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingDeepPlanning1 version / measure
DeepPlanning
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsDeepSearchQA1 version / measure
DeepSearchQA · F1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsDeepShop1 version / measure
DeepShop · Orchard-GUI
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingDeepSWE4 versions / measures
DeepSWE v1.1 · pass@1
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of scored rollout attempts that passed.
Pass@1 is an attempt-level rate. Task coverage, repeated runs, confidence interval, mini-swe-agent version, model variant, and reasoning effort remain material.
Method / sourceDeepSWE v1.1 · pass@4
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of attempted tasks with at least one passing rollout.
Pass@4 is a task-level rate, not the same measurement as pass@1. Task coverage, repeated runs, confidence interval, mini-swe-agent version, model variant, and reasoning effort remain material.
Method / sourceDeepSWE v1.1 · agentic coding
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDeepSWE
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceDNA sequence design for transcription factor binding1 version / measure
DNA sequence design for transcription factor binding · pass@1
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceDNA sequence design for transcription-factor binding1 version / measure
DNA sequence design for transcription-factor binding
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Win rate over the Ledidi baseline reported by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsDOMINO2 versions / measures
DOMINO · manipulation score
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDOMINO · success rate
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsDreamGen Bench1 version / measure
DreamGen Bench · Total
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsDreamGen Bench component6 versions / measures
DreamGen Bench component · GR1-Behavior IF
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDreamGen Bench component · GR1-Behavior PA
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDreamGen Bench component · GR1-Env IF
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDreamGen Bench component · GR1-Env PA
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDreamGen Bench component · GR1-Object IF
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDreamGen Bench component · GR1-Object PA
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningDROP1 version / measure
DROP · F1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingDSBench-FullStack †1 version / measure
DSBench-FullStack †
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingDSBench-Hard †1 version / measure
DSBench-Hard †
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcevisionDynaMath1 version / measure
DynaMath
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyDynamic adversarial user simulations3 versions / measures
Dynamic adversarial user simulations · Emotional reliance
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDynamic adversarial user simulations · Mental health
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceDynamic adversarial user simulations · Self-harm
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsEBench success rate4 versions / measures
EBench success rate · LongHorizon
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench success rate · Overall
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench success rate · SimplePnP
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench success rate · TableTop
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsEBench task score4 versions / measures
EBench task score · LongHorizon
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench task score · Overall
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench task score · SimplePnP
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEBench task score · TableTop
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceEEBench1 version / measure
EEBench
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionEmbSpatialBench1 version / measure
EmbSpatialBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionERQA1 version / measure
ERQA
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsEWMBench1 version / measure
EWMBench · Overall
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsEWMBench component8 versions / measures
EWMBench component · BLEU
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · CLIP
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · Diversity
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · Dyn
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · HSD
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · Logics
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · nDTW
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceEWMBench component · SceneC
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityExploitBench4 versions / measures
ExploitBench · Internal Port success rate
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceExploitBench · Cap Percent
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. OpenAI's Cap Percent averages demonstrated capability flags across 41 V8 vulnerabilities.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceExploitBench v8-bench · capability coverage
How far can the agent progress from reaching a vulnerable V8 code path to arbitrary code execution?
Sixteen deterministically graded capabilities on patched, previously vulnerable V8 builds. Higher is better. Percentage of the 16 capability flags achieved after OR-merging the selected seeds within every evaluated environment.
Control and all are different measurement bases. Agent, AutoNudge state, selected run, turn budget, seed count, failed cells, and environment coverage can all change the value.
Method / sourceExploitBench v8-bench · mean capability score
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Seed-weighted mean of the source's 0–16 per-run capability score across selected environments.
This is not capability coverage. Control and all select runs differently, and the all view may mix turn budgets and seed counts across configurations.
Method / sourcecybersecurityExploitGym2 versions / measures
ExploitGym · reported challenge success
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceExploitGym
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityExploitGym (pass@1)1 version / measure
ExploitGym (pass@1)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workFACTS Parametric1 version / measure
FACTS Parametric · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsFew-shot skill transfer5 versions / measures
Few-shot skill transfer · average success · Fold Towel
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFew-shot skill transfer · average success · Insert Screw
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFew-shot skill transfer · average success · Put Blocks
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFew-shot skill transfer · average success · Put Fruits
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFew-shot skill transfer · average success · Unscrew Cap
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workFinance Agent1 version / measure
Finance Agent v2
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityFORTRESS (ARS)1 version / measure
FORTRESS (ARS)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack-risk score as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyFrontier Red Team16 versions / measures
Frontier Red Team · identity classification · easy
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the easy tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity classification · hard
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the hard tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity classification · medium
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the medium tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity correlation · easy
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the easy tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity correlation · hard
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the hard tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity correlation · medium
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the medium tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · identity correlation · pooled
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Pooled linkage F1 for identity correlation across the 50 reported samples.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · text geolocation · median home-location error
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Lower is better. Median home-location error in kilometers for text geolocation.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · photo geolocation · median distance error
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Lower is better. Median distance error in kilometers for photo geolocation.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · photo geolocation · within 1 km
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of photo-geolocation predictions within 1 kilometer.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · terminal guidance · parked high-visibility target
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Simulated strike rate for terminal guidance to a parked high-visibility target.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · terminal guidance · road-speed target
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Simulated strike rate for terminal guidance to a vehicle moving at road speed.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · payload delivery · target moving at 3 m/s
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a target moving at approximately 3 m/s.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · payload delivery · camouflaged zig-zag target
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a camouflaged zig-zag target.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · payload delivery · hard wind and moving target
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters under the hard wind-and-moving-target setting.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourceFrontier Red Team · payload delivery · stationary target within 5 m
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a stationary target.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Method / sourcecodingFrontier-Bench1 version / measure
Frontier-Bench v0.1 · agentic terminal coding
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingFrontierCode9 versions / measures
FrontierCode 1.1 · Main
Would a repository maintainer merge this software change?
End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Cognition's Score is a weighted aggregate of rubric items; solutions that fail a blocking criterion receive 0.
Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.
Method / sourceFrontierCode 1.1 · Extended
Would a repository maintainer merge this software change?
End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Cognition's Score is a weighted aggregate of rubric items; solutions that fail a blocking criterion receive 0.
Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.
Method / sourceFrontierCode 1.1 · Extended · Pass Rate
Would a repository maintainer merge this software change?
End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Pass Rate is the fraction of trials satisfying every blocker rubric criterion, averaged over the selected tasks.
Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.
Method / sourceFrontierCode 1.1 · Main · Pass Rate
Would a repository maintainer merge this software change?
End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Pass Rate is the fraction of trials satisfying every blocker rubric criterion, averaged over the selected tasks.
Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.
Method / sourceFrontierCode 1.1 · Main · Medium Opus 5.5 comparison
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFrontierCode v1.1 · Main
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFrontierCode 1.0 · Diamond
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceFrontierCode 1.0 · Extended
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cognition reports FrontierCode score as a percentage, not pass rate.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceFrontierCode 1.0 · Main
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cognition reports FrontierCode score as a percentage, not pass rate.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcereasoningFrontierMath Tier 1-3 (v2)1 version / measure
FrontierMath Tier 1-3 (v2)
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningFrontierMath Tier 4 (v2)1 version / measure
FrontierMath Tier 4 (v2)
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingFrontierSWE1 version / measure
FrontierSWE
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsGDM Situational Awareness1 version / measure
GDM Situational Awareness
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsGDM-Stealth (of 4)1 version / measure
GDM-Stealth (of 4)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisiongdp.pdf1 version / measure
gdp.pdf
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceGeneBench Pro1 version / measure
GeneBench Pro
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceGPQA4 versions / measures
GPQA Diamond
Can the model reason through specialist biology, physics, and chemistry questions?
Multiple-choice science questions written and validated by domain experts. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGPQA Diamond · Pass@1
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGPQA Diamond
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGPQA
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextGraphWalks BFS2 versions / measures
GraphWalks BFS · 1M F1
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGraphWalks BFS · 256K F1
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsGraySwan ART (pass@1 ASR)1 version / measure
GraySwan ART (pass@1 ASR)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsGripper dexterity3 versions / measures
Gripper dexterity · diverse tool kitting
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGripper dexterity · general pick and place
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceGripper dexterity · precise insertion tasks
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningGSM8K1 version / measure
GSM8K · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionHallusionBench1 version / measure
HallusionBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceHard-negative protein binding prediction1 version / measure
Hard-negative protein binding prediction · pass@4
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceHard-negative protein binding prediction2 versions / measures
Hard-negative protein binding prediction
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Percentage of protein-binding selections scored correct by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceHard-negative protein binding prediction · helpful-only
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Percentage of protein-binding selections scored correct by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceknowledge workHarvey's Legal Agent Benchmark1 version / measure
Harvey's Legal Agent Benchmark
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workHealthBench5 versions / measures
HealthBench · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench Consensus · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench Hard · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench Professional · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workHealthBench Consensus1 version / measure
HealthBench Consensus · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workHealthBench Hard1 version / measure
HealthBench Hard · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workHealthBench Pro1 version / measure
HealthBench Pro
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workHealthBench Professional3 versions / measures
HealthBench Professional
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench Professional · length-adjusted score
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHealthBench Professional · health
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHellaSwag1 version / measure
HellaSwag · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHMMT3 versions / measures
HMMT 2025 · pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHMMT · February 2026
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHMMT · November 2025
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHMMT 2026 Feb1 version / measure
HMMT 2026 Feb · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHMMT Feb2 versions / measures
HMMT Feb 26
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHMMT Feb 25
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHMMT Nov1 version / measure
HMMT Nov 25
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceHPCT1 version / measure
HPCT
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingHumanEval1 version / measure
HumanEval · Pass@1
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningHumanity’s Last Exam14 versions / measures
Humanity's Last Exam · full
Can the model answer deliberately difficult expert questions across academic fields?
Broad expert-level reasoning and knowledge without external tools. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · text only
Can the model answer the text-only portion of deliberately difficult expert questions?
The publisher's finalized HLE text-only view, preserving the official score, confidence interval, calibration error, and contamination warning fields. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · September 2026 launch table · no tools
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · September 2026 launch table · with tools
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · full with tools
How much do search or code tools help on Humanity's Last Exam?
The full expert-question set with the reporting lab's stated tool setup. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE-Verified · full 1,811-item set
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE · no tools
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE · with tools
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE Calibration
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE with tools · Pass@1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHLE
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · full set
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceHumanity's Last Exam · search + code
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningIFBench1 version / measure
IFBench
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyImage input evaluations4 versions / measures
Image input evaluations · extremism
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceImage input evaluations · harms-erotic
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceImage input evaluations · hate
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceImage input evaluations · self-harm
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningIMOAnswerBench2 versions / measures
IMOAnswerBench
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceIMOAnswerBench · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsIn-distribution manipulation3 versions / measures
In-distribution manipulation · LIBERO
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceIn-distribution manipulation · RT-Easy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceIn-distribution manipulation · RT-Hard
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityInternal Capture-the-Flag challenges1 version / measure
Internal Capture-the-Flag challenges · defender success
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceknowledge workInternal Company Research Reports1 version / measure
Internal Company Research Reports · accepted reports
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. This benchmark reports points rather than percent correct.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingInternal Research Debugging Evaluation1 version / measure
Internal Research Debugging Evaluation
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingInternal Research Debugging Evaluation1 version / measure
Internal Research Debugging Evaluation
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of internal research-debugging tasks solved under OpenAI's evaluation.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcereasoningInternal Sycophancy1 version / measure
Internal Sycophancy
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecybersecurityJailbreak StrongREJECT v2 (ASR)1 version / measure
Jailbreak StrongREJECT v2 (ASR)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workJob Bench1 version / measure
Job Bench
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingKernelBench Hard1 version / measure
KernelBench Hard
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingKernelGen 1P1 version / measure
KernelGen 1P
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsKimi Claw 24/7 Bench (Internal)1 version / measure
Kimi Claw 24/7 Bench (Internal)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingKimi Code Bench 2.0 (Internal)1 version / measure
Kimi Code Bench 2.0 (Internal)
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingKimi Code Bench v2 (Internal)1 version / measure
Kimi Code Bench v2 (Internal)
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceLABBench21 version / measure
LABBench2
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workLegal Agent Benchmark2 versions / measures
Legal Agent Benchmark · Held-out
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLegal Agent Benchmark
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsLIBERO-Plus OOD robustness8 versions / measures
LIBERO-Plus OOD robustness · Background
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Camera
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Language
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Layout
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Light
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Noise
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Robot
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLIBERO-Plus OOD robustness · Total
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceLifeSciBench1 version / measure
LifeSciBench
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingLiveCodeBench3 versions / measures
LiveCodeBench v6
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Arithmetic mean of the official per-problem pass@1 values in the selected date window.
The selected 454-problem window is the official page's default date range, not the full 1,055-problem release_v6 set. Exact source label, model metadata, release date, contamination flag, difficulty counts, and date window remain material; no cost is inferred.
Method / sourceLiveCodeBench v6
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLiveCodeBench · Pass@1
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingLiveCodeBench Pro1 version / measure
LiveCodeBench Pro · Elo
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextLongBench1 version / measure
LongBench-V2 · EM
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionLVBench3 versions / measures
LVBench · agentic
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLVBench · static
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceLVBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workManagement Consulting Tasks (Internal)1 version / measure
Management Consulting Tasks (Internal)
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsManipulation benchmark5 versions / measures
Manipulation benchmark · LIBERO
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceManipulation benchmark · RoboCasa-GR1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceManipulation benchmark · RoboTwin-Easy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceManipulation benchmark · RoboTwin-Hard
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceManipulation benchmark · Simpler-WidowX
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsMASK1 version / measure
MASK
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningMATH1 version / measure
MATH · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningMathArena Apex1 version / measure
MathArena Apex
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMathVision1 version / measure
MathVision
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMathVision with python1 version / measure
MathVision with python
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMathvista(mini)1 version / measure
Mathvista(mini)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceMBCT1 version / measure
MBCT
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsMCP Atlas3 versions / measures
MCP Atlas
Can the model discover and coordinate the right tools to complete a multi-step workflow?
One thousand human-authored tasks spanning 36 real MCP servers and 220 tools. Higher is better. Scale's published MCP Atlas score field is retained in percentage points.
This row set uses the April 2026 methodology revision and its 1,000-task composition. Exact source label, version, effort label, rank, confidence interval, contamination warning, publication flags, and entry timestamp remain material; no model-specific tool setup or cost is inferred.
Method / sourceMCP-Atlas · public set
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMCP-Atlas
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsMCP Mark Verified1 version / measure
MCP Mark Verified
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsMCPAtlas Public1 version / measure
MCPAtlas Public · Pass@1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsMCPMark2 versions / measures
MCPMark · success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMCPMark
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceMedChemBench (Internal)1 version / measure
MedChemBench (Internal)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcereasoningMGSM1 version / measure
MGSM · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingMLE-Bench1 version / measure
MLE-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingMLS Bench1 version / measure
MLS Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingMLS Bench Lite1 version / measure
MLS Bench Lite
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMLVU1 version / measure
MLVU
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMMBench EN-DEV1 version / measure
MMBench EN-DEV-v1.1
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workMMLU5 versions / measures
MMLU-Pro · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMLU · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMLU-Redux · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMLU-Pro
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMLU-Redux
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workMMMLU2 versions / measures
MMMLU · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMMLU
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMMMU5 versions / measures
MMMU-Pro · selected exact setups
Can the model combine visual evidence with college-level knowledge and deliberate reasoning?
Multidiscipline multimodal questions filtered for genuine visual dependence under a no-tools protocol. Higher is better. MMMU Team's pro.overall percentage under the official zero-shot evaluation.
Tool access, exact source setup label, model publication date, model size, author-provided marker, and the maintainer's selected column are material. Do not mix this row with validation, test, or provider-reported values.
Method / sourceMMMU-Pro · with tools
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. MMMU Team's pro.overall percentage for the exact with-tools setup label.
This is a separate tool-enabled setup from no-tools rows. Exact source label, model publication date, model size, author-provided marker, and the maintainer's selected column remain material.
Method / sourceMMMU-Pro with python
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMMU
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMMMU-Pro
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMMStar1 version / measure
MMStar
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextMRCR2 versions / measures
MRCR v2 · 8-needle · 128K average
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMRCR v2 · 8-needle · 1M pointwise
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextMRCR 1M1 version / measure
MRCR 1M · MMR
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextMRCR Long Context (1M context window)1 version / measure
MRCR Long Context (1M context window)
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityMSRC historical cases2 versions / measures
MSRC historical cases · clfs.sys recall
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceMSRC historical cases · tcpip.sys recall
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsMulti-finger dexterity5 versions / measures
Multi-finger dexterity · dustpan
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMulti-finger dexterity · screw bulb
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMulti-finger dexterity · tie trash bag
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMulti-finger dexterity · unscrew bulb
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceMulti-finger dexterity · Ziplock
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingMulti-SWE-Bench1 version / measure
Multi-SWE-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workMultiLoKo1 version / measure
MultiLoKo · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionMVBench1 version / measure
MVBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingNanoGPT1 version / measure
NanoGPT
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingNL2Repo1 version / measure
NL2Repo
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingNL2Repo-Bench1 version / measure
NL2Repo-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionOCRBench1 version / measure
OCRBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionODInW131 version / measure
ODInW13
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workOffice QA Pro1 version / measure
Office QA Pro
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionOmniDocBench1 version / measure
OmniDocBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionOmniDocBench1.51 version / measure
OmniDocBench1.5
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsOnline-Mind2Web1 version / measure
Online-Mind2Web · Orchard-GUI
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcelong contextOpenAI MRCR2 versions / measures
OpenAI MRCR v2 · 8-needle · 256K-512K
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOpenAI MRCR v2 · 8-needle · 512K-1M
Can the model retain and reason over evidence spread across a very long input?
Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsOR-Bench (FRR)1 version / measure
OR-Bench (FRR)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsOrchard-GUI1 version / measure
Orchard-GUI · three-benchmark average
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsOSWorld10 versions / measures
OSWorld 2.0 · partial score
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · offline partial reward
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · August 2026 task release · partial
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · August 2026 task release · strict
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · partial score · batch tool enabled
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · partial score
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · computer use
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld 2.0 · strict binary completion
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceOSWorld-Verified
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. Official OSWorld-Verified success-rate percentage averaged across repeated source rows for the exact Model and Max steps group.
Exact source model label, approach type, institution, maximum steps, source date, 361-or-369 task denominator, per-application category strings, repeated-run count and population standard deviation, augmentation flags, and OSWorld-Verified task revision remain material. Do not merge this workbook with OSWorld 2.0 or provider-reported OSWorld rows.
Method / sourceagentsPBench11 versions / measures
PBench · Aes
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Bg-Con
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Domain
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · I2V-Bg
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · I2V-S
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Img
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Mot
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · O-Con
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Overall
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Quality
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePBench · Sub-Con
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionPerceptionBench1 version / measure
PerceptionBench
Can the model accurately perceive visual evidence before reasoning or outside knowledge is required?
Three thousand verified questions spanning ten atomic visual-perception capabilities. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesciencePhage–plasmid co-evolution2 versions / measures
Phage–plasmid co-evolution · helpful-only negative log-likelihood
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Lower is better. Negative log-likelihood for OpenAI's helpful-only Astra checkpoint.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcePhage–plasmid co-evolution · negative log-likelihood
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Lower is better. Negative log-likelihood of subsequent mutations.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecybersecurityPoly-Guard Bench (ASR)1 version / measure
Poly-Guard Bench (ASR)
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingPostTrain Bench1 version / measure
PostTrain Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingPostTrainBench1 version / measure
PostTrainBench · normalized score
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingPostTrainBench Lite1 version / measure
PostTrainBench Lite
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsPrecision moment-finding2 versions / measures
Precision moment-finding · accuracy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcePrecision moment-finding · mean absolute distance
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. Mean absolute distance in seconds between the predicted and critical video frame.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyProduction Benchmarks8 versions / measures
Production Benchmarks · Extremism
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Gore
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Hate
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Nonviolent illicit behavior
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Self-harm (standard)
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Sexual
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Sexual/minors
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceProduction Benchmarks · Violent Illicit behavior
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingProgramBench1 version / measure
ProgramBench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyPrompt injection attacks in connectors1 version / measure
Prompt injection attacks in connectors
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceProtocolQA1 version / measure
ProtocolQA
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceProtocolQA Open-Ended2 versions / measures
ProtocolQA Open-Ended
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score reported by OpenAI on the open-ended troubleshooting version.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceProtocolQA Open-Ended
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsQwenClawBench3 versions / measures
QwenClawBench · success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceQwenClawBench · success after LWM RL
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceQwenClawBench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingQwenWebBench1 version / measure
QwenWebBench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsR2R-CE2 versions / measures
R2R-CE · validation seen success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceR2R-CE · validation unseen success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsReal-world ALOHA2 versions / measures
Real-world ALOHA · in-domain average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceReal-world ALOHA · out-of-domain average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsReal-world ALOHA dual-arm2 versions / measures
Real-world ALOHA dual-arm · In-domain average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceReal-world ALOHA dual-arm · OOD average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionRealWorldQA1 version / measure
RealWorldQA
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionRefCOCO(avg)1 version / measure
RefCOCO(avg)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionRefSpatialBench1 version / measure
RefSpatialBench
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceRefusals: BioTIER1 version / measure
Refusals: BioTIER
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. The source reports a refusal percentage; interpret with the named safety objective.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceRefusals: Chemical Agents1 version / measure
Refusals: Chemical Agents
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. The source reports a refusal percentage; interpret with the named safety objective.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceReproBAIT1 version / measure
ReproBAIT · published-performance fraction
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Fraction of published performance reached on SecureBio's six-task agentic biology evaluation.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsRISE1 version / measure
RISE
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsRoboCasa3654 versions / measures
RoboCasa365 · Atomic
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboCasa365 · Composite-Seen
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboCasa365 · Composite-Unseen
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboCasa365 · Total
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsRoboChallenge2 versions / measures
RoboChallenge · bimanual coordination average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboChallenge · robust pick-and-place average success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsRoboChallenge Table302 versions / measures
RoboChallenge Table30 v1 · generalist-track success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboChallenge Table30 v1 · process score
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsRoboTwin-Clean2Rand6 versions / measures
RoboTwin-Clean2Rand · Background
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-Clean2Rand · Clutter
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-Clean2Rand · Easy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-Clean2Rand · Hard
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-Clean2Rand · Height
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-Clean2Rand · Light
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsRoboTwin-IF6 versions / measures
RoboTwin-IF · Average
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-IF · Open Microwave Door
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-IF · Open Stapler
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-IF · Operate Table
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-IF · Pick-Diverse
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceRoboTwin-IF · Place-Relative
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingRSI Index1 version / measure
RSI Index
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecybersecuritySandbox Bench1 version / measure
Sandbox Bench · targets exploited
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of private targets on which the model recovered the protected flag.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceagentsSAVE-Bench1 version / measure
SAVE-Bench
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceSciCode2 versions / measures
SciCode · main-problem resolve rate
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Official main-problem resolve rate over the 80 main problems.
This is the repository table's main-problem measurement. The result changelog date is not stated in the pinned README; keep the repository revision, exact source label, and separate subproblem measurement attached.
Method / sourceSciCode · subproblem accuracy
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Official subproblem accuracy over the 338 decomposed subproblems.
This is a separate subproblem measurement from the main-problem resolve rate. The result changelog date is not stated in the pinned README; keep the repository revision and exact source label attached.
Method / sourcecybersecuritySEC-Bench Pro2 versions / measures
SEC-Bench Pro · reported score
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSEC-Bench Pro
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceSecureBio Multimodal Troubleshooting Virology1 version / measure
SecureBio Multimodal Troubleshooting Virology
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Refusal-corrected accuracy reported by OpenAI from SecureBio's evaluation.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceSeqQA (agentic)1 version / measure
SeqQA (agentic)
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsSHADE-Arena1 version / measure
SHADE-Arena
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceSHP2 Protein Function Prediction2 versions / measures
SHP2 Protein Function Prediction · maximum mean R²
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Maximum mean R² across three unpublished assay datasets.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceSHP2 Protein Function Prediction · helpful-only maximum mean R²
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Maximum mean R² for OpenAI's helpful-only Astra checkpoint.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcereasoningSimple-QA verified1 version / measure
Simple-QA verified · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningSimpleQA-Verified1 version / measure
SimpleQA-Verified · Pass@1
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionSimpleVQA1 version / measure
SimpleVQA
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSkillsBench Avg51 version / measure
SkillsBench Avg5
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecuritySocial Engineering1 version / measure
Social Engineering
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workSpreadsheetBench1 version / measure
SpreadsheetBench 2
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSQLite reconstruction1 version / measure
SQLite reconstruction · held-out sqllogictest
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecuritySRE-Bench1 version / measure
SRE-Bench · pass@4
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of reverse-engineering challenges fully solved across four trials.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcesafetyStatic Jailbreak Evaluation5 versions / measures
Static Jailbreak Evaluation · biological high risk defender success
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceStatic Jailbreak Evaluation · biological severe defender success
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceStatic Jailbreak Evaluation · cyber defender success
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceStatic Jailbreak Evaluation · violence moderate defender success
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceStatic Jailbreak Evaluation · violence severe defender success
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecybersecurityStorageDrive1 version / measure
StorageDrive · planted vulnerabilities found
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Number of planted vulnerabilities found out of 21.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceSuperGPQA2 versions / measures
SuperGPQA · EM
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSuperGPQA
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE Marathon1 version / measure
SWE Marathon
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE Multilingual1 version / measure
SWE Multilingual · resolved
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE Pro1 version / measure
SWE Pro · resolved
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE Verified1 version / measure
SWE Verified · resolved
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE-Bench12 versions / measures
SWE-Bench Pro V2 · Public Full
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Resolve Rate on the official 642-task Public Full split across 11 repositories.
Scale's V2 release changes task composition and protocol relative to the 731-task V1 Public set. Compare only rows in this explicit V2 Full view; exact harness, effort, source label, rank, and displayed uncertainty are retained per model.
Method / sourceSWE-Bench Pro · Public
Can an agent resolve difficult software issues from real repositories?
Repository-level issue resolution on the SWE-Bench Pro task set. Higher is better. Resolve Rate score on the official 731-instance Public Set.
This is a benchmark-maintainer leaderboard result with no agent configuration exposed in the embedded entries. Exact source label, effort marker, confidence interval, contamination/deprecation flags, entry timestamp, and task-set revision are material; private and other SWE-Bench Pro splits are not comparable.
Method / sourceSWE-bench Verified
Can an agent fix a real GitHub issue and pass the repository's tests?
The share of engineer-verified software issues resolved by the full model-plus-agent system. Higher is better. Resolved percent on the official verified-500 task set under the bash-only leaderboard block.
This is a model-plus-agent result. Exact agent, model variant, effort, submission date, mini-SWE-agent version, system flags, and task-set revision are material; never compare it with another SWE-bench split or harness as if they were the same test.
Method / sourceSWE-bench Multilingual
Can an agent resolve software issues across multiple programming languages?
The share of 300 multilingual software issues resolved under the benchmark maintainer's published mini-SWE-agent protocol. Higher is better. Resolved percent on the official multilingual-300-9-languages task set under the default leaderboard block.
This is a model-plus-agent result. Exact agent, model variant, effort, submission date, mini-SWE-agent version, system flags, and task-set revision are material; never compare it with another SWE-bench split or harness as if they were the same test.
Method / sourceSWE-Bench Verified · Orchard-SWE Balanced Adaptive Rollout
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-Bench Verified · Orchard-SWE baseline
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-Bench Verified · Orchard-SWE value-model reranking
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-Bench Pro · success
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-Bench Verified · success
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-bench Multilingual
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-bench Pro
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceSWE-bench Verified
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingSWE-fficiency1 version / measure
SWE-fficiency
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceTacit Knowledge and Troubleshooting1 version / measure
Tacit Knowledge and Troubleshooting · refusal-adjusted
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcescienceTacit Knowledge and Troubleshooting2 versions / measures
Tacit Knowledge and Troubleshooting
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Original score reported by OpenAI on an internal multiple-choice evaluation.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceTacit Knowledge and Troubleshooting · refusal-adjusted
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score after counting refusals and safe completions as successes, as reported by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcecodingTAU3-Bench1 version / measure
TAU3-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsTerminal-Bench12 versions / measures
Terminal-Bench 4.0
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 3.0
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.1
Can a model-driven agent complete realistic work inside a terminal?
Multi-step command-line, coding, configuration, debugging, and systems tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal Bench 2.1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.1 · Terminus-2
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.1 · best reported harness
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.0 · success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.0
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal Bench 2.0 · accuracy
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench 2.0 · Codex harness
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench Hard
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench · Cursor report
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceTerminal-Bench Science3 versions / measures
Terminal-Bench-Science 0.1 · official leaderboard
How often an AI agent completes a scientific research workflow using terminal tools.
Resolution rate across 70 research workflows in life, physical, earth, mathematical and engineering sciences, with three trials per task. Higher is better. The maintainer reports the percentage of scored trials resolved successfully.
This is a small, early research-workflow suite, not a general measure of scientific discovery. Model and agent setups differ. Retain source standard errors and exact dataset revision; never mix these official results with provider reports or other benchmark versions.
Method / sourceTerminal-Bench-Science 0.1 · Anthropic Claude Code comparison
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTerminal-Bench-Science 0.1
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsTool-Decathlon2 versions / measures
Tool Decathlon · success
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTool-Decathlon
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsToolathlon2 versions / measures
Toolathlon
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceToolathlon · Pass@1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsToolathlon-Verified1 version / measure
Toolathlon-Verified
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workTriviaQA1 version / measure
TriviaQA · EM
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceTroubleshootingBench2 versions / measures
TroubleshootingBench
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Original score reported by OpenAI on expert-written troubleshooting questions.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceTroubleshootingBench · refusal-adjusted
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score after counting refusals and safe completions as successes, as reported by OpenAI.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourceknowledge workTutorMoments3 versions / measures
TutorMoments · appropriate rigor
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of annotated moments where the model made the appropriate tutoring move.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTutorMoments · appropriate scaffolding
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of annotated moments where the model made the appropriate tutoring move.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceTutorMoments · avoids over-scaffolding
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of moments where the model avoided unnecessary scaffolding.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcesafetyU186 versions / measures
U18 · Age-restricted goods, services, and dangerous challenges/activities
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceU18 · Eating Disorders
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceU18 · Emotional Reliance
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceU18 · Gore
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceU18 · Self Harm
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceU18 · Sexual Content
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionV*1 version / measure
V*
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecybersecurityV8 JavaScript Engine1 version / measure
V8 JavaScript Engine · unique confirmed issues
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Count of unique confirmed vulnerabilities found across the source's fixed number of invocations.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceknowledge workVals Finance Agent1 version / measure
Vals Finance Agent v2
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceVCT1 version / measure
VCT
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingVIBE-Pro1 version / measure
VIBE-Pro · average
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / sourcevisionVideoMME(w sub.)1 version / measure
VideoMME(w sub.)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionVideoMME(w/o sub.)1 version / measure
VideoMME(w/o sub.)
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionVideoMMMU1 version / measure
VideoMMMU
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsVision-language navigation4 versions / measures
Vision-language navigation · R2R Val-Unseen · Oracle Success Rate
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceVision-language navigation · R2R Val-Unseen · Success Rate
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceVision-language navigation · RxR Val-Unseen · SPL
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceVision-language navigation · RxR Val-Unseen · Success Rate
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcecodingVITA-Bench1 version / measure
VITA-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionVlmsAreBlind1 version / measure
VlmsAreBlind
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWebArena-Verified1 version / measure
WebArena-Verified
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWebVoyager1 version / measure
WebVoyager · Orchard-GUI
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWhole-body manipulation3 versions / measures
Whole-body manipulation · pick up from floor
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWhole-body manipulation · pick up from shelf
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWhole-body manipulation · pick up from table
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWide Search1 version / measure
Wide Search
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWideSearch4 versions / measures
WideSearch · F1 by Item
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWideSearch · F1 by Row
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWideSearch · F1 by Item after LWM RL
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWideSearch
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcereasoningWinoGrande1 version / measure
WinoGrande · EM
Can the model solve difficult problems that require more than factual recall?
Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceWMDP-Bio1 version / measure
WMDP-Bio
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcescienceWMDP-Chem1 version / measure
WMDP-Chem
Can the model reason accurately about difficult scientific material?
Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWorldModelBench3 versions / measures
WorldModelBench · Instruction (0-3)
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench · Phys.
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench · Total
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsWorldModelBench physics component7 versions / measures
WorldModelBench physics component · Fluid
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Frame
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Grav.
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Mass
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Newton
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Penetr.
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceWorldModelBench physics component · Temp
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionWorldVQA ForceAnswer1 version / measure
WorldVQA ForceAnswer
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsZero-shot cross-embodiment4 versions / measures
Zero-shot cross-embodiment · ARX
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceZero-shot cross-embodiment · Franka
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceZero-shot cross-embodiment · Total
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceZero-shot cross-embodiment · UR5
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionZeroBench main with python1 version / measure
ZeroBench main with python · pass@5
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionZEROBench_sub1 version / measure
ZEROBench_sub
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourcevisionZeroBench-main2 versions / measures
ZeroBench-main · with tools
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceZeroBench main · pass@5
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceagentsτ²-bench3 versions / measures
τ²-bench · Telecom
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceτ²-bench · Retail
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / sourceτ²-bench · pass@1
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source