Definitions by family

Each family keeps its versions and reported measures together without treating them as the same test.

scienceAAV Capsid Packaging Prediction1 version / measure
Open rankings and effort

AAV Capsid Packaging Prediction · Spearman rank correlation

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Spearman rank correlation reported by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceABC Bench (Fragment Design)1 version / measure
Open rankings and effort

ABC Bench (Fragment Design)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceABC Bench (Liquid Handling)1 version / measure
Open rankings and effort

ABC Bench (Liquid Handling)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceABC Bench (Screening Evasion)1 version / measure
Open rankings and effort

ABC Bench (Screening Evasion)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceAdvanced Screening Evasion1 version / measure
Open rankings and effort

Advanced Screening Evasion

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score reported by OpenAI for the helpful-only checkpoint.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsAgentDojo (pass@1 ASR)1 version / measure
Open rankings and effort

AgentDojo (pass@1 ASR)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAgentHarm (ASR)1 version / measure
Open rankings and effort

AgentHarm (ASR)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAgentHarm Verified (benign) (FRR)1 version / measure
Open rankings and effort

AgentHarm Verified (benign) (FRR)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAgentic Misalignment1 version / measure
Open rankings and effort

Agentic Misalignment

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workAgents’ Last Exam4 versions / measures
Open rankings and effort

Agents' Last Exam · average partial-credit score

Can the model complete long-running professional workflows across many occupations?

Average partial-credit performance on end-to-end agent work across 55 professional subdomains in reproducible desktop sandboxes. Higher is better. Average partial-credit score expressed as a percentage.

ALE-V1 is a living benchmark. Model, harness, effort variant, run count, and task coverage all matter; do not compare this average score with ALE's perfect-run pass rate as if they were the same measurement.

Method / source

Agents' Last Exam · perfect-run pass rate

How often does the model-and-agent configuration complete an entire professional workflow perfectly?

The share of ALE runs that earn a perfect score across long-running professional workflows in reproducible desktop sandboxes. Higher is better. Percentage of runs that earned a perfect score.

ALE-V1 is a living benchmark. Model, harness, effort variant, run count, and task coverage all matter; do not compare this pass rate with ALE's average-score metric as if they were the same measurement.

Method / source

Agents' Last Exam · OpenAI reported score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Agents' Last Exam

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAgentWorldBench1 version / measure
Open rankings and effort

AgentWorldBench · overall rubric mean

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workAGIEval1 version / measure
Open rankings and effort

AGIEval · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionAI2D_TEST1 version / measure
Open rankings and effort

AI2D_TEST

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningAIME2 versions / measures
Open rankings and effort

AIME 2026

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AIME 2025

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningAIME261 version / measure
Open rankings and effort

AIME26

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAIRS-Bench1 version / measure
Open rankings and effort

AIRS-Bench

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAISE-Bench10 versions / measures
Open rankings and effort

AISE-Bench · Answer Content · Completeness

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · Answer Content · Correctness

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · Answer Content · F1-LM

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · Answer Content · Faithful

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · API-based Judge · Paragraph Accuracy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · API-based Judge · Success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · References and Formatting · Edit Dist.

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. Raw edit distance printed by the repository.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · References and Formatting · Format

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · References and Formatting · Precision

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AISE-Bench · References and Formatting · Recall

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionAndroidWorld1 version / measure
Open rankings and effort

AndroidWorld

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningApex1 version / measure
Open rankings and effort

Apex · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningApex Shortlist1 version / measure
Open rankings and effort

Apex Shortlist · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningARC-AGI3 versions / measures
Open rankings and effort

ARC-AGI-3

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.

ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.

Method / source

ARC-AGI-2

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.

ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.

Method / source

ARC-AGI-1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The official ARC score ratio is converted to the catalog percent unit by multiplying by 100.

ARC-AGI version, semi-private split, exact source model label, reasoning setting, provider, release date, display status, and cost field are material. ARC-AGI-1, ARC-AGI-2, and ARC-AGI-3 must remain separate contexts.

Method / source
scienceAstaBench1 version / measure
Open rankings and effort

AstaBench · overall

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAutomation-Bench1 version / measure
Open rankings and effort

Automation-Bench

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsAutomationBench14 versions / measures
Open rankings and effort

AutomationBench · September 2026 launch table

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AutomationBench · public 600-task set

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's public table reports strict task-completion pass rate as a percentage.

This is the public 600-task table, not the private held-out page leaderboard. Exact source label, highest available reasoning effort, repository revision, and the public task-set definition remain material; no cost is inferred.

Method / source

AutomationBench · official private leaderboard

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official strict task-completion rate is expressed as a percentage.

This is the private held-out leaderboard, not the public 600-task README table. Exact model label, reasoning effort, page version, row rank, cost markers, fallback behavior, and the private task-set revision remain material; no rank is treated as a performance metric.

Method / source

AutomationBench · Finance domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Finance domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · HR domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official HR domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · Marketing domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Marketing domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · Operations domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Operations domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · Sales domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Sales domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · Support domain summary

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Zapier's official Support domain summary score is a strict task-completion percentage.

This is a domain subset summary from the private leaderboard, not the overall leaderboard score. Exact top and runner-up labels, effort settings, page version, and private task-set revision remain material; no cost is inferred.

Method / source

AutomationBench · Anthropic release comparison

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AutomationBench · OpenAI launch comparison

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AutomationBench · private set

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AutomationBench Public

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

AutomationBench · business workflows

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionBabyVision1 version / measure
Open rankings and effort

BabyVision · with tools

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionBabyVision with python1 version / measure
Open rankings and effort

BabyVision with python

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningBBH1 version / measure
Open rankings and effort

BBH · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionBenchCAD1 version / measure
Open rankings and effort

BenchCAD

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionBenchCAD (python tool)1 version / measure
Open rankings and effort

BenchCAD (python tool)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsBFCL1 version / measure
Open rankings and effort

BFCL v4 · success after LWM RL

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsBFCL multi-turn1 version / measure
Open rankings and effort

BFCL multi-turn

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workBig Finance Bench1 version / measure
Open rankings and effort

Big Finance Bench

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingBigCodeBench1 version / measure
Open rankings and effort

BigCodeBench · Pass@1

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceBioDesign Tools (avg)1 version / measure
Open rankings and effort

BioDesign Tools (avg)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceBioMysteryBench4 versions / measures
Open rankings and effort

BioMysteryBench · human difficult

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BioMysteryBench · human solvable

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BioMysteryBench · hard

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BioMysteryBench · human solved

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionBlueprint-Bench1 version / measure
Open rankings and effort

Blueprint-Bench 2

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsBrowseComp4 versions / measures
Open rankings and effort

BrowseComp · agentic search

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BrowseComp

Can the model find hard-to-locate facts through multi-step web research?

Persistent browsing, source discovery, and synthesis for intentionally difficult questions. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BrowseComp · Pass@1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

BrowseComp · context management

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workC-Eval2 versions / measures
Open rankings and effort

C-Eval · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

C-Eval

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCapture-the-Flag Challenges1 version / measure
Open rankings and effort

Capture-the-Flag Challenges

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCatastrophic Cyber Misuse (ASR)1 version / measure
Open rankings and effort

Catastrophic Cyber Misuse (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionCC-OCR1 version / measure
Open rankings and effort

CC-OCR

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionChartography1 version / measure
Open rankings and effort

Chartography · with tools

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionCharXiv2 versions / measures
Open rankings and effort

CharXiv · RQ

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

CharXiv · RQ with python

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionCharXiv Reasoning (no tools)1 version / measure
Open rankings and effort

CharXiv Reasoning (no tools)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionCharXiv Reasoning (with tools)1 version / measure
Open rankings and effort

CharXiv Reasoning (with tools)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningChinese-SimpleQA1 version / measure
Open rankings and effort

Chinese-SimpleQA · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsClaw-Eval4 versions / measures
Open rankings and effort

Claw-Eval · Orchard-Claw

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Claw-Eval · Orchard-Claw with ZeroClaw

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Claw-Eval · success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Claw-Eval · success after LWM RL

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingClaw-Eval Avg1 version / measure
Open rankings and effort

Claw-Eval Avg

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingClaw-Eval Pass^31 version / measure
Open rankings and effort

Claw-Eval Pass^3

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningCLUEWSC1 version / measure
Open rankings and effort

CLUEWSC · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningCMath1 version / measure
Open rankings and effort

CMath · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workCMMLU1 version / measure
Open rankings and effort

CMMLU · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingCode Arena1 version / measure
Open rankings and effort

Code Arena · Web development · Elo

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingCodeforces2 versions / measures
Open rankings and effort

Codeforces · rating

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Codeforces · rating

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsCodex harness2 versions / measures
Open rankings and effort

Codex harness · Orchard-trained Orchard-Claw

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Codex harness · untrained Orchard-Claw

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsContinuous progress classification1 version / measure
Open rankings and effort

Continuous progress classification · accuracy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceCoronavirus–ACE2 Cell-Entry Screen2 versions / measures
Open rankings and effort

Coronavirus–ACE2 Cell-Entry Screen · composite score

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Composite score reported by OpenAI for held-out cell-entry measurements.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Coronavirus–ACE2 Cell-Entry Screen · helpful-only composite score

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Composite score for OpenAI's helpful-only Astra checkpoint.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
long contextCorpusQA 1M1 version / measure
Open rankings and effort

CorpusQA 1M · ACC

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionCountBench1 version / measure
Open rankings and effort

CountBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceCritPt1 version / measure
Open rankings and effort

CritPt

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsCross-embodiment skill transfer1 version / measure
Open rankings and effort

Cross-embodiment skill transfer · average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCurated CTFs (pass@1)1 version / measure
Open rankings and effort

Curated CTFs (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingCursorBench4 versions / measures
Open rankings and effort

CursorBench 4.0

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cursor's official CursorBench 4.0 score is a percentage.

Effort and harness are model-specific. The page warns that results have variance; small score differences may not be statistically meaningful. Costs are the publisher's pricing-based average, not observed billed amounts.

Method / source

CursorBench 4.0 · Medium Opus 5.5 comparison

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

CursorBench 3.2.0

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

CursorBench-3

Can the coding agent solve realistic, underspecified tasks drawn from Cursor's own engineering work?

Correctness on private, long-horizon software tasks run in Cursor's production-like harness. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
cybersecurityCybench (pass@1)1 version / measure
Open rankings and effort

Cybench (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyber Misuse Chat (ASR)1 version / measure
Open rankings and effort

Cyber Misuse Chat (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyber Misuse Chat (FRR)1 version / measure
Open rankings and effort

Cyber Misuse Chat (FRR)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports a false-refusal rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyber range exercises1 version / measure
Open rankings and effort

Cyber range exercises · combined pass rate

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyberGym2 versions / measures
Open rankings and effort

CyberGym

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

CyberGym · success rate

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyberGym (pass@1)1 version / measure
Open rankings and effort

CyberGym (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityCyScenarioBench (pass@1)1 version / measure
Open rankings and effort

CyScenarioBench (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningDeceptionBench1 version / measure
Open rankings and effort

DeceptionBench

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workDECK-Bench (Internal)1 version / measure
Open rankings and effort

DECK-Bench (Internal)

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingDeepPlanning1 version / measure
Open rankings and effort

DeepPlanning

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsDeepSearchQA1 version / measure
Open rankings and effort

DeepSearchQA · F1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsDeepShop1 version / measure
Open rankings and effort

DeepShop · Orchard-GUI

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingDeepSWE4 versions / measures
Open rankings and effort

DeepSWE v1.1 · pass@1

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of scored rollout attempts that passed.

Pass@1 is an attempt-level rate. Task coverage, repeated runs, confidence interval, mini-swe-agent version, model variant, and reasoning effort remain material.

Method / source

DeepSWE v1.1 · pass@4

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of attempted tasks with at least one passing rollout.

Pass@4 is a task-level rate, not the same measurement as pass@1. Task coverage, repeated runs, confidence interval, mini-swe-agent version, model variant, and reasoning effort remain material.

Method / source

DeepSWE v1.1 · agentic coding

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DeepSWE

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceDNA sequence design for transcription factor binding1 version / measure
Open rankings and effort

DNA sequence design for transcription factor binding · pass@1

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceDNA sequence design for transcription-factor binding1 version / measure
Open rankings and effort

DNA sequence design for transcription-factor binding

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Win rate over the Ledidi baseline reported by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsDOMINO2 versions / measures
Open rankings and effort

DOMINO · manipulation score

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DOMINO · success rate

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsDreamGen Bench1 version / measure
Open rankings and effort

DreamGen Bench · Total

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsDreamGen Bench component6 versions / measures
Open rankings and effort

DreamGen Bench component · GR1-Behavior IF

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DreamGen Bench component · GR1-Behavior PA

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DreamGen Bench component · GR1-Env IF

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DreamGen Bench component · GR1-Env PA

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DreamGen Bench component · GR1-Object IF

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

DreamGen Bench component · GR1-Object PA

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningDROP1 version / measure
Open rankings and effort

DROP · F1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingDSBench-FullStack †1 version / measure
Open rankings and effort

DSBench-FullStack †

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingDSBench-Hard †1 version / measure
Open rankings and effort

DSBench-Hard †

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
visionDynaMath1 version / measure
Open rankings and effort

DynaMath

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyDynamic adversarial user simulations3 versions / measures
Open rankings and effort

Dynamic adversarial user simulations · Emotional reliance

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Dynamic adversarial user simulations · Mental health

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Dynamic adversarial user simulations · Self-harm

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsEBench success rate4 versions / measures
Open rankings and effort

EBench success rate · LongHorizon

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench success rate · Overall

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench success rate · SimplePnP

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench success rate · TableTop

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsEBench task score4 versions / measures
Open rankings and effort

EBench task score · LongHorizon

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench task score · Overall

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench task score · SimplePnP

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EBench task score · TableTop

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceEEBench1 version / measure
Open rankings and effort

EEBench

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionEmbSpatialBench1 version / measure
Open rankings and effort

EmbSpatialBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionERQA1 version / measure
Open rankings and effort

ERQA

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsEWMBench1 version / measure
Open rankings and effort

EWMBench · Overall

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsEWMBench component8 versions / measures
Open rankings and effort

EWMBench component · BLEU

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · CLIP

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · Diversity

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · Dyn

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · HSD

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · Logics

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · nDTW

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

EWMBench component · SceneC

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityExploitBench4 versions / measures
Open rankings and effort

ExploitBench · Internal Port success rate

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

ExploitBench · Cap Percent

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. OpenAI's Cap Percent averages demonstrated capability flags across 41 V8 vulnerabilities.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

ExploitBench v8-bench · capability coverage

How far can the agent progress from reaching a vulnerable V8 code path to arbitrary code execution?

Sixteen deterministically graded capabilities on patched, previously vulnerable V8 builds. Higher is better. Percentage of the 16 capability flags achieved after OR-merging the selected seeds within every evaluated environment.

Control and all are different measurement bases. Agent, AutoNudge state, selected run, turn budget, seed count, failed cells, and environment coverage can all change the value.

Method / source

ExploitBench v8-bench · mean capability score

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Seed-weighted mean of the source's 0–16 per-run capability score across selected environments.

This is not capability coverage. Control and all select runs differently, and the all view may mix turn budgets and seed counts across configurations.

Method / source
cybersecurityExploitGym2 versions / measures
Open rankings and effort

ExploitGym · reported challenge success

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

ExploitGym

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityExploitGym (pass@1)1 version / measure
Open rankings and effort

ExploitGym (pass@1)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workFACTS Parametric1 version / measure
Open rankings and effort

FACTS Parametric · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsFew-shot skill transfer5 versions / measures
Open rankings and effort

Few-shot skill transfer · average success · Fold Towel

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Few-shot skill transfer · average success · Insert Screw

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Few-shot skill transfer · average success · Put Blocks

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Few-shot skill transfer · average success · Put Fruits

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Few-shot skill transfer · average success · Unscrew Cap

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workFinance Agent1 version / measure
Open rankings and effort

Finance Agent v2

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityFORTRESS (ARS)1 version / measure
Open rankings and effort

FORTRESS (ARS)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack-risk score as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyFrontier Red Team16 versions / measures
Open rankings and effort

Frontier Red Team · identity classification · easy

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the easy tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity classification · hard

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the hard tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity classification · medium

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the medium tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity correlation · easy

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the easy tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity correlation · hard

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the hard tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity correlation · medium

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Linkage F1 for identity correlation on the medium tier.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · identity correlation · pooled

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Pooled linkage F1 for identity correlation across the 50 reported samples.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · text geolocation · median home-location error

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Lower is better. Median home-location error in kilometers for text geolocation.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · photo geolocation · median distance error

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Lower is better. Median distance error in kilometers for photo geolocation.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · photo geolocation · within 1 km

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of photo-geolocation predictions within 1 kilometer.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · terminal guidance · parked high-visibility target

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Simulated strike rate for terminal guidance to a parked high-visibility target.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · terminal guidance · road-speed target

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Simulated strike rate for terminal guidance to a vehicle moving at road speed.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · payload delivery · target moving at 3 m/s

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a target moving at approximately 3 m/s.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · payload delivery · camouflaged zig-zag target

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a camouflaged zig-zag target.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · payload delivery · hard wind and moving target

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters under the hard wind-and-moving-target setting.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source

Frontier Red Team · payload delivery · stationary target within 5 m

How does a model perform on this Anthropic-defined Frontier Red Team evaluation?

Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Percentage of simulated payload sorties with median miss within five meters of a stationary target.

This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.

Method / source
codingFrontier-Bench1 version / measure
Open rankings and effort

Frontier-Bench v0.1 · agentic terminal coding

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingFrontierCode9 versions / measures
Open rankings and effort

FrontierCode 1.1 · Main

Would a repository maintainer merge this software change?

End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Cognition's Score is a weighted aggregate of rubric items; solutions that fail a blocking criterion receive 0.

Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.

Method / source

FrontierCode 1.1 · Extended

Would a repository maintainer merge this software change?

End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Cognition's Score is a weighted aggregate of rubric items; solutions that fail a blocking criterion receive 0.

Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.

Method / source

FrontierCode 1.1 · Extended · Pass Rate

Would a repository maintainer merge this software change?

End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Pass Rate is the fraction of trials satisfying every blocker rubric criterion, averaged over the selected tasks.

Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.

Method / source

FrontierCode 1.1 · Main · Pass Rate

Would a repository maintainer merge this software change?

End-to-end code quality, including correctness, test quality, scope discipline, style, and adherence to repository standards. Higher is better. Pass Rate is the fraction of trials satisfying every blocker rubric criterion, averaged over the selected tasks.

Weighted Score is not all-blocker Pass Rate. Harness and effort vary by row; compare only exact contexts and splits. The feed does not publish attempt counts.

Method / source

FrontierCode 1.1 · Main · Medium Opus 5.5 comparison

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

FrontierCode v1.1 · Main

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

FrontierCode 1.0 · Diamond

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

FrontierCode 1.0 · Extended

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cognition reports FrontierCode score as a percentage, not pass rate.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

FrontierCode 1.0 · Main

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Cognition reports FrontierCode score as a percentage, not pass rate.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
reasoningFrontierMath Tier 1-3 (v2)1 version / measure
Open rankings and effort

FrontierMath Tier 1-3 (v2)

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningFrontierMath Tier 4 (v2)1 version / measure
Open rankings and effort

FrontierMath Tier 4 (v2)

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingFrontierSWE1 version / measure
Open rankings and effort

FrontierSWE

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsGDM Situational Awareness1 version / measure
Open rankings and effort

GDM Situational Awareness

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsGDM-Stealth (of 4)1 version / measure
Open rankings and effort

GDM-Stealth (of 4)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visiongdp.pdf1 version / measure
Open rankings and effort

gdp.pdf

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceGeneBench Pro1 version / measure
Open rankings and effort

GeneBench Pro

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceGPQA4 versions / measures
Open rankings and effort

GPQA Diamond

Can the model reason through specialist biology, physics, and chemistry questions?

Multiple-choice science questions written and validated by domain experts. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

GPQA Diamond · Pass@1

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

GPQA Diamond

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

GPQA

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextGraphWalks BFS2 versions / measures
Open rankings and effort

GraphWalks BFS · 1M F1

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

GraphWalks BFS · 256K F1

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsGraySwan ART (pass@1 ASR)1 version / measure
Open rankings and effort

GraySwan ART (pass@1 ASR)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsGripper dexterity3 versions / measures
Open rankings and effort

Gripper dexterity · diverse tool kitting

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Gripper dexterity · general pick and place

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Gripper dexterity · precise insertion tasks

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningGSM8K1 version / measure
Open rankings and effort

GSM8K · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionHallusionBench1 version / measure
Open rankings and effort

HallusionBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceHard-negative protein binding prediction1 version / measure
Open rankings and effort

Hard-negative protein binding prediction · pass@4

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceHard-negative protein binding prediction2 versions / measures
Open rankings and effort

Hard-negative protein binding prediction

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Percentage of protein-binding selections scored correct by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Hard-negative protein binding prediction · helpful-only

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Percentage of protein-binding selections scored correct by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
knowledge workHarvey's Legal Agent Benchmark1 version / measure
Open rankings and effort

Harvey's Legal Agent Benchmark

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workHealthBench5 versions / measures
Open rankings and effort

HealthBench · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench Consensus · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench Hard · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench Professional · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workHealthBench Consensus1 version / measure
Open rankings and effort

HealthBench Consensus · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workHealthBench Hard1 version / measure
Open rankings and effort

HealthBench Hard · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workHealthBench Pro1 version / measure
Open rankings and effort

HealthBench Pro

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workHealthBench Professional3 versions / measures
Open rankings and effort

HealthBench Professional

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench Professional · length-adjusted score

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HealthBench Professional · health

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHellaSwag1 version / measure
Open rankings and effort

HellaSwag · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHMMT3 versions / measures
Open rankings and effort

HMMT 2025 · pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HMMT · February 2026

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HMMT · November 2025

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHMMT 2026 Feb1 version / measure
Open rankings and effort

HMMT 2026 Feb · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHMMT Feb2 versions / measures
Open rankings and effort

HMMT Feb 26

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HMMT Feb 25

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHMMT Nov1 version / measure
Open rankings and effort

HMMT Nov 25

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceHPCT1 version / measure
Open rankings and effort

HPCT

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingHumanEval1 version / measure
Open rankings and effort

HumanEval · Pass@1

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningHumanity’s Last Exam14 versions / measures
Open rankings and effort

Humanity's Last Exam · full

Can the model answer deliberately difficult expert questions across academic fields?

Broad expert-level reasoning and knowledge without external tools. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · text only

Can the model answer the text-only portion of deliberately difficult expert questions?

The publisher's finalized HLE text-only view, preserving the official score, confidence interval, calibration error, and contamination warning fields. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · September 2026 launch table · no tools

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · September 2026 launch table · with tools

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · full with tools

How much do search or code tools help on Humanity's Last Exam?

The full expert-question set with the reporting lab's stated tool setup. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE-Verified · full 1,811-item set

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE · no tools

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE · with tools

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE Calibration

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE with tools · Pass@1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

HLE

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · full set

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Humanity's Last Exam · search + code

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningIFBench1 version / measure
Open rankings and effort

IFBench

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyImage input evaluations4 versions / measures
Open rankings and effort

Image input evaluations · extremism

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Image input evaluations · harms-erotic

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Image input evaluations · hate

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Image input evaluations · self-harm

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningIMOAnswerBench2 versions / measures
Open rankings and effort

IMOAnswerBench

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

IMOAnswerBench · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsIn-distribution manipulation3 versions / measures
Open rankings and effort

In-distribution manipulation · LIBERO

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

In-distribution manipulation · RT-Easy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

In-distribution manipulation · RT-Hard

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityInternal Capture-the-Flag challenges1 version / measure
Open rankings and effort

Internal Capture-the-Flag challenges · defender success

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
knowledge workInternal Company Research Reports1 version / measure
Open rankings and effort

Internal Company Research Reports · accepted reports

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. This benchmark reports points rather than percent correct.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingInternal Research Debugging Evaluation1 version / measure
Open rankings and effort

Internal Research Debugging Evaluation

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingInternal Research Debugging Evaluation1 version / measure
Open rankings and effort

Internal Research Debugging Evaluation

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of internal research-debugging tasks solved under OpenAI's evaluation.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
reasoningInternal Sycophancy1 version / measure
Open rankings and effort

Internal Sycophancy

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
cybersecurityJailbreak StrongREJECT v2 (ASR)1 version / measure
Open rankings and effort

Jailbreak StrongREJECT v2 (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workJob Bench1 version / measure
Open rankings and effort

Job Bench

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingKernelBench Hard1 version / measure
Open rankings and effort

KernelBench Hard

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingKernelGen 1P1 version / measure
Open rankings and effort

KernelGen 1P

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsKimi Claw 24/7 Bench (Internal)1 version / measure
Open rankings and effort

Kimi Claw 24/7 Bench (Internal)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingKimi Code Bench 2.0 (Internal)1 version / measure
Open rankings and effort

Kimi Code Bench 2.0 (Internal)

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingKimi Code Bench v2 (Internal)1 version / measure
Open rankings and effort

Kimi Code Bench v2 (Internal)

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceLABBench21 version / measure
Open rankings and effort

LABBench2

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workLegal Agent Benchmark2 versions / measures
Open rankings and effort

Legal Agent Benchmark · Held-out

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Legal Agent Benchmark

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsLIBERO-Plus OOD robustness8 versions / measures
Open rankings and effort

LIBERO-Plus OOD robustness · Background

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Camera

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Language

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Layout

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Light

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Noise

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Robot

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LIBERO-Plus OOD robustness · Total

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceLifeSciBench1 version / measure
Open rankings and effort

LifeSciBench

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingLiveCodeBench3 versions / measures
Open rankings and effort

LiveCodeBench v6

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Arithmetic mean of the official per-problem pass@1 values in the selected date window.

The selected 454-problem window is the official page's default date range, not the full 1,055-problem release_v6 set. Exact source label, model metadata, release date, contamination flag, difficulty counts, and date window remain material; no cost is inferred.

Method / source

LiveCodeBench v6

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LiveCodeBench · Pass@1

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingLiveCodeBench Pro1 version / measure
Open rankings and effort

LiveCodeBench Pro · Elo

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextLongBench1 version / measure
Open rankings and effort

LongBench-V2 · EM

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionLVBench3 versions / measures
Open rankings and effort

LVBench · agentic

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LVBench · static

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

LVBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workManagement Consulting Tasks (Internal)1 version / measure
Open rankings and effort

Management Consulting Tasks (Internal)

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsManipulation benchmark5 versions / measures
Open rankings and effort

Manipulation benchmark · LIBERO

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Manipulation benchmark · RoboCasa-GR1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Manipulation benchmark · RoboTwin-Easy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Manipulation benchmark · RoboTwin-Hard

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Manipulation benchmark · Simpler-WidowX

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsMASK1 version / measure
Open rankings and effort

MASK

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningMATH1 version / measure
Open rankings and effort

MATH · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningMathArena Apex1 version / measure
Open rankings and effort

MathArena Apex

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMathVision1 version / measure
Open rankings and effort

MathVision

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMathVision with python1 version / measure
Open rankings and effort

MathVision with python

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMathvista(mini)1 version / measure
Open rankings and effort

Mathvista(mini)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceMBCT1 version / measure
Open rankings and effort

MBCT

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsMCP Atlas3 versions / measures
Open rankings and effort

MCP Atlas

Can the model discover and coordinate the right tools to complete a multi-step workflow?

One thousand human-authored tasks spanning 36 real MCP servers and 220 tools. Higher is better. Scale's published MCP Atlas score field is retained in percentage points.

This row set uses the April 2026 methodology revision and its 1,000-task composition. Exact source label, version, effort label, rank, confidence interval, contamination warning, publication flags, and entry timestamp remain material; no model-specific tool setup or cost is inferred.

Method / source

MCP-Atlas · public set

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MCP-Atlas

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsMCP Mark Verified1 version / measure
Open rankings and effort

MCP Mark Verified

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsMCPAtlas Public1 version / measure
Open rankings and effort

MCPAtlas Public · Pass@1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsMCPMark2 versions / measures
Open rankings and effort

MCPMark · success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MCPMark

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceMedChemBench (Internal)1 version / measure
Open rankings and effort

MedChemBench (Internal)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
reasoningMGSM1 version / measure
Open rankings and effort

MGSM · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingMLE-Bench1 version / measure
Open rankings and effort

MLE-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingMLS Bench1 version / measure
Open rankings and effort

MLS Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingMLS Bench Lite1 version / measure
Open rankings and effort

MLS Bench Lite

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMLVU1 version / measure
Open rankings and effort

MLVU

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMMBench EN-DEV1 version / measure
Open rankings and effort

MMBench EN-DEV-v1.1

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workMMLU5 versions / measures
Open rankings and effort

MMLU-Pro · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMLU · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMLU-Redux · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMLU-Pro

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMLU-Redux

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workMMMLU2 versions / measures
Open rankings and effort

MMMLU · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMMLU

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMMMU5 versions / measures
Open rankings and effort

MMMU-Pro · selected exact setups

Can the model combine visual evidence with college-level knowledge and deliberate reasoning?

Multidiscipline multimodal questions filtered for genuine visual dependence under a no-tools protocol. Higher is better. MMMU Team's pro.overall percentage under the official zero-shot evaluation.

Tool access, exact source setup label, model publication date, model size, author-provided marker, and the maintainer's selected column are material. Do not mix this row with validation, test, or provider-reported values.

Method / source

MMMU-Pro · with tools

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. MMMU Team's pro.overall percentage for the exact with-tools setup label.

This is a separate tool-enabled setup from no-tools rows. Exact source label, model publication date, model size, author-provided marker, and the maintainer's selected column remain material.

Method / source

MMMU-Pro with python

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMMU

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MMMU-Pro

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMMStar1 version / measure
Open rankings and effort

MMStar

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextMRCR2 versions / measures
Open rankings and effort

MRCR v2 · 8-needle · 128K average

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

MRCR v2 · 8-needle · 1M pointwise

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextMRCR 1M1 version / measure
Open rankings and effort

MRCR 1M · MMR

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextMRCR Long Context (1M context window)1 version / measure
Open rankings and effort

MRCR Long Context (1M context window)

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityMSRC historical cases2 versions / measures
Open rankings and effort

MSRC historical cases · clfs.sys recall

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

MSRC historical cases · tcpip.sys recall

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsMulti-finger dexterity5 versions / measures
Open rankings and effort

Multi-finger dexterity · dustpan

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Multi-finger dexterity · screw bulb

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Multi-finger dexterity · tie trash bag

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Multi-finger dexterity · unscrew bulb

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Multi-finger dexterity · Ziplock

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingMulti-SWE-Bench1 version / measure
Open rankings and effort

Multi-SWE-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workMultiLoKo1 version / measure
Open rankings and effort

MultiLoKo · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionMVBench1 version / measure
Open rankings and effort

MVBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingNanoGPT1 version / measure
Open rankings and effort

NanoGPT

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingNL2Repo1 version / measure
Open rankings and effort

NL2Repo

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingNL2Repo-Bench1 version / measure
Open rankings and effort

NL2Repo-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionOCRBench1 version / measure
Open rankings and effort

OCRBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionODInW131 version / measure
Open rankings and effort

ODInW13

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workOffice QA Pro1 version / measure
Open rankings and effort

Office QA Pro

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionOmniDocBench1 version / measure
Open rankings and effort

OmniDocBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionOmniDocBench1.51 version / measure
Open rankings and effort

OmniDocBench1.5

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsOnline-Mind2Web1 version / measure
Open rankings and effort

Online-Mind2Web · Orchard-GUI

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
long contextOpenAI MRCR2 versions / measures
Open rankings and effort

OpenAI MRCR v2 · 8-needle · 256K-512K

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OpenAI MRCR v2 · 8-needle · 512K-1M

Can the model retain and reason over evidence spread across a very long input?

Long-context retrieval and reasoning, not context-window size alone. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsOR-Bench (FRR)1 version / measure
Open rankings and effort

OR-Bench (FRR)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsOrchard-GUI1 version / measure
Open rankings and effort

Orchard-GUI · three-benchmark average

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsOSWorld10 versions / measures
Open rankings and effort

OSWorld 2.0 · partial score

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · offline partial reward

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · August 2026 task release · partial

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · August 2026 task release · strict

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · partial score · batch tool enabled

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · partial score

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · computer use

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld 2.0 · strict binary completion

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

OSWorld-Verified

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. Official OSWorld-Verified success-rate percentage averaged across repeated source rows for the exact Model and Max steps group.

Exact source model label, approach type, institution, maximum steps, source date, 361-or-369 task denominator, per-application category strings, repeated-run count and population standard deviation, augmentation flags, and OSWorld-Verified task revision remain material. Do not merge this workbook with OSWorld 2.0 or provider-reported OSWorld rows.

Method / source
agentsPBench11 versions / measures
Open rankings and effort

PBench · Aes

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Bg-Con

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Domain

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · I2V-Bg

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · I2V-S

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Img

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Mot

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · O-Con

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Overall

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Quality

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

PBench · Sub-Con

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionPerceptionBench1 version / measure
Open rankings and effort

PerceptionBench

Can the model accurately perceive visual evidence before reasoning or outside knowledge is required?

Three thousand verified questions spanning ten atomic visual-perception capabilities. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
sciencePhage–plasmid co-evolution2 versions / measures
Open rankings and effort

Phage–plasmid co-evolution · helpful-only negative log-likelihood

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Lower is better. Negative log-likelihood for OpenAI's helpful-only Astra checkpoint.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Phage–plasmid co-evolution · negative log-likelihood

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Lower is better. Negative log-likelihood of subsequent mutations.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
cybersecurityPoly-Guard Bench (ASR)1 version / measure
Open rankings and effort

Poly-Guard Bench (ASR)

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingPostTrain Bench1 version / measure
Open rankings and effort

PostTrain Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingPostTrainBench1 version / measure
Open rankings and effort

PostTrainBench · normalized score

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingPostTrainBench Lite1 version / measure
Open rankings and effort

PostTrainBench Lite

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsPrecision moment-finding2 versions / measures
Open rankings and effort

Precision moment-finding · accuracy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Precision moment-finding · mean absolute distance

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. Mean absolute distance in seconds between the predicted and critical video frame.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyProduction Benchmarks8 versions / measures
Open rankings and effort

Production Benchmarks · Extremism

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Gore

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Hate

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Nonviolent illicit behavior

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Self-harm (standard)

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Sexual

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Sexual/minors

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Production Benchmarks · Violent Illicit behavior

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingProgramBench1 version / measure
Open rankings and effort

ProgramBench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyPrompt injection attacks in connectors1 version / measure
Open rankings and effort

Prompt injection attacks in connectors

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceProtocolQA1 version / measure
Open rankings and effort

ProtocolQA

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceProtocolQA Open-Ended2 versions / measures
Open rankings and effort

ProtocolQA Open-Ended

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score reported by OpenAI on the open-ended troubleshooting version.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

ProtocolQA Open-Ended

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsQwenClawBench3 versions / measures
Open rankings and effort

QwenClawBench · success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

QwenClawBench · success after LWM RL

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

QwenClawBench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingQwenWebBench1 version / measure
Open rankings and effort

QwenWebBench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsR2R-CE2 versions / measures
Open rankings and effort

R2R-CE · validation seen success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

R2R-CE · validation unseen success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsReal-world ALOHA2 versions / measures
Open rankings and effort

Real-world ALOHA · in-domain average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Real-world ALOHA · out-of-domain average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsReal-world ALOHA dual-arm2 versions / measures
Open rankings and effort

Real-world ALOHA dual-arm · In-domain average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Real-world ALOHA dual-arm · OOD average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionRealWorldQA1 version / measure
Open rankings and effort

RealWorldQA

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionRefCOCO(avg)1 version / measure
Open rankings and effort

RefCOCO(avg)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionRefSpatialBench1 version / measure
Open rankings and effort

RefSpatialBench

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceRefusals: BioTIER1 version / measure
Open rankings and effort

Refusals: BioTIER

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. The source reports a refusal percentage; interpret with the named safety objective.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceRefusals: Chemical Agents1 version / measure
Open rankings and effort

Refusals: Chemical Agents

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. The source reports a refusal percentage; interpret with the named safety objective.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceReproBAIT1 version / measure
Open rankings and effort

ReproBAIT · published-performance fraction

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Fraction of published performance reached on SecureBio's six-task agentic biology evaluation.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsRISE1 version / measure
Open rankings and effort

RISE

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsRoboCasa3654 versions / measures
Open rankings and effort

RoboCasa365 · Atomic

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboCasa365 · Composite-Seen

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboCasa365 · Composite-Unseen

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboCasa365 · Total

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsRoboChallenge2 versions / measures
Open rankings and effort

RoboChallenge · bimanual coordination average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboChallenge · robust pick-and-place average success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsRoboChallenge Table302 versions / measures
Open rankings and effort

RoboChallenge Table30 v1 · generalist-track success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboChallenge Table30 v1 · process score

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsRoboTwin-Clean2Rand6 versions / measures
Open rankings and effort

RoboTwin-Clean2Rand · Background

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-Clean2Rand · Clutter

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-Clean2Rand · Easy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-Clean2Rand · Hard

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-Clean2Rand · Height

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-Clean2Rand · Light

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsRoboTwin-IF6 versions / measures
Open rankings and effort

RoboTwin-IF · Average

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-IF · Open Microwave Door

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-IF · Open Stapler

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-IF · Operate Table

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-IF · Pick-Diverse

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

RoboTwin-IF · Place-Relative

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingRSI Index1 version / measure
Open rankings and effort

RSI Index

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
cybersecuritySandbox Bench1 version / measure
Open rankings and effort

Sandbox Bench · targets exploited

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of private targets on which the model recovered the protected flag.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
agentsSAVE-Bench1 version / measure
Open rankings and effort

SAVE-Bench

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceSciCode2 versions / measures
Open rankings and effort

SciCode · main-problem resolve rate

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Official main-problem resolve rate over the 80 main problems.

This is the repository table's main-problem measurement. The result changelog date is not stated in the pinned README; keep the repository revision, exact source label, and separate subproblem measurement attached.

Method / source

SciCode · subproblem accuracy

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Official subproblem accuracy over the 338 decomposed subproblems.

This is a separate subproblem measurement from the main-problem resolve rate. The result changelog date is not stated in the pinned README; keep the repository revision and exact source label attached.

Method / source
cybersecuritySEC-Bench Pro2 versions / measures
Open rankings and effort

SEC-Bench Pro · reported score

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SEC-Bench Pro

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceSecureBio Multimodal Troubleshooting Virology1 version / measure
Open rankings and effort

SecureBio Multimodal Troubleshooting Virology

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Refusal-corrected accuracy reported by OpenAI from SecureBio's evaluation.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceSeqQA (agentic)1 version / measure
Open rankings and effort

SeqQA (agentic)

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsSHADE-Arena1 version / measure
Open rankings and effort

SHADE-Arena

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceSHP2 Protein Function Prediction2 versions / measures
Open rankings and effort

SHP2 Protein Function Prediction · maximum mean R²

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Maximum mean R² across three unpublished assay datasets.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

SHP2 Protein Function Prediction · helpful-only maximum mean R²

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Maximum mean R² for OpenAI's helpful-only Astra checkpoint.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
reasoningSimple-QA verified1 version / measure
Open rankings and effort

Simple-QA verified · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningSimpleQA-Verified1 version / measure
Open rankings and effort

SimpleQA-Verified · Pass@1

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionSimpleVQA1 version / measure
Open rankings and effort

SimpleVQA

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSkillsBench Avg51 version / measure
Open rankings and effort

SkillsBench Avg5

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecuritySocial Engineering1 version / measure
Open rankings and effort

Social Engineering

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workSpreadsheetBench1 version / measure
Open rankings and effort

SpreadsheetBench 2

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSQLite reconstruction1 version / measure
Open rankings and effort

SQLite reconstruction · held-out sqllogictest

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecuritySRE-Bench1 version / measure
Open rankings and effort

SRE-Bench · pass@4

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of reverse-engineering challenges fully solved across four trials.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
safetyStatic Jailbreak Evaluation5 versions / measures
Open rankings and effort

Static Jailbreak Evaluation · biological high risk defender success

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Static Jailbreak Evaluation · biological severe defender success

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Static Jailbreak Evaluation · cyber defender success

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Static Jailbreak Evaluation · violence moderate defender success

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Static Jailbreak Evaluation · violence severe defender success

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
cybersecurityStorageDrive1 version / measure
Open rankings and effort

StorageDrive · planted vulnerabilities found

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Number of planted vulnerabilities found out of 21.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceSuperGPQA2 versions / measures
Open rankings and effort

SuperGPQA · EM

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SuperGPQA

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE Marathon1 version / measure
Open rankings and effort

SWE Marathon

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE Multilingual1 version / measure
Open rankings and effort

SWE Multilingual · resolved

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE Pro1 version / measure
Open rankings and effort

SWE Pro · resolved

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE Verified1 version / measure
Open rankings and effort

SWE Verified · resolved

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE-Bench12 versions / measures
Open rankings and effort

SWE-Bench Pro V2 · Public Full

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Resolve Rate on the official 642-task Public Full split across 11 repositories.

Scale's V2 release changes task composition and protocol relative to the 731-task V1 Public set. Compare only rows in this explicit V2 Full view; exact harness, effort, source label, rank, and displayed uncertainty are retained per model.

Method / source

SWE-Bench Pro · Public

Can an agent resolve difficult software issues from real repositories?

Repository-level issue resolution on the SWE-Bench Pro task set. Higher is better. Resolve Rate score on the official 731-instance Public Set.

This is a benchmark-maintainer leaderboard result with no agent configuration exposed in the embedded entries. Exact source label, effort marker, confidence interval, contamination/deprecation flags, entry timestamp, and task-set revision are material; private and other SWE-Bench Pro splits are not comparable.

Method / source

SWE-bench Verified

Can an agent fix a real GitHub issue and pass the repository's tests?

The share of engineer-verified software issues resolved by the full model-plus-agent system. Higher is better. Resolved percent on the official verified-500 task set under the bash-only leaderboard block.

This is a model-plus-agent result. Exact agent, model variant, effort, submission date, mini-SWE-agent version, system flags, and task-set revision are material; never compare it with another SWE-bench split or harness as if they were the same test.

Method / source

SWE-bench Multilingual

Can an agent resolve software issues across multiple programming languages?

The share of 300 multilingual software issues resolved under the benchmark maintainer's published mini-SWE-agent protocol. Higher is better. Resolved percent on the official multilingual-300-9-languages task set under the default leaderboard block.

This is a model-plus-agent result. Exact agent, model variant, effort, submission date, mini-SWE-agent version, system flags, and task-set revision are material; never compare it with another SWE-bench split or harness as if they were the same test.

Method / source

SWE-Bench Verified · Orchard-SWE Balanced Adaptive Rollout

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-Bench Verified · Orchard-SWE baseline

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-Bench Verified · Orchard-SWE value-model reranking

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-Bench Pro · success

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-Bench Verified · success

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-bench Multilingual

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-bench Pro

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

SWE-bench Verified

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingSWE-fficiency1 version / measure
Open rankings and effort

SWE-fficiency

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceTacit Knowledge and Troubleshooting1 version / measure
Open rankings and effort

Tacit Knowledge and Troubleshooting · refusal-adjusted

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
scienceTacit Knowledge and Troubleshooting2 versions / measures
Open rankings and effort

Tacit Knowledge and Troubleshooting

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Original score reported by OpenAI on an internal multiple-choice evaluation.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

Tacit Knowledge and Troubleshooting · refusal-adjusted

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score after counting refusals and safe completions as successes, as reported by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
codingTAU3-Bench1 version / measure
Open rankings and effort

TAU3-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsTerminal-Bench12 versions / measures
Open rankings and effort

Terminal-Bench 4.0

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 3.0

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.1

Can a model-driven agent complete realistic work inside a terminal?

Multi-step command-line, coding, configuration, debugging, and systems tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal Bench 2.1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.1 · Terminus-2

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.1 · best reported harness

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.0 · success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.0

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal Bench 2.0 · accuracy

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench 2.0 · Codex harness

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench Hard

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench · Cursor report

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceTerminal-Bench Science3 versions / measures
Open rankings and effort

Terminal-Bench-Science 0.1 · official leaderboard

How often an AI agent completes a scientific research workflow using terminal tools.

Resolution rate across 70 research workflows in life, physical, earth, mathematical and engineering sciences, with three trials per task. Higher is better. The maintainer reports the percentage of scored trials resolved successfully.

This is a small, early research-workflow suite, not a general measure of scientific discovery. Model and agent setups differ. Retain source standard errors and exact dataset revision; never mix these official results with provider reports or other benchmark versions.

Method / source

Terminal-Bench-Science 0.1 · Anthropic Claude Code comparison

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Terminal-Bench-Science 0.1

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsTool-Decathlon2 versions / measures
Open rankings and effort

Tool Decathlon · success

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Tool-Decathlon

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsToolathlon2 versions / measures
Open rankings and effort

Toolathlon

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Toolathlon · Pass@1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsToolathlon-Verified1 version / measure
Open rankings and effort

Toolathlon-Verified

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workTriviaQA1 version / measure
Open rankings and effort

TriviaQA · EM

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceTroubleshootingBench2 versions / measures
Open rankings and effort

TroubleshootingBench

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Original score reported by OpenAI on expert-written troubleshooting questions.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source

TroubleshootingBench · refusal-adjusted

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Score after counting refusals and safe completions as successes, as reported by OpenAI.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
knowledge workTutorMoments3 versions / measures
Open rankings and effort

TutorMoments · appropriate rigor

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of annotated moments where the model made the appropriate tutoring move.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

TutorMoments · appropriate scaffolding

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of annotated moments where the model made the appropriate tutoring move.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

TutorMoments · avoids over-scaffolding

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Share of moments where the model avoided unnecessary scaffolding.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
safetyU186 versions / measures
Open rankings and effort

U18 · Age-restricted goods, services, and dangerous challenges/activities

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

U18 · Eating Disorders

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

U18 · Emotional Reliance

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

U18 · Gore

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

U18 · Self Harm

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

U18 · Sexual Content

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionV*1 version / measure
Open rankings and effort

V*

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
cybersecurityV8 JavaScript Engine1 version / measure
Open rankings and effort

V8 JavaScript Engine · unique confirmed issues

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Count of unique confirmed vulnerabilities found across the source's fixed number of invocations.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
knowledge workVals Finance Agent1 version / measure
Open rankings and effort

Vals Finance Agent v2

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceVCT1 version / measure
Open rankings and effort

VCT

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingVIBE-Pro1 version / measure
Open rankings and effort

VIBE-Pro · average

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source
visionVideoMME(w sub.)1 version / measure
Open rankings and effort

VideoMME(w sub.)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionVideoMME(w/o sub.)1 version / measure
Open rankings and effort

VideoMME(w/o sub.)

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionVideoMMMU1 version / measure
Open rankings and effort

VideoMMMU

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsVision-language navigation4 versions / measures
Open rankings and effort

Vision-language navigation · R2R Val-Unseen · Oracle Success Rate

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Vision-language navigation · R2R Val-Unseen · Success Rate

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Vision-language navigation · RxR Val-Unseen · SPL

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Vision-language navigation · RxR Val-Unseen · Success Rate

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
codingVITA-Bench1 version / measure
Open rankings and effort

VITA-Bench

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionVlmsAreBlind1 version / measure
Open rankings and effort

VlmsAreBlind

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWebArena-Verified1 version / measure
Open rankings and effort

WebArena-Verified

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWebVoyager1 version / measure
Open rankings and effort

WebVoyager · Orchard-GUI

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWhole-body manipulation3 versions / measures
Open rankings and effort

Whole-body manipulation · pick up from floor

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Whole-body manipulation · pick up from shelf

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Whole-body manipulation · pick up from table

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWide Search1 version / measure
Open rankings and effort

Wide Search

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWideSearch4 versions / measures
Open rankings and effort

WideSearch · F1 by Item

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WideSearch · F1 by Row

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WideSearch · F1 by Item after LWM RL

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WideSearch

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
reasoningWinoGrande1 version / measure
Open rankings and effort

WinoGrande · EM

Can the model solve difficult problems that require more than factual recall?

Multi-step reasoning on academic, mathematical, or abstract problems. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceWMDP-Bio1 version / measure
Open rankings and effort

WMDP-Bio

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
scienceWMDP-Chem1 version / measure
Open rankings and effort

WMDP-Chem

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWorldModelBench3 versions / measures
Open rankings and effort

WorldModelBench · Instruction (0-3)

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench · Phys.

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench · Total

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsWorldModelBench physics component7 versions / measures
Open rankings and effort

WorldModelBench physics component · Fluid

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Frame

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Grav.

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Mass

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Newton

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Penetr.

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

WorldModelBench physics component · Temp

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionWorldVQA ForceAnswer1 version / measure
Open rankings and effort

WorldVQA ForceAnswer

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsZero-shot cross-embodiment4 versions / measures
Open rankings and effort

Zero-shot cross-embodiment · ARX

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Zero-shot cross-embodiment · Franka

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Zero-shot cross-embodiment · Total

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

Zero-shot cross-embodiment · UR5

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionZeroBench main with python1 version / measure
Open rankings and effort

ZeroBench main with python · pass@5

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionZEROBench_sub1 version / measure
Open rankings and effort

ZEROBench_sub

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
visionZeroBench-main2 versions / measures
Open rankings and effort

ZeroBench-main · with tools

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

ZeroBench main · pass@5

Can the model extract and reason over information in images or documents?

Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source
agentsτ²-bench3 versions / measures
Open rankings and effort

τ²-bench · Telecom

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

τ²-bench · Retail

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source

τ²-bench · pass@1

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source