Benchmark familiesBrowse 296

Benchmark overview

Five leading models in each audited category, using one exact representative benchmark per category.

Families
296
Categories
9
Updated
Sep 23, 2026

Coding

Repository code-change mergeability

FrontierCode 1.1 · Main
  1. 1Claude Opus 5.5medium effort · Cognition leaderboard54.64%
  2. 2Claude Fable 5xhigh effort · Cognition leaderboard53.48%
  3. 3Claude Opus 5medium effort · Cognition leaderboard53.38%
  4. 4GPT-6 Astramax effort · Cognition leaderboard53.26%
  5. 5Claude Fable 5.1medium effort · Cognition leaderboard50.91%
0Higher is better100%

Agents

Terminal agent work

Terminal-Bench 4.0
  1. 1Claude Opus 5.5xhigh effort · Anthropic report66.40%
  2. 2Mythos 5.1Reported configuration · Anthropic report60.90%
  3. 3GPT-6 Astramax effort · Terminal-Bench 4.0 · official leaderboard58.18%
  4. 4Claude Fable 5.1max effort · Terminal-Bench 4.0 · official leaderboard57.88%
  5. 5Claude Opus 5xhigh effort · Terminal-Bench 4.0 · official leaderboard53.94%
0Higher is better100%

Expert reasoning

Broad expert questions

Humanity's Last Exam · full
  1. 1Claude Mythos PreviewReported configuration · Anthropic report56.80%
  2. 2GPT-6 AstraReported configuration · HLE · full54.80%
  3. 3Claude Fable 5max effort · Moonshot AI report53.30%
  4. 4Muse Spark 1.1xhigh effort · Meta report52.20%
  5. 5Claude Opus 4.8max effort · Moonshot AI report49.80%
0Higher is better100%

Science

Scientific research workflows

TB-Science 0.1 · official
  1. 1GPT-6 Astramax effort · Terminal-Bench-Science 0.1 · official leaderboard68.1%
  2. 2Claude Fable 5.1max effort · Terminal-Bench-Science 0.1 · official leaderboard40.0%
  3. 3Claude Opus 5max effort · Terminal-Bench-Science 0.1 · official leaderboard30.0%
  4. 4GPT-5.6 Solmax effort · Terminal-Bench-Science 0.1 · official leaderboard22.4%
  5. 5Claude Fable 5max effort · Terminal-Bench-Science 0.1 · official leaderboard21.4%
0Higher is better100%

Visual perception

Atomic visual perception

PerceptionBench
  1. 1GPT-5.6 Solmax effort · Moonshot AI report59.7%
  2. 2Kimi K3max effort · Moonshot AI report58.5%
  3. 3Claude Fable 5max effort · Moonshot AI report57.2%
  4. 4GPT-5.5xhigh effort · Moonshot AI report55.8%
  5. 5Claude Opus 4.8max effort · Moonshot AI report47.2%
0Higher is better100%

Knowledge work

Professional deliverables

ALE · pass rate
  1. 1GPT-6 Astramax effort · ALE API snapshot34.2%
  2. 2Claude Sonnet 5Reported configuration · Google report33.3%
  3. 3Claude Opus 5high effort · ALE API snapshot32.2%
  4. 4Muse Spark 1.3xhigh effort · ALE API snapshot32.2%
  5. 5DeepSeek V4.1 FlashReported configuration · DeepSeek report31.8%
0Higher is better100%

Multimodal

Visual knowledge and reasoning

MMMU-Pro · selected setups
  1. 1Gemini 3.5 FlashReported configuration · Google report83.6%
  2. 2GPT-5.6 Solmax effort · Moonshot AI report83.0%
  3. 3Kimi K3max effort · Moonshot AI report81.6%
  4. 4Claude Fable 5max effort · Moonshot AI report81.2%
  5. 5GPT-5.5xhigh effort · Moonshot AI report81.2%
0Higher is better100%

Exploit capability

Offensive exploitation

ExploitBench · coverage
  1. 1Claude Mythos 5Reported configuration · OpenAI report78.0%
  2. 2Claude Mythos PreviewReported configuration · ExploitBench · all77.7%
  3. 3GPT-5.6 SolReported configuration · OpenAI report73.5%
  4. 4GPT-5.5Reported configuration · ExploitBench · all72.3%
  5. 5GPT-5.6 TerraReported configuration · OpenAI report52.9%
0Higher is better100%

Tool use

Multi-tool workflows

MCP Atlas
  1. 1Muse Spark 1.1Reported configuration · MCP Atlas88.1%
  2. 2Claude Fable 5.1Reported configuration · MCP Atlas87.2%
  3. 3Claude Opus 5xhigh effort · MCP Atlas85.8%
  4. 4Claude Fable 5max effort · Moonshot AI report84.7%
  5. 5Qwen3.8-2.4T-A95Bxhigh effort · MCP Atlas84.5%
0Higher is better100%

Category lists are independent. Scores from different benchmarks and versions are never combined into a universal ranking.