Benchmark overview
Five leading models in each audited category, using one exact representative benchmark per category.
- Families
- 296
- Categories
- 9
- Updated
- Sep 23, 2026
Coding
Repository code-change mergeability
- 1Claude Opus 5.5medium effort · Cognition leaderboard54.64%
- 2Claude Fable 5xhigh effort · Cognition leaderboard53.48%
- 3Claude Opus 5medium effort · Cognition leaderboard53.38%
- 4GPT-6 Astramax effort · Cognition leaderboard53.26%
- 5Claude Fable 5.1medium effort · Cognition leaderboard50.91%
0Higher is better100%
Agents
Terminal agent work
- 1Claude Opus 5.5xhigh effort · Anthropic report66.40%
- 2Mythos 5.1Reported configuration · Anthropic report60.90%
- 3GPT-6 Astramax effort · Terminal-Bench 4.0 · official leaderboard58.18%
- 4Claude Fable 5.1max effort · Terminal-Bench 4.0 · official leaderboard57.88%
- 5Claude Opus 5xhigh effort · Terminal-Bench 4.0 · official leaderboard53.94%
0Higher is better100%
Expert reasoning
Broad expert questions
- 1Claude Mythos PreviewReported configuration · Anthropic report56.80%
- 2GPT-6 AstraReported configuration · HLE · full54.80%
- 3Claude Fable 5max effort · Moonshot AI report53.30%
- 4Muse Spark 1.1xhigh effort · Meta report52.20%
- 5Claude Opus 4.8max effort · Moonshot AI report49.80%
0Higher is better100%
Science
Scientific research workflows
- 1GPT-6 Astramax effort · Terminal-Bench-Science 0.1 · official leaderboard68.1%
- 2Claude Fable 5.1max effort · Terminal-Bench-Science 0.1 · official leaderboard40.0%
- 3Claude Opus 5max effort · Terminal-Bench-Science 0.1 · official leaderboard30.0%
- 4GPT-5.6 Solmax effort · Terminal-Bench-Science 0.1 · official leaderboard22.4%
- 5Claude Fable 5max effort · Terminal-Bench-Science 0.1 · official leaderboard21.4%
0Higher is better100%
Visual perception
Atomic visual perception
- 1GPT-5.6 Solmax effort · Moonshot AI report59.7%
- 2Kimi K3max effort · Moonshot AI report58.5%
- 3Claude Fable 5max effort · Moonshot AI report57.2%
- 4GPT-5.5xhigh effort · Moonshot AI report55.8%
- 5Claude Opus 4.8max effort · Moonshot AI report47.2%
0Higher is better100%
Knowledge work
Professional deliverables
- 1GPT-6 Astramax effort · ALE API snapshot34.2%
- 2Claude Sonnet 5Reported configuration · Google report33.3%
- 3Claude Opus 5high effort · ALE API snapshot32.2%
- 4Muse Spark 1.3xhigh effort · ALE API snapshot32.2%
- 5DeepSeek V4.1 FlashReported configuration · DeepSeek report31.8%
0Higher is better100%
Multimodal
Visual knowledge and reasoning
- 1Gemini 3.5 FlashReported configuration · Google report83.6%
- 2GPT-5.6 Solmax effort · Moonshot AI report83.0%
- 3Kimi K3max effort · Moonshot AI report81.6%
- 4Claude Fable 5max effort · Moonshot AI report81.2%
- 5GPT-5.5xhigh effort · Moonshot AI report81.2%
0Higher is better100%
Exploit capability
Offensive exploitation
- 1Claude Mythos 5Reported configuration · OpenAI report78.0%
- 2Claude Mythos PreviewReported configuration · ExploitBench · all77.7%
- 3GPT-5.6 SolReported configuration · OpenAI report73.5%
- 4GPT-5.5Reported configuration · ExploitBench · all72.3%
- 5GPT-5.6 TerraReported configuration · OpenAI report52.9%
0Higher is better100%
Tool use
Multi-tool workflows
- 1Muse Spark 1.1Reported configuration · MCP Atlas88.1%
- 2Claude Fable 5.1Reported configuration · MCP Atlas87.2%
- 3Claude Opus 5xhigh effort · MCP Atlas85.8%
- 4Claude Fable 5max effort · Moonshot AI report84.7%
- 5Qwen3.8-2.4T-A95Bxhigh effort · MCP Atlas84.5%
0Higher is better100%
Category lists are independent. Scores from different benchmarks and versions are never combined into a universal ranking.