GPQA
Can the model reason through specialist biology, physics, and chemistry questions?
Science27 ranked models41 reported values7 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Mythos PreviewAnthropic | 94.6% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 2 | GPT-5.6 SolOpenAI | 94.6% | Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source |
| 3 | Gemini 3.1 Pro PreviewGoogle | 94.3% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 4 | Gemini 3.1 ProGoogle | 94.3% | Reported configurationZ.ai report3 values · 3 reportsJun 16, 2026 · source |
| 5 | Claude Mythos 5Anthropic | 94.1% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 6 | GPT-5.5OpenAI | 93.6% | Reported configurationOpenAI report3 values · 3 reportsJul 9, 2026 · source |
| 7 | Claude Opus 4.8Anthropic | 93.6% | Reported configurationZ.ai report3 values · 3 reportsJun 16, 2026 · source |
| 8 | Kimi K3Moonshot AI | 93.5% | max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source |
| 9 | MiniMax M3MiniMax | 93.0% | Reported configurationZ.ai report1 value · 1 reportJun 16, 2026 · source |
| 10 | GPT-5.6 TerraOpenAI | 92.9% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 11 | Claude Fable 5Anthropic | 92.6% | max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source |
| 12 | GPT-5.2OpenAI | 92.4% | xhigh effortGoogle report2 values · 2 reportsFeb 19, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology
Can the model reason through specialist biology, physics, and chemistry questions?
Multiple-choice science questions written and validated by domain experts. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
DeepSeek's official V4.1-Flash release table; exact source label GPQA Diamond; harness and effort are not printed. No tools.
Method / source