GPQABrowse 296

GPQA

Can the model reason through specialist biology, physics, and chemistry questions?

Canonical static setDiamond
Science27 ranked models41 reported values7 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

27
RankModelBest scoreBest reported setting
1Claude Mythos PreviewAnthropic94.6%Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source
2GPT-5.6 SolOpenAI94.6%Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source
3Gemini 3.1 Pro PreviewGoogle94.3%Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source
4Gemini 3.1 ProGoogle94.3%Reported configurationZ.ai report3 values · 3 reportsJun 16, 2026 · source
5Claude Mythos 5Anthropic94.1%Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source
6GPT-5.5OpenAI93.6%Reported configurationOpenAI report3 values · 3 reportsJul 9, 2026 · source
7Claude Opus 4.8Anthropic93.6%Reported configurationZ.ai report3 values · 3 reportsJun 16, 2026 · source
8Kimi K3Moonshot AI93.5%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
9MiniMax M3MiniMax93.0%Reported configurationZ.ai report1 value · 1 reportJun 16, 2026 · source
10GPT-5.6 TerraOpenAI92.9%Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source
11Claude Fable 5Anthropic92.6%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
12GPT-5.2OpenAI92.4%xhigh effortGoogle report2 values · 2 reportsFeb 19, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model reason through specialist biology, physics, and chemistry questions?

Multiple-choice science questions written and validated by domain experts. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

DeepSeek's official V4.1-Flash release table; exact source label GPQA Diamond; harness and effort are not printed. No tools.

Method / source