HealthBenchBrowse 296

HealthBench

Can the model produce useful work in a professional knowledge-work setting?

Latest stableGPT-6 system-card appendix · September 2026 length-adjusted
Knowledge Work7 ranked models7 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

7
RankModelBest scoreBest reported setting
1GPT-6 AstraOpenAI58.3%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
2GPT-5.6 SolOpenAI57.0%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
3GPT-5.6 TerraOpenAI57.0%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
4GPT-5.5OpenAI56.5%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
5GPT-5.6 LunaOpenAI55.8%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
6GPT-6 LunaOpenAI54.5%Reported configurationOpenAI report1 value · 1 reportSep 22, 2026 · source
7GPT-6 SolOpenAI53.2%Reported configurationOpenAI report1 value · 1 reportSep 22, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. Length-adjusted score on the 0-100 scale reported by OpenAI.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

OpenAI's September 22 appendix corrects the earlier Astra value to 58.3. The table prints length-adjusted score followed by unadjusted score and mean response length in characters: GPT-6 Astra 58.3 (60.8, 977); GPT-6 Sol 53.2 (47.1, 977); GPT-6 Luna 54.5 (50.0, 1255). Earlier GPT-5.5 and GPT-5.6 comparison values are retained from the same report: 56.5 (58.4, 2313); 57.0 (55.6, 1764); 57.0 (58.7, 2285); 55.8 (55.4, 1930).

Method / source