Management Consulting Tasks (Internal)
Can the model produce useful work in a professional knowledge-work setting?
Knowledge Work7 ranked models7 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.6 SolOpenAI | 43.2% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 2 | GPT-5.6 TerraOpenAI | 37.2% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 3 | Claude Fable 5Anthropic | 35.5% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 4 | GPT-5.6 LunaOpenAI | 35.4% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 5 | Claude Opus 4.8Anthropic | 31.6% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 6 | GPT-5.5OpenAI | 31.3% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 7 | Gemini 3.1 Pro PreviewGoogle | 13.2% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source