DECK-Bench (Internal)
Can the model produce useful work in a professional knowledge-work setting?
Knowledge Work6 ranked models6 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.6 SolOpenAI | 74.7% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 2 | Kimi K3Moonshot AI | 73.5% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 3 | Claude Fable 5Anthropic | 73.0% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 4 | GLM-5.2Z.ai | 68.6% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 5 | GPT-5.5OpenAI | 68.2% | xhigh effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 6 | Claude Opus 4.8Anthropic | 66.9% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source