DECK-Bench (Internal)Browse 296

DECK-Bench (Internal)

Can the model produce useful work in a professional knowledge-work setting?

Version not specifiedExact reported variant
Knowledge Work6 ranked models6 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

6
RankModelBest scoreBest reported setting
1GPT-5.6 SolOpenAI74.7%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
2Kimi K3Moonshot AI73.5%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
3Claude Fable 5Anthropic73.0%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
4GLM-5.2Z.ai68.6%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
5GPT-5.5OpenAI68.2%xhigh effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
6Claude Opus 4.8Anthropic66.9%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source