Internal Company Research Reports
Can the model produce useful work in a professional knowledge-work setting?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Opus 5.5Anthropic | 16 | Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source |
| 2 | Claude Fable 5.1Anthropic | 0 | Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source |
| 3 | Claude Opus 5Anthropic | 0 | Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundaryinternal methodology
Can the model produce useful work in a professional knowledge-work setting?
Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. This benchmark reports points rather than percent correct.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Across different effort settings, Anthropic says 16 of 18 Opus 5.5 reports cleared its quality bar; Claude Fable 5.1 and Claude Opus 5 cleared no attempt. The report describes this internal task and grader but does not publish per-effort results.
Method / source