Internal Company Research ReportsBrowse 296

Internal Company Research Reports

Can the model produce useful work in a professional knowledge-work setting?

Version not specifiedExact reported variant
Knowledge Work3 ranked models3 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

3
RankModelBest scoreBest reported setting
1Claude Opus 5.5Anthropic16Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source
2Claude Fable 5.1Anthropic0Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source
3Claude Opus 5Anthropic0Reported configurationAnthropic report1 value · 1 reportSep 22, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model produce useful work in a professional knowledge-work setting?

Professional analysis, document, finance, office, or domain-specific tasks. Higher is better. This benchmark reports points rather than percent correct.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Across different effort settings, Anthropic says 16 of 18 Opus 5.5 reports cleared its quality bar; Claude Fable 5.1 and Claude Opus 5 cleared no attempt. The report describes this internal task and grader but does not publish per-effort results.

Method / source