Frontier Red Team
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Mythos PreviewAnthropic | 0.819891 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 2 | Claude Opus 5Anthropic | 0.807631 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 3 | Kimi K3Moonshot AI | 0.806478 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 4 | Claude Opus 4.8Anthropic | 0.805502 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 5 | Claude Opus 4.5Anthropic | 0.805056 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 6 | Claude Opus 4.6Anthropic | 0.803937 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 7 | Claude Mythos 5Anthropic | 0.797919 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 8 | Claude Opus 4.7Anthropic | 0.795413 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
| 9 | Claude Sonnet 5Anthropic | 0.785770 | Reported configurationAnthropic report1 value · 1 reportSep 10, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarylimited methodology
How does a model perform on this Anthropic-defined Frontier Red Team evaluation?
Aggregate capability measurements from Anthropic's publisher-defined intelligence and simulated weapons-development evaluations. Higher is better. Cell-member classification F1 on the easy tier.
This is an Anthropic-defined evaluation rather than a common benchmark leaderboard. The task set, harness, model variant, tool access, reasoning setting, and safety constraints materially affect interpretation; compare only within this reporting context.
Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.78577; source interval 0.753319–0.818222; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.805056; source interval 0.782432–0.827679; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.803937; source interval 0.777644–0.83023; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.795413; source interval 0.76412–0.826706; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.805502; source interval 0.772156–0.838848; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.807631; source interval 0.773428–0.841835; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.819891; source interval 0.795897–0.843885; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.797919; source interval 0.765779–0.830058; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained. Anthropic's embedded identity-classification artifact reports cell-member classification F1 by difficulty tier; this row retains exact means, source intervals, and n=18. The artifact's upper-bound caveat is preserved in the context note. Identity classification · easy; exact aggregate mean 0.806478; source interval 0.777111–0.835844; n=18. Opus 4.5 and Opus 4.6 used a fixed 12k-token thinking budget where shown; all other models used adaptive thinking. Raw task material is not retained.
Method / source