Production Benchmarks
Does the model follow the publisher's safety policy under challenging prompts?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.3 InstantOpenAI | 0.943 | Reported configurationOpenAI report1 value · 1 reportAug 6, 2026 · source |
| 2 | GPT-5.5OpenAI | 0.943 | Reported configurationOpenAI report3 values · 1 reportAug 6, 2026 · source |
| 3 | GPT-5.6 LunaOpenAI | 0.925 | Reported configurationOpenAI report1 value · 1 reportAug 6, 2026 · source |
| 4 | GPT-5.6 SolOpenAI | 0.906 | Reported configurationOpenAI report1 value · 1 reportAug 6, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarylimited methodology
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. Raw not_unsafe share in the publisher's [0,1] scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Production Benchmarks with Challenging Prompts; evaluations run without system-level safeguards on difficult production-derived examples.
Method / source