Static Jailbreak Evaluation
Does the model follow the publisher's safety policy under challenging prompts?
Top rankings
One row per model, using its best reported score across effort settings.
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundaryinternal methodology
Does the model follow the publisher's safety policy under challenging prompts?
Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
OpenAI's system card prints defender success with 95% confidence intervals: GPT-6 Sol 85.8% (81.8–88.9), GPT-6 Luna 73.8% (69.1–77.9). The card cautions that Luna's higher scores may reflect broader refusal behavior, including on legitimate requests.
Method / source