Static Jailbreak EvaluationBrowse 296

Static Jailbreak Evaluation

Does the model follow the publisher's safety policy under challenging prompts?

Latest stableGPT-6 System Card · biological high-risk category
Safety2 ranked models2 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

2
RankModelBest scoreBest reported setting
1GPT-6 SolOpenAI85.8%Reported configurationOpenAI report1 value · 1 reportSep 22, 2026 · source
2GPT-6 LunaOpenAI73.8%Reported configurationOpenAI report1 value · 1 reportSep 22, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Does the model follow the publisher's safety policy under challenging prompts?

Publisher-defined safety behavior, robustness, or preparedness evaluations under the stated test protocol. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

OpenAI's system card prints defender success with 95% confidence intervals: GPT-6 Sol 85.8% (81.8–88.9), GPT-6 Luna 73.8% (69.1–77.9). The card cautions that Luna's higher scores may reflect broader refusal behavior, including on legitimate requests.

Method / source