SRE-Bench
Can the model diagnose or complete difficult cybersecurity tasks?
Top rankings
One row per model, using its best reported score across effort settings.
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundaryinternal methodology
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of reverse-engineering challenges fully solved across four trials.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
OpenAI reports pass@4 on 262 binary instances derived from 19 privately developed programs; a challenge is fully solved only when all six task-specific objectives are satisfied. GPT-6 Astra is 99.2%; GPT-5.6 Sol is 68.7%. Output-token efficiency is described qualitatively and is not quantified here.
Method / source