SRE-BenchBrowse 296

SRE-Bench

Can the model diagnose or complete difficult cybersecurity tasks?

Latest stableGPT-6 Astra System Card · pass@4
Cybersecurity2 ranked models2 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

2
RankModelBest scoreBest reported setting
1GPT-6 AstraOpenAI99.2%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
2GPT-5.6 SolOpenAI68.7%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of reverse-engineering challenges fully solved across four trials.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

OpenAI reports pass@4 on 262 binary instances derived from 19 privately developed programs; a challenge is fully solved only when all six task-specific objectives are satisfied. GPT-6 Astra is 99.2%; GPT-5.6 Sol is 68.7%. Output-token efficiency is described qualitatively and is not quantified here.

Method / source