Sandbox BenchBrowse 296

Sandbox Bench

Can the model diagnose or complete difficult cybersecurity tasks?

Latest stableGPT-6 Astra System Card · 22-target private evaluation
Cybersecurity2 ranked models2 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

2
RankModelBest scoreBest reported setting
1GPT-6 AstraOpenAI45.5%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source
2GPT-5.6 SolOpenAI4.5%Reported configurationOpenAI report1 value · 1 reportSep 3, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. Percentage of private targets on which the model recovered the protected flag.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

OpenAI reports Astra on 10 of 22 targets (45.5%) and GPT-5.6 Sol on 1 of 22 (4.5%); Astra's successes were four of five runtime targets, five of fourteen parser targets, and one of three egress proxies. The affected software details are withheld during coordinated disclosure.

Method / source