ExploitBenchBrowse 296

ExploitBench

Can the model diagnose or complete difficult cybersecurity tasks?

Latest stableInternal Port · June–August 2026
Cybersecurity2 ranked models2 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

2
RankModelBest scoreBest reported setting
1GPT-6 SolOpenAI5.5%max effortOpenAI report1 value · 1 reportSep 22, 2026 · source
2GPT-6 LunaOpenAI0.0%Reported configurationOpenAI report1 value · 1 reportSep 22, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model diagnose or complete difficult cybersecurity tasks?

Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

OpenAI reports maximum success rates from its June–August 2026 Internal Port evaluation: GPT-6 Sol at Max 5.5%, GPT-6 Luna 0%. This private Port result is separate from public ExploitBench v8 and must not be compared as the same task set or metric.

Method / source
ExploitBench · Benchmaxxer