ExploitBench
Can the model diagnose or complete difficult cybersecurity tasks?
Top rankings
One row per model, using its best reported score across effort settings.
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundaryinternal methodology
Can the model diagnose or complete difficult cybersecurity tasks?
Security analysis or exploitation under a controlled evaluation environment. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
OpenAI reports maximum success rates from its June–August 2026 Internal Port evaluation: GPT-6 Sol at Max 5.5%, GPT-6 Luna 0%. This private Port result is separate from public ExploitBench v8 and must not be compared as the same task set or metric.
Method / source