SWE-Bench
Can the model complete substantial programming work under this benchmark's agent setup?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Opus 5Anthropic | 99.40% | xhigh effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 2 | Claude Fable 5.1Anthropic | 99.10% | high effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 3 | Kimi K3Moonshot AI | 97.70% | max effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 4 | GPT-6 AstraOpenAI | 96.90% | high effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 5 | GLM-5.3Z.ai | 95.60% | max effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 6 | GPT-5.6 SolOpenAI | 95.50% | xhigh effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 7 | Gemini 3.8 FlashGoogle | 94.86% | high effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 8 | Claude Sonnet 5Anthropic | 93.15% | xhigh effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 9 | GPT-5.6 TerraOpenAI | 92.37% | xhigh effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
| 10 | InklingThinking Machines Lab | 89.88% | xhigh effortSWE-Bench Pro V2 · Public Full1 value · 1 reportSep 22, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Resolve Rate on the official 642-task Public Full split across 11 repositories.
Scale's V2 release changes task composition and protocol relative to the 731-task V1 Public set. Compare only rows in this explicit V2 Full view; exact harness, effort, source label, rank, and displayed uncertainty are retained per model.
Exact Scale AI label Claude Opus 5 (Claude Code) xhigh; harness Claude Code; effort xhigh; rank 1; displayed score 99.40%; displayed uncertainty ±0.40; swe-bench-pro-v2-public-642. Exact Scale AI label Fable 5.1 (Claude Code) high; harness Claude Code; effort high; rank 2; displayed score 99.10%; displayed uncertainty ±0.50; swe-bench-pro-v2-public-642. Exact Scale AI label Kimi-K3 (mini-swe-agent) max; harness mini-swe-agent; effort max; rank 3; displayed score 97.70%; displayed uncertainty ±0.90; swe-bench-pro-v2-public-642. Exact Scale AI label GPT - 6- Astra (Codex) high; harness Codex; effort high; rank 4; displayed score 96.90%; displayed uncertainty ±1.10; swe-bench-pro-v2-public-642. Exact Scale AI label GLM-5.3 (mini-swe-agent) max; harness mini-swe-agent; effort max; rank 5; displayed score 95.60%; displayed uncertainty ±1.30; swe-bench-pro-v2-public-642. Exact Scale AI label GPT-5.6-sol (codex) xhigh; harness Codex; effort xhigh; rank 6; displayed score 95.50%; displayed uncertainty ±1.40; swe-bench-pro-v2-public-642. Exact Scale AI label Gemini 3.8 Flash (mini-swe-agent) high; harness mini-swe-agent; effort high; rank 7; displayed score 94.86%; displayed uncertainty ±1.46; swe-bench-pro-v2-public-642. Exact Scale AI label Claude Sonnet 5 (Claude Code) xhigh; harness Claude Code; effort xhigh; rank 8; displayed score 93.15%; displayed uncertainty ±1.71; swe-bench-pro-v2-public-642. Exact Scale AI label GPT-5.6-Terra (codex) xhigh; harness Codex; effort xhigh; rank 9; displayed score 92.37%; displayed uncertainty ±1.81; swe-bench-pro-v2-public-642. Exact Scale AI label Inkling (mini-swe-agent) xhigh; harness mini-swe-agent; effort xhigh; rank 10; displayed score 89.88%; displayed uncertainty ±2.10; swe-bench-pro-v2-public-642.
Method / source