DeepSWE
Can the model complete substantial programming work under this benchmark's agent setup?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | DeepSeek V4.1 FlashDeepSeek | 74.2% | Reported configurationDeepSeek report1 value · 1 reportSep 10, 2026 · source |
| 2 | GPT-6 AstraOpenAI | 74.1% | xhigh effortDeepSWE v1.1 feed5 values · 1 reportSep 22, 2026 · source |
| 3 | Claude Opus 5Anthropic | 74.0% | Reported configurationGoogle DeepMind report6 values · 2 reportsSep 2, 2026 · source |
| 4 | Gemini 3.8 FlashGoogle | 73.8% | high effortDeepSWE v1.1 feed3 values · 2 reportsSep 22, 2026 · source |
| 5 | GPT-5.6 SolOpenAI | 72.7% | Reported configurationGoogle DeepMind report7 values · 3 reportsSep 2, 2026 · source |
| 6 | Grok 4.7SpaceXAI | 71.0% | high effortSpaceXAI report1 value · 1 reportSep 21, 2026 · source |
| 7 | Claude Fable 5Anthropic | 70.0% | max effortSpaceXAI report7 values · 3 reportsJul 16, 2026 · source |
| 8 | GPT-5.5OpenAI | 70.0% | Reported configurationZ.ai report8 values · 5 reportsJun 16, 2026 · source |
| 9 | GPT-5.6 TerraOpenAI | 69.6% | max effortDeepSWE v1.1 feed8 values · 4 reportsSep 22, 2026 · source |
| 10 | GLM-5.3Z.ai | 69.0% | max effortDeepSWE v1.1 feed1 value · 1 reportSep 22, 2026 · source |
| 11 | GPT-6 SolOpenAI | 68.8% | max effortOpenAI report1 value · 1 reportSep 22, 2026 · source |
| 12 | Kimi K3Moonshot AI | 68.5% | max effortDeepSWE v1.1 feed1 value · 1 reportSep 22, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
Can the model complete substantial programming work under this benchmark's agent setup?
Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Percentage of scored rollout attempts that passed.
Pass@1 is an attempt-level rate. Task coverage, repeated runs, confidence interval, mini-swe-agent version, model variant, and reasoning effort remain material.
Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 295/430 scored rollout attempts passed; 98/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.674838084453603–0.6972549388022109 with half-width 0.011208427174303877. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $9.177638214788733, median rollout cost $7.322686000000002, mean input tokens 4833178.365116279, mean cache tokens 4711714.3162790695, and mean output tokens 57287.06976744186; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 258/433 scored rollout attempts passed; 92/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5679029016687951–0.6237830105713897 with half-width 0.027940054451297294. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.757873125874126, median rollout cost $2.7559204999999998, mean input tokens 1691250.452655889, mean cache tokens 1624197.8868360277, and mean output tokens 25242.833718244805; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 304/436 scored rollout attempts passed; 95/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.656912963913939–0.7375824489300977 with half-width 0.04033474250807936. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $21.634702092592594, median rollout cost $19.23457225, mean input tokens 12638832.12614679, mean cache tokens 12388674.334862385, and mean output tokens 118592.77293577982; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 285/436 scored rollout attempts passed; 94/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6094504359782605–0.6978890135630239 with half-width 0.044219288792381656. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $6.088186528935186, median rollout cost $4.62221125, mean input tokens 2942251.607798165, mean cache tokens 2848228.9174311925, and mean output tokens 40201.35321100918; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 316/452 scored rollout attempts passed; 100/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6666658684930744–0.7315642200025008 with half-width 0.03244917575471319. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $13.414521495535714, median rollout cost $11.341082999999998, mean input tokens 7329160.601769911, mean cache tokens 7159496.181415929, and mean output tokens 80352.1703539823; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 234/452 scored rollout attempts passed; 88/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4720830235866894–0.5633152065018062 with half-width 0.04561609145755843. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.281940017699115, median rollout cost $3.6635345, mean input tokens 4908687.0265486725, mean cache tokens 4807501.1283185845, and mean output tokens 50063.52654867257; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 184/451 scored rollout attempts passed; 77/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3933464140551067–0.42261810922648974 with half-width 0.014635847585691517. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.293369216740577, median rollout cost $1.8240152500000004, mean input tokens 2417987.955654102, mean cache tokens 2354415.691796009, and mean output tokens 28922.736141906873; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 253/429 scored rollout attempts passed; 88/111 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5720954362692973–0.6073917432178823 with half-width 0.01764815347429254. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $13.222583593240094, median rollout cost $12.353505999999996, mean input tokens 17047938.554778554, mean cache tokens 16817766.494172495, and mean output tokens 135031.67132867133; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 220/452 scored rollout attempts passed; 86/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4643336050994396–0.5091177223341886 with half-width 0.02239205861737452. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.4438571028761062, median rollout cost $2.84043425, mean input tokens 3843602.6681415928, mean cache tokens 3757520.9115044246, and mean output tokens 41313.45132743363; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 243/447 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5064898547928808–0.5807584673547702 with half-width 0.03713430628094476. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $8.006355704138702, median rollout cost $7.049357750000001, mean input tokens 9854652.170022372, mean cache tokens 9693014.304250559, and mean output tokens 86088.7874720358; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 327/449 scored rollout attempts passed; 99/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7088307563794007–0.7477393995226038 with half-width 0.019454321571601495. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $6.076084489977728, median rollout cost $4.96882375, mean input tokens 7233879.0824053455, mean cache tokens 7085027.516703786, and mean output tokens 64207.43429844098; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 261/449 scored rollout attempts passed; 96/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5579848799986361–0.6045986389323216 with half-width 0.023306879466842723. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.6626223919821828, median rollout cost $1.2162672499999996, mean input tokens 1564287.6570155902, mean cache tokens 1497598.9933184856, and mean output tokens 19884.3140311804; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 327/444 scored rollout attempts passed; 100/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6977633822227692–0.7752095907502038 with half-width 0.03872310426371729. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $11.837583271396396, median rollout cost $10.4281505, mean input tokens 15025834.39864865, mean cache tokens 14784785.963963963, and mean output tokens 117565.6936936937; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 308/447 scored rollout attempts passed; 101/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6773124446609795–0.7007636179788415 with half-width 0.011725586658930944. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.28978764541387, median rollout cost $2.4819885000000004, mean input tokens 3579815.3601789707, mean cache tokens 3479530.087248322, and mean output tokens 36981.530201342284; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 327/447 scored rollout attempts passed; 97/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7009338059446892–0.7621534423774585 with half-width 0.030609818216384643. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $9.07220232326622, median rollout cost $7.928816750000001, mean input tokens 11292729.463087248, mean cache tokens 11094990.232662192, and mean output tokens 91672.01565995526; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 135/451 scored rollout attempts passed; 64/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.2584133988380421–0.34025622422182483 with half-width 0.04092141269189136. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $5.5223517987804875, median rollout cost $4.867440150000001, mean input tokens 12870627.738359202, mean cache tokens 12719994.494456762, and mean output tokens 76160.31263858093; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 218/452 scored rollout attempts passed; 90/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4372377362536038–0.5273640336579006 with half-width 0.045063148702148406. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $7.425560161197339, median rollout cost $5.833649100000001, mean input tokens 18254491.34368071, mean cache tokens 18068889.902439024, and mean output tokens 87294.75609756098; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 137/449 scored rollout attempts passed; 65/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.2938623190176438–0.31638266984649877 with half-width 0.011260175414427497. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.1865550832962137, median rollout cost $1.6764955500000003, mean input tokens 4534475.699331849, mean cache tokens 4449592.527839644, and mean output tokens 35595.08240534521; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 238/442 scored rollout attempts passed; 89/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4960923737597946–0.5808307031632823 with half-width 0.04236916470174386. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $26.399858950791852, median rollout cost $23.277376725000003, mean input tokens 72415756.48642534, mean cache tokens 71991474.90950227, and mean output tokens 214117.54977375566; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 179/450 scored rollout attempts passed; 73/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3665020157144805–0.4290535398410751 with half-width 0.03127576206329727. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.079036229, median rollout cost $3.3573110999999995, mean input tokens 9286962.313333333, mean cache tokens 9158953.993333334, and mean output tokens 56816.753333333334; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 224/451 scored rollout attempts passed; 85/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4621285677328888–0.5312195475664461 with half-width 0.034545489916778645. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $11.890634540022173, median rollout cost $9.965361750000001, mean input tokens 30695627.87804878, mean cache tokens 30442947.99556541, and mean output tokens 120698.59201773835; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 241/452 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4975163385576403–0.5688553428582889 with half-width 0.035669502150324287. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.10023560370265487, median rollout cost $0.09124898719999996, mean input tokens 19829931.451327432, mean cache tokens 19719518.867256638, and mean output tokens 107687.02212389381; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 284/452 scored rollout attempts passed; 100/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5649842780985077–0.6916528900430852 with half-width 0.06333430597228876. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.24138687892256638, median rollout cost $0.2232531505000001, mean input tokens 24191606.099557523, mean cache tokens 24049100.743362833, and mean output tokens 105998.91814159292; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 53/452 scored rollout attempts passed; 32/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.10244568255705315–0.13206759177923005 with half-width 0.01481095461108845. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.1434045568888886, median rollout cost $1.7235064000000002, mean input tokens 2403677.331111111, mean cache tokens 1678992.7311111111, and mean output tokens 28368.884444444444; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 163/452 scored rollout attempts passed; 72/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3209564389213443–0.40028249913175307 with half-width 0.03966303010520439. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.4467114588691796, median rollout cost $2.8624087500000015, mean input tokens 7629071.858093127, mean cache tokens 6428494.862527716, and mean output tokens 75730.19290465632; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 211/452 scored rollout attempts passed; 85/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4297656202995626–0.5038626982845081 with half-width 0.037048538992472776. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.419013277333334, median rollout cost $3.444695325000001, mean input tokens 12595957.344444444, mean cache tokens 11254636.448888889, and mean output tokens 95844.86222222222; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 295/452 scored rollout attempts passed; 93/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6347762421887939–0.6705334923244803 with half-width 0.017878625067843174. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.1762860351216813, median rollout cost $1.9044440625, mean input tokens 16999853.159292035, mean cache tokens 16260422.17920354, and mean output tokens 107248.30309734514; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 243/452 scored rollout attempts passed; 88/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5117141286507079–0.5635071102873452 with half-width 0.02589649081831869. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.8322645629424779, median rollout cost $1.3913454375, mean input tokens 15056565.075221239, mean cache tokens 14422633.818584071, and mean output tokens 73364.9557522124; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 296/452 scored rollout attempts passed; 94/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.624001933873013–0.6857325794013234 with half-width 0.030865322764155222. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.0251351826880533, median rollout cost $1.65422055, mean input tokens 15332065.907079646, mean cache tokens 14557599.913716814, and mean output tokens 93990.8517699115; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 330/447 scored rollout attempts passed; 97/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7240819014906883–0.7524281656234058 with half-width 0.014173132066358783. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.362349413758389, median rollout cost $2.1098912999999997, mean input tokens 21734449.25279642, mean cache tokens 21456039.230425056, and mean output tokens 143242.6644295302; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 321/452 scored rollout attempts passed; 94/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6873689454216633–0.7329850368792218 with half-width 0.02280804572877925. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.9671024789823006, median rollout cost $1.7342966249999998, mean input tokens 16863939.814159293, mean cache tokens 16519023.008849557, and mean output tokens 124684.36283185841; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 164/452 scored rollout attempts passed; 77/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3153311289278631–0.41033258788629623 with half-width 0.047500729479216554. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.8355164725663715, median rollout cost $2.25012998, mean input tokens 9070605.707964601, mean cache tokens 8861413.805309735, and mean output tokens 54245.50442477876; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 197/450 scored rollout attempts passed; 87/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.42052148157196967–0.45503407398358586 with half-width 0.01725629620580811. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.9199075743111114, median rollout cost $3.4637479200000003, mean input tokens 12606560.215555556, mean cache tokens 12344954.453333333, and mean output tokens 78175.30666666667; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 284/448 scored rollout attempts passed; 96/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.590148863911635–0.6777082789455078 with half-width 0.04377970751693644. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.48196371245535713, median rollout cost $0.389975125, mean input tokens 12387025.401785715, mean cache tokens 11770874.857142856, and mean output tokens 72829.77008928571; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 311/451 scored rollout attempts passed; 99/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6593922012533708–0.7197652266845449 with half-width 0.030186512715587. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.9933584893126386, median rollout cost $3.1315945999999992, mean input tokens 13272344.266075388, mean cache tokens 13106008.833702883, and mean output tokens 80435.60975609756; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 234/452 scored rollout attempts passed; 88/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.502678065476865–0.5327201646116306 with half-width 0.015021049567382833. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $5.6524577245575225, median rollout cost $4.3528565, mean input tokens 9418920.763274336, mean cache tokens 8613618.407079646, and mean output tokens 71408.87389380531; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 291/452 scored rollout attempts passed; 102/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6126368832394585–0.6749737362295679 with half-width 0.031168426495054677. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $5.100437004424779, median rollout cost $4.582879500000001, mean input tokens 4617781.876106195, mean cache tokens 4205870.442477876, and mean output tokens 31159.49778761062; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 122/452 scored rollout attempts passed; 54/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.24696647220470694–0.29285653664485056 with half-width 0.02294503222007181. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.2002025442477877, median rollout cost $1.108106, mean input tokens 776751.5840707965, mean cache tokens 658930.407079646, and mean output tokens 9442.774336283186; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 244/452 scored rollout attempts passed; 88/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5142921338966476–0.5653538838024673 with half-width 0.02553087495290979. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.749342734513274, median rollout cost $2.436942, mean input tokens 2438539.053097345, mean cache tokens 2229359.0088495575, and mean output tokens 19625.433628318584; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 303/452 scored rollout attempts passed; 100/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6056975188917963–0.7350104457099735 with half-width 0.06465646340908864. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $7.226236674778761, median rollout cost $6.110482499999999, mean input tokens 8372927.630530974, mean cache tokens 8032415.716814159, and mean output tokens 46294.72345132743; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 200/452 scored rollout attempts passed; 85/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.4132822036270294–0.47167354858536004 with half-width 0.029195672479165317. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.7779007499999999, median rollout cost $0.6635542999999999, mean input tokens 3372477.7876106193, mean cache tokens 3055611.469026549, and mean output tokens 25778.274336283186; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 7/452 scored rollout attempts passed; 5/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.0071835281016603–0.023789923225773328 with half-width 0.008303197562056514. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.07240623938053097, median rollout cost $0.0697851, mean input tokens 150031.52876106196, mean cache tokens 107102.01769911505, and mean output tokens 3127.754424778761; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 301/448 scored rollout attempts passed; 102/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6319549876432895–0.7117950123567105 with half-width 0.03992001235671055. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.028116810267857, median rollout cost $2.2920917999999997, mean input tokens 15443718.330357144, mean cache tokens 14666893.714285715, and mean output tokens 73399.70758928571; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 51/452 scored rollout attempts passed; 31/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.10452866084502313–0.12113505596913615 with half-width 0.008303197562056514. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.21630978982300886, median rollout cost $0.1994702, mean input tokens 617827.4800884955, mean cache tokens 500661.2389380531, and mean output tokens 8179.570796460177; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 257/452 scored rollout attempts passed; 89/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5469030532683623–0.5902650883245582 with half-width 0.021681017528097975. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.5356192685840708, median rollout cost $1.2449332, mean input tokens 7610577.53539823, mean cache tokens 7096152.3539823005, and mean output tokens 44677.900442477876; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 313/451 scored rollout attempts passed; 98/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.679691851565629–0.7083347559731736 with half-width 0.014321452203772355. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.4698331108647453, median rollout cost $2.979421, mean input tokens 2713572.2505543237, mean cache tokens 2435251.370288248, and mean output tokens 28450.31929046563; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 205/452 scored rollout attempts passed; 81/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.42965787629426033–0.47742176972343875 with half-width 0.0238819467145892. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.0742664646017699, median rollout cost $0.944557, mean input tokens 690833.657079646, mean cache tokens 599908.814159292, and mean output tokens 10579.141592920354; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 327/450 scored rollout attempts passed; 97/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6983684441682949–0.7549648891650385 with half-width 0.02829822249837175. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $8.386436346666667, median rollout cost $6.8414715, mean input tokens 7907652.344444444, mean cache tokens 7419435.235555556, and mean output tokens 60013.64444444444; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 276/452 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5947858925334765–0.6264530455196208 with half-width 0.01583357649307215. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.8620318075221238, median rollout cost $1.545414, mean input tokens 1505793.9557522123, mean cache tokens 1382962.9734513275, and mean output tokens 18425.216814159292; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 319/451 scored rollout attempts passed; 97/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6991259559533886–0.7155081903880748 with half-width 0.008191117217343103. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.703663860310422, median rollout cost $4.117039, mean input tokens 4259318.379157428, mean cache tokens 3967289.3303769403, and mean output tokens 40744.59201773836; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 243/452 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.49432091479689105–0.5809003241411621 with half-width 0.04328970467213552. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.1343762765486725, median rollout cost $1.031709, mean input tokens 1557052.774336283, mean cache tokens 1369338.3362831858, and mean output tokens 21517.03982300885; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 108/449 scored rollout attempts passed; 50/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.23276690748868134–0.24830213482757701 with half-width 0.007767613669447833. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.4277452605790646, median rollout cost $0.3989895, mean input tokens 480565.34521158127, mean cache tokens 401000.90868596884, and mean output tokens 8572.26280623608; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 314/451 scored rollout attempts passed; 100/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6706612091651647–0.7217999881740815 with half-width 0.025569389504458404. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.945847263858093, median rollout cost $4.114557, mean input tokens 9230560.541019956, mean cache tokens 8652201.720620843, and mean output tokens 71938.62527716186; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 158/450 scored rollout attempts passed; 68/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3172914680770448–0.3849307541451774 with half-width 0.03381964303406632. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $0.5832493688888889, median rollout cost $0.546099, mean input tokens 725100.8355555555, mean cache tokens 624756.0533333333, and mean output tokens 11746.56; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 272/452 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5805269394851535–0.6230128835236961 with half-width 0.021242972019271257. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.127191725663717, median rollout cost $1.8670102499999999, mean input tokens 3246355.4668141594, mean cache tokens 2926653.1681415928, and mean output tokens 39616.53982300885; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 331/452 scored rollout attempts passed; 93/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6980659243770221–0.7665358455344823 with half-width 0.03423496057873005. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.9237237909292038, median rollout cost $3.454243, mean input tokens 1283270.8030973452, mean cache tokens 1168811.0420353983, and mean output tokens 26505.853982300883; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 303/452 scored rollout attempts passed; 90/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6573453717840262–0.6833625928177437 with half-width 0.013008610516858731. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.595223439159292, median rollout cost $1.26315875, mean input tokens 494506.91150442476, mean cache tokens 444727.3495575221, and mean output tokens 10579.53982300885; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 331/452 scored rollout attempts passed; 90/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7239976873936956–0.7406040825178087 with half-width 0.008303197562056526. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $7.497827996681416, median rollout cost $6.712564, mean input tokens 2183033.681415929, mean cache tokens 1986676.7079646017, and mean output tokens 61148.50663716814; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 329/452 scored rollout attempts passed; 93/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7019796153763715–0.7537725970130089 with half-width 0.025896490818318695. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.0754742975663714, median rollout cost $2.6204002500000003, mean input tokens 1008238.2898230088, mean cache tokens 916944.3473451327, and mean output tokens 20361.809734513274; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 335/452 scored rollout attempts passed; 91/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.7124964807371247–0.7698044042186275 with half-width 0.02865396174075141. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.429117203539823, median rollout cost $3.8406789999999997, mean input tokens 1456927.0929203539, mean cache tokens 1326921.6836283186, and mean output tokens 29557.327433628318; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 243/452 scored rollout attempts passed; 88/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5148025737402474–0.5604186651978057 with half-width 0.0228080457287792. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.4157481725663716, median rollout cost $2.02311, mean input tokens 4059182.975663717, mean cache tokens 3943846.5132743362, and mean output tokens 35525.33185840708; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; high reasoning effort; 113 task set; 294/451 scored rollout attempts passed; 96/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6365486924786155–0.6672207088517614 with half-width 0.015336008186572951. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.3848732239467845, median rollout cost $3.693604, mean input tokens 7348488.252771619, mean cache tokens 7119379.299334812, and mean output tokens 61160.94456762749; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; low reasoning effort; 113 task set; 187/449 scored rollout attempts passed; 78/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3932945513980172–0.43966758668661526 with half-width 0.023186517644299014. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $1.0423604231625836, median rollout cost $0.742614, mean input tokens 1646919.8017817372, mean cache tokens 1566819.4922048997, and mean output tokens 16458.3429844098; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; medium reasoning effort; 113 task set; 305/452 scored rollout attempts passed; 95/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6519707153331677–0.697586806790726 with half-width 0.0228080457287792. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.4489837522123894, median rollout cost $2.945834, mean input tokens 5725195.139380531, mean cache tokens 5533327.0088495575, and mean output tokens 49763.99778761062; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 301/451 scored rollout attempts passed; 96/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.64564892272269–0.689162607210791 with half-width 0.02175684224405049. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $5.497668745011087, median rollout cost $4.63937, mean input tokens 9344022.093126386, mean cache tokens 9079197.80044346, and mean output tokens 71403.54323725056; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; default reasoning effort; 113 task set; 138/452 scored rollout attempts passed; 69/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.3003027179908134–0.31031675103573525 with half-width 0.005007016522460913. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.815536243274336, median rollout cost $2.1955079100000003, mean input tokens 13008041.482300885, mean cache tokens 12857596.451327434, and mean output tokens 59297.24778761062; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; max reasoning effort; 113 task set; 309/451 scored rollout attempts passed; 101/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.6397739833756912–0.7305142649613376 with half-width 0.045370140792823206. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $4.654682129933482, median rollout cost $3.6024987, mean input tokens 9695436.246119734, mean cache tokens 9501527.498891352, and mean output tokens 81499.84257206209; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 241/452 scored rollout attempts passed; 90/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5028324161686275–0.5635392652473017 with half-width 0.030353424539337114. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $2.361142644469026, median rollout cost $1.7171158, mean input tokens 12005399.261061948, mean cache tokens 11781947.471238937, and mean output tokens 74008.4203539823; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 248/452 scored rollout attempts passed; 92/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5274295943524101–0.5699155383909527 with half-width 0.02124297201927127. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.6955048180309733, median rollout cost $2.6840506, mean input tokens 13693824.435840707, mean cache tokens 12584988.949115044, and mean output tokens 99226.38053097345; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. Official DeepSWE v1.1 aggregate row; mini-swe-agent harness; xhigh reasoning effort; 113 task set; 258/449 scored rollout attempts passed; 94/113 tasks passed at least once; 4 repeated whole-benchmark runs; 95% run-to-run confidence interval 0.5479823182637887–0.6012381717139396 with half-width 0.026627926725075454. The feed reports pass@1 and pass@4 as ratios, plus mean rollout cost $3.7290646500445437, median rollout cost $3.132689060000001, mean input tokens 12996066.461024499, mean cache tokens 12477251.240534522, and mean output tokens 95075.18262806236; these costs and token aggregates are retained as source aggregates and are not normalized to cost per task. DeepSeek's official V4.1-Flash release table; exact v1.1 label retained. OpenAI reports GPT-6 Sol Max 68.8%, Claude Fable 5 xHigh 69.9%, and GPT-6 Luna Max 66.6%. No cost-per-task values are printed. The source prints an asterisk; its footnote identifies this Grok 4.7 cell as high effort. Non-Gemini comparison values are provider-reported unless otherwise stated in Google's methodology. Google self-computed this cell with mini-swe-agent, LiteLLM 1.96, and high thinking. Non-Gemini comparison values are provider-reported unless otherwise stated in Google's methodology. Highest-scoring public-leaderboard configuration, using high thinking. The model-card table prints 48.6 while Google's companion launch post prints 49.0; this context retains the model-card value without reconciling them. Non-Gemini comparison values are provider-reported unless otherwise stated in Google's methodology. Exact percentage cells from Google's model card; no effort or harness is assigned beyond the source's benchmark label. 113 tasks; internal bash-only mini-swe-agent fork, no internet, 64GB VMs, pass@1 averaged over five attempts. mini-swe-agent harness run by Datacurve. Official pier framework with mini-swe-agent.
Method / source