Terminal-Bench
Can the model plan, use tools, and finish a multi-step task?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Opus 5.5Anthropic | 66.40% | xhigh effortAnthropic report1 value · 1 reportSep 22, 2026 · source |
| 2 | Mythos 5.1Anthropic | 60.90% | Reported configurationAnthropic report1 value · 1 reportSep 1, 2026 · source |
| 3 | GPT-6 AstraOpenAI | 58.18% | max effortTerminal-Bench 4.0 · official leaderboard6 values · 2 reportsSep 3, 2026 · source |
| 4 | Claude Fable 5.1Anthropic | 57.88% | max effortTerminal-Bench 4.0 · official leaderboard7 values · 3 reportsSep 1, 2026 · source |
| 5 | Claude Opus 5Anthropic | 53.94% | xhigh effortTerminal-Bench 4.0 · official leaderboard8 values · 4 reportsJul 24, 2026 · source |
| 6 | Claude Fable 5Anthropic | 44.55% | max effortTerminal-Bench 4.0 · official leaderboard2 values · 2 reportsJun 9, 2026 · source |
| 7 | GLM-5.3Z.ai | 41.82% | max effortTerminal-Bench 4.0 · official leaderboard1 value · 1 reportAug 14, 2026 · source |
| 8 | Grok 4.7SpaceXAI | 38.00% | xhigh effortSpaceXAI report2 values · 2 reportsSep 21, 2026 · source |
| 9 | GPT-5.6 SolOpenAI | 37.30% | Reported configurationAnthropic report4 values · 4 reportsSep 22, 2026 · source |
| 10 | DeepSeek V4.1 FlashDeepSeek | 31.20% | Reported configurationDeepSeek report1 value · 1 reportSep 10, 2026 · source |
| 11 | Claude Opus 4.8Anthropic | 23.64% | max effortTerminal-Bench 4.0 · official leaderboard1 value · 1 reportMay 28, 2026 · source |
| 12 | GPT-5.6 TerraOpenAI | 23.60% | Reported configurationGoogle DeepMind report2 values · 2 reportsSep 2, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
DeepSeek's official V4.1-Flash release table; exact source label and printed percentage retained; no maintainer revision is asserted for this provider cell. Anthropic says Claude Opus 5.5 used xHigh. The Claude Code public leaderboard setup uses five trials per task; Anthropic's reproduced Opus 5 value is 52.3% versus 51.8% on the leaderboard. GPT-6 Astra and GPT-5.6 Sol comparison cells are attributed to OpenAI in Anthropic's footnote. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5.1; agent organization Anthropic; model organization Anthropic; reasoning effort high; source date 2026-09-01; accuracy 54.55% with 95% confidence-interval half-width 3.44 percentage points; row n_trials 330 and metric n_trials 330; successes 180; pass@2=0.647, pass@3=0.697, pass@4=0.7303, pass@5=0.7576; total tokens 2233621484, cached input tokens 2064726577, uncached input tokens 2196192440, output tokens 37429044, and average trial duration 3017.7 seconds; explicitly reported total cost $3985.38; no cost per task is inferred. Source row 4643b23c-53e8-47ae-a3bd-9abc900e9588; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5.1; agent organization Anthropic; model organization Anthropic; reasoning effort low; source date 2026-09-01; accuracy 43.33% with 95% confidence-interval half-width 3.61 percentage points; row n_trials 330 and metric n_trials 330; successes 143; pass@2=0.5455, pass@3=0.6, pass@4=0.6303, pass@5=0.6515; total tokens 1334012516, cached input tokens 1225657226, uncached input tokens 1314077505, output tokens 19935011, and average trial duration 2290.8 seconds; explicitly reported total cost $2358.72; no cost per task is inferred. Source row e18d02ab-0973-4d86-872f-a353c8a7f344; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5.1; agent organization Anthropic; model organization Anthropic; reasoning effort max; source date 2026-09-01; accuracy 57.88% with 95% confidence-interval half-width 3.76 percentage points; row n_trials 330 and metric n_trials 330; successes 191; pass@2=0.7, pass@3=0.747, pass@4=0.7727, pass@5=0.7879; total tokens 2746221560, cached input tokens 2467056906, uncached input tokens 2683099992, output tokens 63121568, and average trial duration 3893.5 seconds; explicitly reported total cost $6243.5; no cost per task is inferred. Source row c741608e-c94e-417d-b7a8-e67111a0c887; createdAt 2026-09-03T01:51:14.437314+00:00; updatedAt 2026-09-03T02:56:08.149542+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5.1; agent organization Anthropic; model organization Anthropic; reasoning effort medium; source date 2026-09-01; accuracy 53.94% with 95% confidence-interval half-width 3.39 percentage points; row n_trials 330 and metric n_trials 330; successes 178; pass@2=0.6379, pass@3=0.6788, pass@4=0.7, pass@5=0.7121; total tokens 1620460859, cached input tokens 1498585169, uncached input tokens 1594469059, output tokens 25991800, and average trial duration 2528.3 seconds; explicitly reported total cost $2832.9; no cost per task is inferred. Source row 83e6bf7e-aeb4-4116-9f30-ca8de96603fb; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5.1; agent organization Anthropic; model organization Anthropic; reasoning effort xhigh; source date 2026-09-01; accuracy 57.88% with 95% confidence-interval half-width 3.36 percentage points; row n_trials 330 and metric n_trials 330; successes 191; pass@2=0.6758, pass@3=0.7152, pass@4=0.7394, pass@5=0.7576; total tokens 2343414000, cached input tokens 2140665130, uncached input tokens 2292987069, output tokens 50426931, and average trial duration 3202.1 seconds; explicitly reported total cost $4872.04; no cost per task is inferred. Source row 9e1c5e11-748f-4f2e-911a-070a396d5bf8; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Fable 5; agent organization Anthropic; model organization Anthropic; reasoning effort max; source date 2026-06-09; accuracy 44.55% with 95% confidence-interval half-width 3.85 percentage points; row n_trials 330 and metric n_trials 330; successes 147; pass@2=0.5727, pass@3=0.6318, pass@4=0.6636, pass@5=0.6818; total tokens 3785220369, cached input tokens 3564065093, uncached input tokens 3726600464, output tokens 58619905, and average trial duration 4202.8 seconds; explicitly reported total cost $7265.01; no cost per task is inferred. Source row 36c077e0-4879-4444-b315-8532d66401d6; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:00.046808+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label GLM-5.3; agent organization Anthropic; model organization Z.ai; reasoning effort max; source date 2026-08-14; accuracy 41.82% with 95% confidence-interval half-width 3.23 percentage points; row n_trials 330 and metric n_trials 330; successes 138; pass@2=0.5076, pass@3=0.547, pass@4=0.5667, pass@5=0.5758; total tokens 8677039373, cached input tokens 8444041600, uncached input tokens 8608379326, output tokens 68660047, and average trial duration 5831 seconds; explicitly reported total cost $2727.63; no cost per task is inferred. Source row d72e8775-f4f1-4313-99e7-35b1cb499f24; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:08:03.780431+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 4.8; agent organization Anthropic; model organization Anthropic; reasoning effort max; source date 2026-05-28; accuracy 23.64% with 95% confidence-interval half-width 3.56 percentage points; row n_trials 330 and metric n_trials 330; successes 78; pass@2=0.3455, pass@3=0.4045, pass@4=0.4424, pass@5=0.4697; total tokens 6415823126, cached input tokens 6136275149, uncached input tokens 6327009817, output tokens 88813309, and average trial duration 5163 seconds; explicitly reported total cost $6481.26; no cost per task is inferred. Source row 952b4217-421f-4d14-849b-9968fcb2063b; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:38.524394+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 5; agent organization Anthropic; model organization Anthropic; reasoning effort high; source date 2026-07-24; accuracy 50.3% with 95% confidence-interval half-width 3.73 percentage points; row n_trials 330 and metric n_trials 330; successes 166; pass@2=0.6227, pass@3=0.6803, pass@4=0.7182, pass@5=0.7424; total tokens 5482033334, cached input tokens 5303103991, uncached input tokens 5434433560, output tokens 47599774, and average trial duration 3822.2 seconds; explicitly reported total cost $4662.27; no cost per task is inferred. Source row 0b2908ee-3035-447c-9876-e68fecb3f413; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 5; agent organization Anthropic; model organization Anthropic; reasoning effort low; source date 2026-07-24; accuracy 34.85% with 95% confidence-interval half-width 3.94 percentage points; row n_trials 330 and metric n_trials 330; successes 115; pass@2=0.4818, pass@3=0.547, pass@4=0.5879, pass@5=0.6212; total tokens 2721178403, cached input tokens 2619174445, uncached input tokens 2697342992, output tokens 23835411, and average trial duration 2947.3 seconds; explicitly reported total cost $2393.88; no cost per task is inferred. Source row 03fe4f10-b683-427b-a820-61e38145d28b; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 5; agent organization Anthropic; model organization Anthropic; reasoning effort max; source date 2026-07-24; accuracy 51.82% with 95% confidence-interval half-width 3.39 percentage points; row n_trials 330 and metric n_trials 330; successes 171; pass@2=0.6167, pass@3=0.6576, pass@4=0.6818, pass@5=0.697; total tokens 6527147637, cached input tokens 6271909956, uncached input tokens 6461135481, output tokens 66012156, and average trial duration 4792.8 seconds; explicitly reported total cost $5969.11; no cost per task is inferred. Source row d71ac3d0-da36-49cb-89d9-323136e77111; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:44.281605+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 5; agent organization Anthropic; model organization Anthropic; reasoning effort medium; source date 2026-07-24; accuracy 44.85% with 95% confidence-interval half-width 3.83 percentage points; row n_trials 330 and metric n_trials 330; successes 148; pass@2=0.5742, pass@3=0.6379, pass@4=0.6727, pass@5=0.697; total tokens 3668060107, cached input tokens 3541133339, uncached input tokens 3634581597, output tokens 33478510, and average trial duration 3292.9 seconds; explicitly reported total cost $3191.59; no cost per task is inferred. Source row 8c102a88-0789-4ce7-b503-bd39e84e157b; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Opus 5; agent organization Anthropic; model organization Anthropic; reasoning effort xhigh; source date 2026-07-24; accuracy 53.94% with 95% confidence-interval half-width 3.17 percentage points; row n_trials 330 and metric n_trials 330; successes 178; pass@2=0.6258, pass@3=0.6621, pass@4=0.6848, pass@5=0.697; total tokens 6897141993, cached input tokens 6631366566, uncached input tokens 6837979419, output tokens 59162574, and average trial duration 4514.7 seconds; explicitly reported total cost $6086.22; no cost per task is inferred. Source row f7844abc-3109-4d5f-91ad-481df7b9fac2; createdAt 2026-09-17T23:52:40.302826+00:00; updatedAt 2026-09-17T23:52:40.302826+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Claude Code; exact model label Sonnet 5; agent organization Anthropic; model organization Anthropic; reasoning effort max; source date 2026-06-30; accuracy 12.42% with 95% confidence-interval half-width 3.06 percentage points; row n_trials 330 and metric n_trials 330; successes 41; pass@2=0.2045, pass@3=0.2621, pass@4=0.3091, pass@5=0.3485; total tokens 21562180742, cached input tokens 21027162464, uncached input tokens 21445581876, output tokens 116598866, and average trial duration 6510.3 seconds; explicitly reported total cost $9603.86; no cost per task is inferred. Source row 8180b9e4-4990-43d0-b906-cd8b43adaaa2; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:50.318045+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-5.6 Luna; agent organization OpenAI; model organization OpenAI; reasoning effort max; source date 2026-06-26; accuracy 17.27% with 95% confidence-interval half-width 2.85 percentage points; row n_trials 330 and metric n_trials 330; successes 57; pass@2=0.2424, pass@3=0.2848, pass@4=0.3121, pass@5=0.3333; total tokens 11557463640, cached input tokens 11305212851, uncached input tokens 11496180968, output tokens 61282672, and average trial duration 4088.3 seconds; explicitly reported total cost $346.67; no cost per task is inferred. Source row 51c6d76e-5baa-48c1-97b3-88656c620eef; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:06.717559+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-5.6 Sol; agent organization OpenAI; model organization OpenAI; reasoning effort max; source date 2026-06-26; accuracy 37.27% with 95% confidence-interval half-width 3.78 percentage points; row n_trials 330 and metric n_trials 330; successes 123; pass@2=0.4955, pass@3=0.55, pass@4=0.5818, pass@5=0.6061; total tokens 4409948969, cached input tokens 4316454466, uncached input tokens 4386497554, output tokens 23451415, and average trial duration 2388.4 seconds; explicitly reported total cost $2541.7; no cost per task is inferred. Source row 0e349bde-b264-494d-a853-fada9c696192; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:13.502808+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-5.6 Terra; agent organization OpenAI; model organization OpenAI; reasoning effort max; source date 2026-06-26; accuracy 21.52% with 95% confidence-interval half-width 3.25 percentage points; row n_trials 330 and metric n_trials 330; successes 71; pass@2=0.3061, pass@3=0.3636, pass@4=0.4061, pass@5=0.4394; total tokens 5683805798, cached input tokens 5560907776, uncached input tokens 5650580965, output tokens 33224833, and average trial duration 2514.6 seconds; explicitly reported total cost $1733.52; no cost per task is inferred. Source row e53da412-5e92-408c-b369-c767924c7c1c; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:20.912429+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-6 Astra; agent organization OpenAI; model organization OpenAI; reasoning effort high; source date 2026-09-03; accuracy 57.88% with 95% confidence-interval half-width 2.97 percentage points; row n_trials 330 and metric n_trials 330; successes 191; pass@2=0.6545, pass@3=0.6864, pass@4=0.703, pass@5=0.7121; total tokens 1208224034, cached input tokens 1151219647, uncached input tokens 1194520091, output tokens 13703943, and average trial duration 2113.7 seconds; explicitly reported total cost $2269.42; no cost per task is inferred. Source row 3475050c-bf5e-4261-a3f6-0af5350af13f; createdAt 2026-09-03T20:53:40.128995+00:00; updatedAt 2026-09-10T21:57:54.867871+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-6 Astra; agent organization OpenAI; model organization OpenAI; reasoning effort low; source date 2026-09-03; accuracy 50.61% with 95% confidence-interval half-width 2.75 percentage points; row n_trials 330 and metric n_trials 330; successes 167; pass@2=0.5712, pass@3=0.6015, pass@4=0.6212, pass@5=0.6364; total tokens 889780710, cached input tokens 851115630, uncached input tokens 881792373, output tokens 7988337, and average trial duration 1683.7 seconds; explicitly reported total cost $1557.3; no cost per task is inferred. Source row b3ad58f3-b311-4d4d-875b-3158cf0309d6; createdAt 2026-09-03T20:53:40.128995+00:00; updatedAt 2026-09-10T21:57:49.894535+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-6 Astra; agent organization OpenAI; model organization OpenAI; reasoning effort max; source date 2026-09-03; accuracy 58.18% with 95% confidence-interval half-width 2.79 percentage points; row n_trials 330 and metric n_trials 330; successes 192; pass@2=0.6485, pass@3=0.6788, pass@4=0.697, pass@5=0.7121; total tokens 1529778322, cached input tokens 1443351702, uncached input tokens 1505789330, output tokens 23988992, and average trial duration 2796.3 seconds; explicitly reported total cost $3267.18; no cost per task is inferred. Source row 5c537be4-7fc3-449b-8bfc-ceb9061c2535; createdAt 2026-09-03T20:53:40.128995+00:00; updatedAt 2026-09-10T21:58:00.001022+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-6 Astra; agent organization OpenAI; model organization OpenAI; reasoning effort medium; source date 2026-09-03; accuracy 54.24% with 95% confidence-interval half-width 2.66 percentage points; row n_trials 330 and metric n_trials 330; successes 179; pass@2=0.603, pass@3=0.6318, pass@4=0.6515, pass@5=0.6667; total tokens 1051885207, cached input tokens 1003874758, uncached input tokens 1041114602, output tokens 10770605, and average trial duration 1885.9 seconds; explicitly reported total cost $1914.8; no cost per task is inferred. Source row f3c3d5a6-6424-4acb-bfcc-615c3f79f6dd; createdAt 2026-09-03T20:53:40.128995+00:00; updatedAt 2026-09-10T21:57:52.562258+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Codex; exact model label GPT-6 Astra; agent organization OpenAI; model organization OpenAI; reasoning effort xhigh; source date 2026-09-03; accuracy 57.88% with 95% confidence-interval half-width 2.72 percentage points; row n_trials 330 and metric n_trials 330; successes 191; pass@2=0.6424, pass@3=0.6697, pass@4=0.6848, pass@5=0.697; total tokens 1201907334, cached input tokens 1140730194, uncached input tokens 1186957047, output tokens 14950287, and average trial duration 2208.5 seconds; explicitly reported total cost $2350.51; no cost per task is inferred. Source row 16db8ad5-84aa-4588-b660-1ce68c0d45e2; createdAt 2026-09-03T20:53:40.128995+00:00; updatedAt 2026-09-10T21:57:57.437008+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Grok Build; exact model label Grok 4.5; agent organization xAI; model organization xAI; reasoning effort high; source date 2026-07-16; accuracy 12.42% with 95% confidence-interval half-width 2.62 percentage points; row n_trials 330 and metric n_trials 330; successes 41; pass@2=0.1833, pass@3=0.2227, pass@4=0.2515, pass@5=0.2727; total tokens 3402536479, cached input tokens 3248613376, uncached input tokens 3370147151, output tokens 32389328, and average trial duration 3162.1 seconds; explicitly reported total cost $2094.11; no cost per task is inferred. Source row 0eed5a0d-96b6-491b-aeb7-cd5e4700c2d7; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:26.760784+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Grok Build; exact model label Grok 4.6; agent organization xAI; model organization xAI; reasoning effort high; source date 2026-08-12; accuracy 20.3% with 95% confidence-interval half-width 3.09 percentage points; row n_trials 330 and metric n_trials 330; successes 67; pass@2=0.2848, pass@3=0.3303, pass@4=0.3636, pass@5=0.3939; total tokens 4001923428, cached input tokens 3877638528, uncached input tokens 3969305207, output tokens 32618221, and average trial duration 2351.5 seconds; explicitly reported total cost $3591.58; no cost per task is inferred. Source row 26354542-edc0-40cd-8f8d-9fe6fbe92ac3; createdAt 2026-08-27T18:30:27.559933+00:00; updatedAt 2026-09-03T00:09:32.707079+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label Grok Build; exact model label Grok 4.7; agent organization xAI; model organization xAI; reasoning effort xhigh; source date 2026-09-21; accuracy 37.58% with 95% confidence-interval half-width 3.54 percentage points; row n_trials 330 and metric n_trials 330; successes 124; pass@2=0.4833, pass@3=0.5394, pass@4=0.5758, pass@5=0.6061; total tokens 5485862389, cached input tokens 5207370880, uncached input tokens 255514442, output tokens 22977067, and average trial duration 5700.9 seconds; explicitly reported total cost $3683.287146; no cost per task is inferred. Source row 84b39f56-fe3c-46c3-919f-3b67b73f9b49; createdAt 2026-09-21T21:28:10.802412+00:00; updatedAt 2026-09-21T21:28:10.802412+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label mini-SWE-agent; exact model label Gemini 3.7 Flash; agent organization SWE-agent; model organization Google; reasoning effort high; source date 2026-08-13; accuracy 11.21% with 95% confidence-interval half-width 2.45 percentage points; row n_trials 330 and metric n_trials 330; successes 37; pass@2=0.1636, pass@3=0.2, pass@4=0.2303, pass@5=0.2576; total tokens 11137189802, cached input tokens 10736925351, uncached input tokens 11085055801, output tokens 52134001, and average trial duration 1688.9 seconds; explicitly reported total cost $1261.87; no cost per task is inferred. Source row 14f4da14-b86a-4dca-92d7-11178d4fe064; createdAt 2026-09-03T01:51:30.749146+00:00; updatedAt 2026-09-03T02:56:10.037161+00:00. Official Terminal-Bench 4.0 leaderboard aggregate row; exact agent label mini-SWE-agent; exact model label Gemini 3.8 Flash; agent organization SWE-agent; model organization Google; reasoning effort high; source date 2026-09-02; accuracy 19.09% with 95% confidence-interval half-width 3.36 percentage points; row n_trials 330 and metric n_trials 330; successes 63; pass@2=0.2879, pass@3=0.353, pass@4=0.4, pass@5=0.4394; total tokens 17189729495, cached input tokens 16692887408, uncached input tokens 17121673011, output tokens 68056484, and average trial duration 1995.1 seconds; explicitly reported total cost $1828.77; no cost per task is inferred. Source row 06850434-507d-4dbd-b74a-09d52519ee35; createdAt 2026-09-03T02:56:12.472078+00:00; updatedAt 2026-09-03T18:03:18.272136+00:00. SpaceXAI's table reports 38.0% for Grok 4.7 at xhigh; maintainer-reported Terminal-Bench rows remain in a separate context. Anthropic reports this Terminal-Bench 4.0 value in its launch table; the table does not print a leaderboard dataset revision or task-level artifacts. Anthropic states that production safeguards were enabled for these comparisons; safeguard interventions can lower scores and intervention cases can be scored as zero. Anthropic reports this Terminal-Bench 4.0 value in its launch table; the table does not print a leaderboard dataset revision or task-level artifacts. General agent capabilities; exact model-card cells, with no additional harness detail inferred.
Method / source