Terminal-Bench Science
How often an AI agent completes a scientific research workflow using terminal tools.
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 68.1% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 22, 2026 · source |
| 2 | Claude Fable 5.1Anthropic | 40.0% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 3 | Claude Opus 5Anthropic | 30.0% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 4 | GPT-5.6 SolOpenAI | 22.4% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 5 | Claude Fable 5Anthropic | 21.4% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 6 | DeepSeek V4.1 FlashDeepSeek | 15.7% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 7 | Grok 4.7SpaceXAI | 14.3% | xhigh effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 21, 2026 · source |
| 8 | Gemini 3.8 FlashGoogle | 12.4% | high effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 9 | Claude Opus 4.8Anthropic | 10.5% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 10 | GPT-5.6 TerraOpenAI | 8.6% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 11 | GLM-5.3Z.ai | 8.1% | max effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
| 12 | Grok 4.6SpaceXAI | 7.1% | xhigh effortTerminal-Bench-Science 0.1 · official leaderboard1 value · 1 reportSep 16, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
How often an AI agent completes a scientific research workflow using terminal tools.
Resolution rate across 70 research workflows in life, physical, earth, mathematical and engineering sciences, with three trials per task. Higher is better. The maintainer reports the percentage of scored trials resolved successfully.
This is a small, early research-workflow suite, not a general measure of scientific discovery. Model and agent setups differ. Retain source standard errors and exact dataset revision; never mix these official results with provider reports or other benchmark versions.
Official Terminal-Bench-Science 0.1; exact model DeepSeek V4 Pro; agent Codex (OpenAI); model organization DeepSeek; reasoning effort max; model release date 2026-08-13; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=8; resolution rate 3.8095238095238093%; standard error 1.3209662936780717 percentage points (not a 95% confidence interval); total tokens 11201441597; reported total cost $1120.15076932; no cost per task inferred. Source rank 15; row 00b5eabf-caca-435c-b8a8-458633192aa8; created 2026-09-05T19:44:39.989994+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Fable 5.1; agent Claude Code (Anthropic); model organization Anthropic; reasoning effort max; model release date 2026-09-01; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=84; resolution rate 40%; standard error 3.380617018914066 percentage points (not a 95% confidence interval); total tokens 3753292530; reported total cost $6247.1398385; no cost per task inferred. Source rank 2; row 1264a994-632f-4742-a143-a848a6ad4114; created 2026-09-13T19:56:07.340022+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Grok 4.6; agent Grok Build (xAI); model organization xAI; reasoning effort xhigh; model release date 2026-08-12; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=15; resolution rate 7.142857142857142%; standard error 1.7771905411718545 percentage points (not a 95% confidence interval); total tokens 3709385002; reported total cost $3344.187332; no cost per task inferred. Source rank 13; row 231bc6c3-8f2b-455e-8c87-5a08ec200b05; created 2026-08-24T10:25:39.579396+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Grok 4.7; agent Grok Build (xAI); model organization xAI; reasoning effort xhigh; model release date 2026-09-21; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=30; resolution rate 14.285714285714286%; standard error 2.414726442081476 percentage points (not a 95% confidence interval); total tokens 4206226774; reported total cost $3070.816824; no cost per task inferred. Source rank 7; row 360317df-baea-4ead-8a11-09db70a27d15; created 2026-09-21T20:47:03.503383+00:00; updated 2026-09-21T20:47:10.285461+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model GPT-5.6 Sol; agent Codex (OpenAI); model organization OpenAI; reasoning effort max; model release date 2026-07-09; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=47; resolution rate 22.380952380952383%; standard error 2.876164947101791 percentage points (not a 95% confidence interval); total tokens 8406870193; reported total cost $4221.8683028; no cost per task inferred. Source rank 4; row 378f8b49-dbd8-4acf-a065-8c07e1712de2; created 2026-08-23T21:46:16.093869+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Gemini 3.8 Flash; agent mini-SWE-agent (SWE-agent); model organization Google DeepMind; reasoning effort high; model release date 2026-09-02; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=26; resolution rate 12.380952380952381%; standard error 2.2728283787427124 percentage points (not a 95% confidence interval); total tokens 10499649912; reported total cost $1116.97645035; no cost per task inferred. Source rank 8; row 409f68a6-d6a4-43c9-adb8-63679dc17a45; created 2026-09-03T08:32:35.782397+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model DeepSeek V4.1 Flash; agent Codex (OpenAI); model organization DeepSeek; reasoning effort max; model release date 2026-09-10; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=33; resolution rate 15.714285714285714%; standard error 2.511392893650442 percentage points (not a 95% confidence interval); total tokens 14859405919; reported total cost $385.933474254; no cost per task inferred. Source rank 6; row 58d4c48d-791b-4873-ba3e-af1cf997a79c; created 2026-09-11T18:11:57.008377+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Gemini 3.7 Flash; agent mini-SWE-agent (SWE-agent); model organization Google DeepMind; reasoning effort high; model release date 2026-08-13; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=12; resolution rate 5.714285714285714%; standard error 1.6017483159468233 percentage points (not a 95% confidence interval); total tokens 12023077591; reported total cost $1296.5871279; no cost per task inferred. Source rank 14; row 5a71d088-1bfd-4bbe-a65a-973a4fffd50f; created 2026-09-03T08:32:35.782397+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Kimi K3; agent Claude Code (Anthropic); model organization Moonshot AI; reasoning effort max; model release date 2026-07-16; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=15; resolution rate 7.142857142857142%; standard error 1.7771905411718545 percentage points (not a 95% confidence interval); total tokens 3191602969; reported total cost $1523.7995976; no cost per task inferred. Source rank 12; row 66e64f64-abbd-40a6-a470-2151df25472e; created 2026-08-25T19:57:08.310275+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model GPT-6 Astra; agent Codex (OpenAI); model organization OpenAI; reasoning effort max; model release date 2026-09-03; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=143; resolution rate 68.0952380952381%; standard error 3.2164475807033743 percentage points (not a 95% confidence interval); total tokens 2288326796; reported total cost $4997.461181; no cost per task inferred. Source rank 1; row 76867ebc-8f35-43c5-a8b0-0cee9db46c96; created 2026-09-22T11:45:58.377345+00:00; updated 2026-09-22T11:46:04.426644+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model GLM 5.3; agent Claude Code (Anthropic); model organization Z.AI; reasoning effort max; model release date 2026-08-14; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=17; resolution rate 8.095238095238095%; standard error 1.8822364227102866 percentage points (not a 95% confidence interval); total tokens 8486097200; reported total cost $2733.16185912; no cost per task inferred. Source rank 11; row 78f18248-3ee9-4eb4-b7af-19019f1fc64f; created 2026-08-26T07:35:19.166068+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model GPT-5.6 Luna; agent Codex (OpenAI); model organization OpenAI; reasoning effort max; model release date 2026-07-09; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=7; resolution rate 3.3333333333333335%; standard error 1.2387055882620108 percentage points (not a 95% confidence interval); total tokens 14246922352; reported total cost $383.19963959999995; no cost per task inferred. Source rank 16; row 7cf10b06-3e5b-4723-b98f-c679e773959b; created 2026-08-25T19:25:57.794681+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Opus 5; agent Claude Code (Anthropic); model organization Anthropic; reasoning effort max; model release date 2026-07-24; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=63; resolution rate 30%; standard error 3.162277660168379 percentage points (not a 95% confidence interval); total tokens 7266668313; reported total cost $6992.67732375; no cost per task inferred. Source rank 3; row bcf50bd8-8555-48a6-8cae-9f697af54051; created 2026-08-24T07:50:37.026092+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model GPT-5.6 Terra; agent Codex (OpenAI); model organization OpenAI; reasoning effort max; model release date 2026-07-09; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=18; resolution rate 8.571428571428571%; standard error 1.9317811536651808 percentage points (not a 95% confidence interval); total tokens 7618801416; reported total cost $1983.9189368999998; no cost per task inferred. Source rank 10; row c79d5ef6-b3c6-40da-84b5-946c11099426; created 2026-08-25T08:52:26.313446+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Fable 5; agent Claude Code (Anthropic); model organization Anthropic; reasoning effort max; model release date 2026-06-09; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=45; resolution rate 21.428571428571427%; standard error 2.831517739900328 percentage points (not a 95% confidence interval); total tokens 6363562181; reported total cost $14175.948185250001; no cost per task inferred. Source rank 5; row e3f2861a-97db-4730-a0ff-99f08d8c7299; created 2026-08-24T14:02:29.713688+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate. Official Terminal-Bench-Science 0.1; exact model Opus 4.8; agent Claude Code (Anthropic); model organization Anthropic; reasoning effort max; model release date 2026-05-28; 70 unique tasks, 3 trials per task; source tasks=210, n_trials=210, passes=22; resolution rate 10.476190476190476%; standard error 2.1133008267654967 percentage points (not a 95% confidence interval); total tokens 6770231957; reported total cost $5837.716905; no cost per task inferred. Source rank 9; row e8bf7739-3c94-46a3-95a5-fa4d4444034e; created 2026-08-25T19:25:57.794681+00:00; updated 2026-09-16T13:45:27.369938+00:00; dataset terminal-bench-science/terminal-bench-science/2b817f26-dc4f-4477-8032-2218dcc553b5. Agent setups differ; provider-reported results remain separate.
Method / source