MCP AtlasBrowse 296

MCP Atlas

Can the model discover and coordinate the right tools to complete a multi-step workflow?

Latest stableApril 2026 methodology
Agents40 ranked models61 reported values7 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

40
RankModelBest scoreBest reported setting
1Muse Spark 1.1Meta88.1%Reported configurationMCP Atlas2 values · 2 reportsJul 9, 2026 · source
2Claude Fable 5.1Anthropic87.2%Reported configurationMCP Atlas1 value · 1 reportSep 4, 2026 · source
3Claude Opus 5Anthropic85.8%xhigh effortMCP Atlas1 value · 1 reportAug 3, 2026 · source
4Claude Fable 5Anthropic84.7%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
5Qwen3.8-2.4T-A95BQwen84.5%xhigh effortMCP Atlas1 value · 1 reportSep 17, 2026 · source
6GLM-5.3Z.ai84.2%Reported configurationMCP Atlas1 value · 1 reportSep 17, 2026 · source
7Kimi K3Moonshot AI84.2%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
8Claude Opus 4.8Anthropic83.6%max effortMoonshot AI report4 values · 4 reportsJul 17, 2026 · source
9GPT-5.6 SolOpenAI83.6%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
10Gemini 3.5 FlashGoogle83.6%high effortMCP Atlas2 values · 2 reportsMay 19, 2026 · source
11GPT-5.5OpenAI82.8%xhigh effortMoonshot AI report5 values · 5 reportsJul 17, 2026 · source
12GLM-5.2Z.ai82.6%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model discover and coordinate the right tools to complete a multi-step workflow?

One thousand human-authored tasks spanning 36 real MCP servers and 220 tools. Higher is better. Scale's published MCP Atlas score field is retained in percentage points.

This row set uses the April 2026 methodology revision and its 1,000-task composition. Exact source label, version, effort label, rank, confidence interval, contamination warning, publication flags, and entry timestamp remain material; no model-specific tool setup or cost is inferred.

Official Scale AI MCP Atlas aggregate row; exact source model label "Claude Fable 5"; source version not reported; company anthropic; rank 2; score 83.3 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.25; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-06-09T17:51:31.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-haiku-4-5"; source version not reported; company anthropic; rank 26; score 40.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-opus-4-5 (high)"; source version not reported; company anthropic; rank 13; score 69.8 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.9; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:12.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-opus-4-6 (max)"; source version not reported; company anthropic; rank 3; score 76.8 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.7; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-12-18T18:01:17.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-opus-4-7 (max)"; source version not reported; company anthropic; rank 3; score 79.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.5; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-08T17:02:16.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-opus-4-8 (max)"; source version not reported; company anthropic; rank 2; score 82.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.4; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-05-28T17:37:01.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-opus-5 (xhigh)"; source version not reported; company anthropic; rank 2; score 85.8 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.1; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-08-03T15:52:23.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-sonnet-4-5 (thinking)"; source version not reported; company anthropic; rank 19; score 59.5 percentage points on publisher max score 92.7515; confidenceInterval_upper 3.1; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-12-16T18:09:14.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "claude-sonnet-4-6"; source version not reported; company anthropic; rank 13; score 69.5 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.9; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-06T22:29:16.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Fable 5.1"; source version not reported; company anthropic; rank 1; score 87.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.05; contamination message none reported; isNew=true; new=true; deprecated=false; entry timestamp 2026-09-04T16:47:29.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gemini-3.1-flash-lite (high)"; source version not reported; company google; rank 19; score 57.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-01-08T19:31:06.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gemini-3.1-pro-preview (high)"; source version not reported; company google; rank 3; score 78.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.5; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-24T18:59:59.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Gemini 3.5 Flash (high)"; source version not reported; company google; rank 2; score 83.6 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-05-19T18:56:30.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gemini-3-flash-preview"; source version not reported; company google; rank 18; score 62 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-18T06:16:07.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gemini-3-pro-preview"; source version not reported; company google; rank 13; score 70.3 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.8; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-12-18T18:01:43.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "glm-4p7"; source version not reported; company zai; rank 19; score 58.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "GLM 5.3"; source version not reported; company zai; rank 2; score 84.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.15; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-09-17T14:59:54.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "glm-5p1"; source version not reported; company zai; rank 3; score 75.6 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.7; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-12-17T17:20:04.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "glm-5p2"; source version not reported; company zai; rank 3; score 77.8 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.6; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-20T17:09:57.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.1 (high)"; source version not reported; company openai; rank 24; score 50.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 3.1; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.2 (xhigh)"; source version not reported; company openai; rank 17; score 67.6 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.9; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.4-mini (xhigh)"; source version not reported; company openai; rank 19; score 56.7 percentage points on publisher max score 92.7515; confidenceInterval_upper 3.1; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.4 (xhigh)"; source version not reported; company openai; rank 13; score 70.6 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.8; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-30T18:41:10.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.5 (xhigh)"; source version not reported; company openai; rank 3; score 75.3 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.7; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-28T17:25:29.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "gpt-5.6 (sol)"; source version not reported; company openai; rank 2; score 81.8 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.4; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-15T19:04:43.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Inkling-small"; source version not reported; company thinkingmachines; rank 2; score 79.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.5; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-30T19:17:22.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Inkling (xHigh)"; source version not reported; company thinkingmachines; rank 3; score 76 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.6; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-15T18:19:13.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "kimi-k2p5"; source version not reported; company kimi; rank 18; score 64.4 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "kimi-k3 (max)"; source version not reported; company kimi; rank 2; score 82.3 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.35; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-20T17:00:18.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Muse Spark"; source version not reported; company meta; rank 2; score 82.2 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-08T16:48:43.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Muse Spark 1.1"; source version not reported; company meta; rank 1; score 88.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 1.95; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-07-09T19:46:54.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Nemotron 3 Ultra (thinking)"; source version not reported; company nvidia; rank 18; score 63.1 percentage points on publisher max score 92.7515; confidenceInterval_upper 3; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-09-17T15:00:58.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "o3-pro"; source version not reported; company openai; rank 25; score 44.5 percentage points on publisher max score 92.7515; confidenceInterval_upper 3.1; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2025-09-10T15:38:26.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. Official Scale AI MCP Atlas aggregate row; exact source model label "Qwen3.8-2.4T-A95B (xHigh)"; source version not reported; company alibaba; rank 2; score 84.5 percentage points on publisher max score 92.7515; confidenceInterval_upper 2.25; contamination message none reported; isNew=false; new=false; deprecated=false; entry timestamp 2026-09-17T14:59:32.000Z; methodology revision april-2026-updated-evaluation. The public aggregate feed does not expose a reproducible per-row harness or cost series, so no tool setup or cost is inferred. 1,000 tasks across 36 MCP servers and 220 tools; values are sourced from the report's cited evaluator. Official MCP Atlas codebase; public-set scoring model and setup are described in the source methodology. Kimi Code CLI with thinking enabled, temperature 1.0, top-p 0.95, and 262,144-token context; GPT-5.5 used Codex xhigh and Claude Opus 4.8 used Claude Code xhigh.

Method / source