GDM-Stealth (of 4)
Can the model plan, use tools, and finish a multi-step task?
Agents5 ranked models5 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Gemini 3.1 ProGoogle | 3 | high effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 2 | Muse Spark 1.1Meta | 2 | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 3 | GPT-5.5OpenAI | 1 | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 4 | Muse Spark 1.0Meta | 1 | Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source |
| 5 | Claude Opus 4.8Anthropic | 0 | max effortMeta report1 value · 1 reportJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundarylimited methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source