GDM-Stealth (of 4)Browse 296

GDM-Stealth (of 4)

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents5 ranked models5 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

5
RankModelBest scoreBest reported setting
1Gemini 3.1 ProGoogle3high effortMeta report1 value · 1 reportJul 9, 2026 · source
2Muse Spark 1.1Meta2xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
3GPT-5.5OpenAI1xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
4Muse Spark 1.0Meta1Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source
5Claude Opus 4.8Anthropic0max effortMeta report1 value · 1 reportJul 9, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarylimited methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The source prints a numerator out of four scenarios; the denominator is preserved in the benchmark name and scale.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source