GraySwan ART (pass@1 ASR)
Can the model plan, use tools, and finish a multi-step task?
Agents5 ranked models5 reported values1 reportsLower is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Opus 4.8Anthropic | 0.1% | max effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 2 | Muse Spark 1.1Meta | 0.3% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 3 | GPT-5.5OpenAI | 0.8% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 4 | Gemini 3.1 ProGoogle | 1.0% | high effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 5 | Muse Spark 1.0Meta | 6.1% | Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source