GraySwan ART (pass@1 ASR)Browse 296

GraySwan ART (pass@1 ASR)

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents5 ranked models5 reported values1 reportsLower is better

Top rankings

One row per model, using its best reported score across effort settings.

5
RankModelBest scoreBest reported setting
1Claude Opus 4.8Anthropic0.1%max effortMeta report1 value · 1 reportJul 9, 2026 · source
2Muse Spark 1.1Meta0.3%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
3GPT-5.5OpenAI0.8%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
4Gemini 3.1 ProGoogle1.0%high effortMeta report1 value · 1 reportJul 9, 2026 · source
5Muse Spark 1.0Meta6.1%Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports an attack success rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source