GDM Situational Awareness
Can the model plan, use tools, and finish a multi-step task?
Agents5 ranked models5 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.5OpenAI | 60.0% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 2 | Muse Spark 1.1Meta | 55.1% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 3 | Gemini 3.1 ProGoogle | 54.5% | high effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 4 | Claude Opus 4.8Anthropic | 53.1% | max effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 5 | Muse Spark 1.0Meta | 29.3% | Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source