GDM Situational AwarenessBrowse 296

GDM Situational Awareness

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents5 ranked models5 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

5
RankModelBest scoreBest reported setting
1GPT-5.5OpenAI60.0%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
2Muse Spark 1.1Meta55.1%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
3Gemini 3.1 ProGoogle54.5%high effortMeta report1 value · 1 reportJul 9, 2026 · source
4Claude Opus 4.8Anthropic53.1%max effortMeta report1 value · 1 reportJul 9, 2026 · source
5Muse Spark 1.0Meta29.3%Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source