DreamGen BenchBrowse 296

DreamGen Bench

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents6 ranked models6 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

6
RankModelBest scoreBest reported setting
1Qwen-RobotWorldQwen4.952Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
2LVPExternal baseline4.758Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
3WowExternal baseline4.728Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
4GigaWorldExternal baseline4.216Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
5CosmosNVIDIA4.129Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
6VidarExternal baseline3.341Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

DreamGen source-defined aggregate across six component metrics.

Method / source