PBenchBrowse 296

PBench

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents11 ranked models11 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

11
RankModelBest scoreBest reported setting
1Veo 3Google0.526Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
2KlingExternal baseline0.521Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
3WowExternal baseline0.517Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
4LVPExternal baseline0.515Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
5Wan2.6External baseline0.514Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
6LTX-2External baseline0.506Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
7VidarExternal baseline0.501Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
8GigaWorldExternal baseline0.495Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
9Sora2OpenAI0.487Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
10CosmosNVIDIA0.470Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source
11Qwen-RobotWorldQwen0.455Reported configurationQwen report1 value · 1 reportJun 15, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. This benchmark reports points rather than percent correct.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

PBench source-defined quality, consistency, motion, and domain metrics.

Method / source
PBench · Benchmaxxer