OR-Bench (FRR)
Can the model plan, use tools, and finish a multi-step task?
Agents5 ranked models5 reported values1 reportsLower is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Gemini 3.1 ProGoogle | 2.5% | high effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 2 | Claude Opus 4.8Anthropic | 3.3% | max effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 3 | Muse Spark 1.1Meta | 4.8% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 4 | GPT-5.5OpenAI | 6.7% | xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source |
| 5 | Muse Spark 1.0Meta | 8.0% | Reported configurationMeta report1 value · 1 reportJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Method / source