AgentHarm Verified (benign) (FRR)Browse 296

AgentHarm Verified (benign) (FRR)

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents4 ranked models4 reported values1 reportsLower is better

Top rankings

One row per model, using its best reported score across effort settings.

4
RankModelBest scoreBest reported setting
1Gemini 3.1 ProGoogle5.7%high effortMeta report1 value · 1 reportJul 9, 2026 · source
2Muse Spark 1.1Meta16.5%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
3GPT-5.5OpenAI21.6%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
4Claude Opus 4.8Anthropic37.5%max effortMeta report1 value · 1 reportJul 9, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Lower is better. The source reports a false-refusal rate as a percentage.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

Method / source