RISEBrowse 296

RISE

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents6 ranked models6 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

6
RankModelBest scoreBest reported setting
1Claude Opus 4.6Anthropic62.5%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
2Claude Opus 4.5Anthropic50.5%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
3MiniMax M2.5MiniMax50.2%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
4GPT-5.2OpenAI50.0%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
5Gemini 3 ProGoogle36.8%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
6MiniMax M2.1MiniMax34.0%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Method / source