RISE
Can the model plan, use tools, and finish a multi-step task?
Agents6 ranked models6 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Claude Opus 4.6Anthropic | 62.5% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
| 2 | Claude Opus 4.5Anthropic | 50.5% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
| 3 | MiniMax M2.5MiniMax | 50.2% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
| 4 | GPT-5.2OpenAI | 50.0% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
| 5 | Gemini 3 ProGoogle | 36.8% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
| 6 | MiniMax M2.1MiniMax | 34.0% | Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Method / source