AISE-Bench
Can the model plan, use tools, and finish a multi-step task?
Agents6 ranked models6 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Gemini 3 ProGoogle | 0.5495 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
| 2 | Qwen3-235B-A22BQwen | 0.4778 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
| 3 | DeepSeek-V3.2DeepSeek | 0.4729 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
| 4 | GPT-5.2OpenAI | 0.4562 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
| 5 | GLM-4.7Z.ai | 0.3510 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
| 6 | Claude-4.5Anthropic | 0.3328 | Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
CAW model row; source prints GLM-4.7 as 0.3510.
Method / source