AISE-BenchBrowse 296

AISE-Bench

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents6 ranked models6 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

6
RankModelBest scoreBest reported setting
1Gemini 3 ProGoogle0.5495Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source
2Qwen3-235B-A22BQwen0.4778Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source
3DeepSeek-V3.2DeepSeek0.4729Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source
4GPT-5.2OpenAI0.4562Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source
5GLM-4.7Z.ai0.3510Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source
6Claude-4.5Anthropic0.3328Reported configurationZ.ai report1 value · 1 reportAug 7, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is a ratio on a 0–1 scale reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

CAW model row; source prints GLM-4.7 as 0.3510.

Method / source