Kimi Claw 24/7 Bench (Internal)Browse 296

Kimi Claw 24/7 Bench (Internal)

Can the model plan, use tools, and finish a multi-step task?

Version not specifiedExact reported variant
Agents4 ranked models4 reported values1 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

4
RankModelBest scoreBest reported setting
1GPT-5.5OpenAI52.8%xhigh effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source
2Claude Opus 4.8Anthropic50.4%xhigh effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source
3Kimi K2.7 CodeMoonshot AI46.9%thinking effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source
4Kimi K2.6Moonshot AI42.9%thinking effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model plan, use tools, and finish a multi-step task?

Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Kimi Code CLI with thinking enabled, temperature 1.0, top-p 0.95, and 262,144-token context; GPT-5.5 used Codex xhigh and Claude Opus 4.8 used Claude Code xhigh.

Method / source