Kimi Claw 24/7 Bench (Internal)
Can the model plan, use tools, and finish a multi-step task?
Agents4 ranked models4 reported values1 reportsHigher is better
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.5OpenAI | 52.8% | xhigh effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source |
| 2 | Claude Opus 4.8Anthropic | 50.4% | xhigh effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source |
| 3 | Kimi K2.7 CodeMoonshot AI | 46.9% | thinking effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source |
| 4 | Kimi K2.6Moonshot AI | 42.9% | thinking effortMoonshot AI report1 value · 1 reportJul 22, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology
Can the model plan, use tools, and finish a multi-step task?
Agentic execution across tools, environments, or long-running workflows. Higher is better. The value is the percentage reported in this lab's table.
This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.
Kimi Code CLI with thinking enabled, temperature 1.0, top-p 0.95, and 262,144-token context; GPT-5.5 used Codex xhigh and Claude Opus 4.8 used Claude Code xhigh.
Method / source