Follow the evidence for the work. Start with a capability, inspect every model on that benchmark, then turn the view around and inspect the model.
Reason about science Scientific reasoning Solve expert problems Hard reasoning Fix real software bugs Repository fixes Build with an agent Coding agents Operate tools and systems Terminal work Reason across long inputs Long documents Survey real workflows Professional work
Models by benchmark Benchmarks by model Capability matrix Effort curves
Benchmark Agents' Last Exam Management Consulting Tasks (Internal) Big Finance Bench SWE-Bench Pro DeepSWE v1.1 Terminal-Bench 2.1 GeneBench Pro LifeSciBench MedChemBench (Internal) HealthBench Professional OSWorld 2.0 BrowseComp BenchCAD BenchCAD (python tool) Capture-the-Flag Challenges SEC-Bench Pro ExploitBench ExploitGym Internal Research Debugging Evaluation KernelGen 1P NanoGPT PostTrainBench Lite RSI Index MMMU Pro (no tools) MMMU Pro (with tools) gdp.pdf GPQA Diamond FrontierMath Tier 1-3 (v2) FrontierMath Tier 4 (v2) AutomationBench Toolathlon OpenAI MRCR v2 · 8-needle · 256K-512K OpenAI MRCR v2 · 8-needle · 512K-1M GraphWalks BFS · 256K F1 GraphWalks BFS · 1M F1 ARC-AGI-3 DeepSWE Program Bench FrontierSWE SWE Marathon PostTrain Bench MLS Bench Kimi Code Bench 2.0 (Internal) DeepSearchQA · F1 Toolathlon-Verified MCP Atlas Job Bench Office QA Pro SpreadsheetBench 2 DECK-Bench (Internal) HLE-Full HLE-Full with tools MMMU-Pro with python CharXiv · RQ CharXiv · RQ with python MathVision MathVision with python BabyVision with python ZeroBench main · pass@5 ZeroBench main with python · pass@5 WorldVQA ForceAnswer OmniDocBench PerceptionBench FrontierCode · Diamond Blueprint-Bench 2 OSWorld-Verified Legal Agent Benchmark BioMysteryBench · hard BioMysteryBench · human solved Humanity's Last Exam · search + code ARC-AGI-2 · verified Terminal-Bench 2.0 Terminal-Bench 2.0 · Codex harness SWE-Bench Verified LiveCodeBench Pro · Elo SciCode τ²-bench · Retail τ²-bench · Telecom MMMLU MRCR v2 · 8-needle · 128K average MRCR v2 · 8-needle · 1M pointwise Finance Agent v2 CharXiv Reasoning (no tools) MLE-Bench CharXiv Reasoning (with tools) MRCR Long Context (1M context window) OSWorld 2.0 · strict binary completion OSWorld 2.0 · partial score WebArena-Verified HealthBench Pro BabyVision (with tools) MBCT VCT HPCT WMDP-Bio WMDP-Chem ProtocolQA SeqQA (agentic) ABC Bench (Fragment Design) ABC Bench (Liquid Handling) ABC Bench (Screening Evasion) BioDesign Tools (avg) Cybench (pass@1) Curated CTFs (pass@1) CyberGym (pass@1) ExploitGym (pass@1) CyScenarioBench (pass@1) Social Engineering AIRS-Bench SHADE-Arena GDM-Stealth (of 4) GDM Situational Awareness Refusals: BioTIER Refusals: Chemical Agents Cyber Misuse Chat (ASR) Catastrophic Cyber Misuse (ASR) Poly-Guard Bench (ASR) MASK Agentic Misalignment Jailbreak StrongREJECT v2 (ASR) FORTRESS (ARS) AgentHarm (ASR) AgentDojo (pass@1 ASR) GraySwan ART (pass@1 ASR) OR-Bench (FRR) Cyber Misuse Chat (FRR) AgentHarm Verified (benign) (FRR) SAVE-Bench Internal Sycophancy DeceptionBench HLE Calibration CursorBench-3 SWE-bench Multilingual Terminal-Bench · Cursor report CritPt AIME 2026 HMMT · November 2025 HMMT · February 2026 IMOAnswerBench NL2Repo Terminal-Bench 2.1 · Terminus-2 Terminal-Bench 2.1 · best reported harness MCP-Atlas · public set Tool-Decathlon SWE-fficiency KernelBench Hard PostTrainBench · normalized score Multi-SWE-Bench VIBE-Pro · average BrowseComp · context management Wide Search RISE BFCL multi-turn AIME 2025 IFBench Humanity's Last Exam · full set Kimi Code Bench v2 (Internal) MLS Bench Lite Kimi Claw 24/7 Bench (Internal) MCP Mark Verified Terminal-Bench Hard LiveCodeBench v6 AstaBench · overall HMMT 2025 · pass@1 Codeforces · rating τ²-bench · pass@1 Sort models Highest median first Lowest median first Model name Reported values 43
agents % success
“Can a model-driven agent complete realistic work inside a terminal?”
Measures Multi-step command-line, coding, configuration, debugging, and systems tasks.
Use it for agents tool use workflow automation
Boundary Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
Report setup Kimi Code harness. Codex CLI. Gemini CLI. Terminus-2 harness. 89 tasks, bash-only agent harness, 6 CPU cores, 8GB RAM, pass@1 averaged over five attempts. 8C16G sandbox, two-hour timeout, 128K max output tokens, Terminus 2 scaffolding. Evidence mode All reported values 9 first-party reporting contexts are plotted together. A report label stays on every value; Benchmaxxer does not combine them into a synthetic score.
Method / reporting source ↗