LiveCodeBenchBrowse 296

LiveCodeBench

Can the model complete substantial programming work under this benchmark's agent setup?

Latest stablerelease_v6 · default window
Coding21 ranked models29 reported values2 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

21
RankModelBest scoreBest reported setting
1o4-miniOpenAI80.2%high effortLiveCodeBench v63 values · 1 reportSep 5, 2026 · source
2o3OpenAI75.8%high effortLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
3Gemini 2.5 ProGoogle73.6%Reported configurationLiveCodeBench v62 values · 1 reportSep 5, 2026 · source
4DeepSeek-R1-0528DeepSeek73.1%Reported configurationLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
5North Mini CodeCohere70.3%Reported configurationCohere report1 value · 1 reportJun 9, 2026 · source
6EXAONE-4.0-32BLG AI Research70.0%Reported configurationLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
7OpenReasoning-Nemotron-32BNVIDIA69.8%Reported configurationLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
8o3 miniOpenAI67.4%high effortLiveCodeBench v63 values · 1 reportSep 5, 2026 · source
9OpenCodeReasoning-Nemotron-1.1-32BNVIDIA66.8%Reported configurationLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
10Grok 3 MiniSpaceXAI66.7%high effortLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
11Qwen3-235B-A22BQwen65.9%Reported configurationLiveCodeBench v61 value · 1 reportSep 5, 2026 · source
12XBai o4-mediumXiaobai65.0%medium effortLiveCodeBench v61 value · 1 reportSep 5, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. Arithmetic mean of the official per-problem pass@1 values in the selected date window.

The selected 454-problem window is the official page's default date range, not the full 1,055-problem release_v6 set. Exact source label, model metadata, release date, contamination flag, difficulty counts, and date window remain material; no cost is inferred.

Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-3.5-Sonnet-20241022"; source model name "claude-3-5-sonnet-20241022"; model style Claude3; release date 2024-03-31; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 36.4, easy 91.2, medium 34.3, hard 8.2; source model URL https://www.anthropic.com/news/claude-3-5-sonnet. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-3-Haiku"; source model name "claude-3-haiku-20240307"; model style Claude3; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 20.2, easy 62.3, medium 12.6, hard 2.8; source model URL https://www.anthropic.com/index/claude-3. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-Opus-4"; source model name "claude-opus-4-20250514_nothink"; model style Claude3Thinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 46.9, easy 94.5, medium 52.5, hard 17.2; source model URL https://www.anthropic.com/claude/sonnet. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-Opus-4 (Thinking)"; source model name "claude-opus-4-20250514"; model style Claude3Thinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 56.6, easy 98.2, medium 70.9, hard 24.1; source model URL https://www.anthropic.com/claude/sonnet. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-Sonnet-4"; source model name "claude-sonnet-4-20250514_nothink"; model style Claude3; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 47.1, easy 96.4, medium 53.9, hard 15.8; source model URL https://www.anthropic.com/claude/sonnet. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Claude-Sonnet-4 (Thinking)"; source model name "claude-sonnet-4-20250514"; model style Claude3Thinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 55.9, easy 97.3, medium 66, hard 26.6; source model URL https://www.anthropic.com/claude/sonnet. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "DeepSeek-R1-0528"; source model name "deepseek-reasoner"; model style DeepSeekR1; release date 2024-06-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 73.1, easy 98.7, medium 85.2, hard 50.7; source model URL https://huggingface.co/deepseek-ai/DeepSeek-R1-0528. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "DeepSeek-V3"; source model name "deepseek-chat"; model style DeepSeekAPI; release date 2024-06-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 27.2, easy 64.3, medium 27.9, hard 6.7; source model URL https://huggingface.co/deepseek-ai/DeepSeek-V3. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "EXAONE-4.0-32B"; source model name "EXAONE-4.0-32B"; model style EXAONE; release date 2024-04-01; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 70, easy 98.4, medium 82.3, hard 46.2; source model URL https://www.wenxiaobai.com/. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Gemini-2.5-Flash-04-17"; source model name "gemini-2.5-flash-preview-04-17"; model style GeminiThinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 60.6, easy 99.1, medium 70.2, hard 33; source model URL https://developers.googleblog.com/en/start-building-with-gemini-25-flash/. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Gemini-2.5-Flash-05-20"; source model name "gemini-2.5-flash-preview-05-20"; model style GeminiThinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 61.9, easy 99.1, medium 71.6, hard 35; source model URL https://developers.googleblog.com/en/start-building-with-gemini-25-flash/. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Gemini-2.5-Pro-05-06"; source model name "gemini-2.5-pro-preview-05-06"; model style GeminiThinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 71.8, easy 98.2, medium 82.3, hard 50.2; source model URL https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#advanced-coding. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Gemini-2.5-Pro-06-05"; source model name "gemini-2.5-pro-preview-06-05"; model style GeminiThinking; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 73.6, easy 99.1, medium 87.2, hard 50.2; source model URL https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#advanced-coding. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "GPT-4-Turbo-2024-04-09"; source model name "gpt-4-turbo-2024-04-09"; model style OpenAIChat; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 28.7, easy 81, medium 24.8, hard 3.1; source model URL https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "GPT-4O-2024-08-06"; source model name "gpt-4o-2024-08-06"; model style OpenAIChat; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 29.5, easy 82.5, medium 26, hard 3.3; source model URL https://openai.com/index/spring-update. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "GPT-4O-mini-2024-07-18"; source model name "gpt-4o-mini-2024-07-18"; model style OpenAIChat; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 27.5, easy 81.9, medium 18.9, hard 3.9; source model URL https://openai.com/index/spring-update. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Grok-3-Mini (High)"; source model name "grok-3-mini-beta_high"; model style Grok; release date 2024-03-01; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 66.7, easy 97.3, medium 74.5, hard 44.8; source model URL https://x.com/i/grok. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O3 (High)"; source model name "o3__high"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 75.8, easy 99.1, medium 84.4, hard 57.1; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O3-Mini-2025-01-31 (High)"; source model name "o3-mini-2025-01-31__high"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 67.4, easy 99.1, medium 84.4, hard 38.4; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O3-Mini-2025-01-31 (Low)"; source model name "o3-mini-2025-01-31__low"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 57, easy 99.1, medium 68.8, hard 26.1; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O3-Mini-2025-01-31 (Med)"; source model name "o3-mini-2025-01-31__medium"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 63, easy 99.1, medium 80.1, hard 31.5; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O4-Mini (High)"; source model name "o4-mini__high"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 80.2, easy 99.1, medium 89.4, hard 63.5; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O4-Mini (Low)"; source model name "o4-mini__low"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 65.9, easy 98.2, medium 80.1, hard 38.4; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "O4-Mini (Medium)"; source model name "o4-mini__medium"; model style OpenAIReason; release date 2023-04-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 74.2, easy 98.2, medium 86.5, hard 52.7; source model URL https://platform.openai.com/docs/api-reference/chat/create#chat-create-reasoning_effort. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "OpenCodeReasoning-Nemotron-1.1-32B"; source model name "nvidia/OpenCodeReasoning-Nemotron-1.1-32B"; model style DeepSeekR1; release date 2024-04-01; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 66.8, easy 97.9, medium 79.6, hard 41.1; source model URL https://huggingface.co/nvidia/OpenCodeReasoning-Nemotron-1.1-32B. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "OpenReasoning-Nemotron-32B"; source model name "nvidia/OpenReasoning-Nemotron-32B"; model style DeepSeekR1; release date 2024-04-01; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 69.8, easy 98.3, medium 81.4, hard 46.3; source model URL https://huggingface.co/nvidia/OpenReasoning-Nemotron-32B. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "Qwen3-235B-A22B"; source model name "Qwen/Qwen3-235B-A22B"; model style CodeQwenInstruct; release date 2024-06-30; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 65.9, easy 99.1, medium 80.1, hard 37.9; source model URL https://huggingface.co/Qwen/Qwen3-235B-A22B. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained. Official LiveCodeBench generation leaderboard aggregate; exact source model label "XBai-o4-medium"; source model name "XBai o4-medium"; model style XBai; release date 2024-04-01; page contamination flag false; selected window 2024-08-01 through 2025-05-01; 454 selected problems with difficulty counts {"easy":110,"medium":141,"hard":203}; pass@1 overall 65, easy 98.2, medium 82.3, hard 35; source model URL https://www.wenxiaobai.com/. The official page computes arithmetic means from per-problem pass@1 values; no per-problem rows, prompts, answers, code, or cost is retained.

Method / source