SciCodeBrowse 296

SciCode

Can the model reason accurately about difficult scientific material?

Latest stablerepository README leaderboard · main problems
Science27 ranked models33 reported values4 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

27
RankModelBest scoreBest reported setting
1Gemini 3.1 ProGoogle59.0%high effortGoogle report1 value · 1 reportFeb 19, 2026 · source
2Gemini 3 ProGoogle56.0%high effortGoogle report2 values · 2 reportsFeb 19, 2026 · source
3Claude Opus 4.6Anthropic52.0%max effortGoogle report2 values · 2 reportsFeb 19, 2026 · source
4GPT-5.2OpenAI52.0%xhigh effortGoogle report2 values · 2 reportsFeb 19, 2026 · source
5Claude Opus 4.5Anthropic50.0%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
6Claude Sonnet 4.6Anthropic47.0%max effortGoogle report1 value · 1 reportFeb 19, 2026 · source
7Claude Sonnet 4.5Anthropic45.0%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
8MiniMax M2.5MiniMax44.4%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
9MiniMax M2.1MiniMax41.0%Reported configurationMiniMax report1 value · 1 reportFeb 12, 2026 · source
10North Mini CodeCohere38.2%Reported configurationCohere report1 value · 1 reportJun 9, 2026 · source
11o3 miniOpenAI10.8%low effortSciCode repository leaderboard3 values · 1 reportOct 7, 2025 · source
12o1OpenAI7.7%Reported configurationSciCode repository leaderboard1 value · 1 reportOct 7, 2025 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model reason accurately about difficult scientific material?

Scientific knowledge and reasoning under the benchmark's published question set. Higher is better. Official main-problem resolve rate over the 80 main problems.

This is the repository table's main-problem measurement. The result changelog date is not stated in the pinned README; keep the repository revision, exact source label, and separate subproblem measurement attached.

Official SciCode repository README aggregate row; exact source model label "Claude3.5-Sonnet"; source table rank not ranked; main-problem resolve rate 4.6% over 80 main problems; subproblem accuracy 26% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Claude3.5-Sonnet (new)"; source table rank not ranked; main-problem resolve rate 4.6% over 80 main problems; subproblem accuracy 25.3% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Claude3-Opus"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 21.5% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Claude3-Sonnet"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 17% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Deepseek-Coder-v2"; source table rank not ranked; main-problem resolve rate 3.1% over 80 main problems; subproblem accuracy 21.2% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Deepseek-R1"; source table rank not ranked; main-problem resolve rate 4.6% over 80 main problems; subproblem accuracy 28.5% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Deepseek-v3"; source table rank not ranked; main-problem resolve rate 3.1% over 80 main problems; subproblem accuracy 23.7% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Gemini 1.5 Pro"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 21.9% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "GPT-4-Turbo"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 22.9% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "GPT-4o"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 25% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Llama-3.1-405B-Chat"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 19.8% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Llama-3.1-70B-Chat"; source table rank not ranked; main-problem resolve rate 0% over 80 main problems; subproblem accuracy 17% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Llama-3-70B-Chat"; source table rank not ranked; main-problem resolve rate 0% over 80 main problems; subproblem accuracy 14.6% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Mixtral-8x22B-Instruct"; source table rank not ranked; main-problem resolve rate 0% over 80 main problems; subproblem accuracy 16.3% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "OpenAI o1-mini"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 22.2% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "OpenAI o1-preview"; source table rank not ranked; main-problem resolve rate 7.7% over 80 main problems; subproblem accuracy 28.5% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "OpenAI o3-mini-high"; source table rank 2; main-problem resolve rate 9.2% over 80 main problems; subproblem accuracy 34.4% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "OpenAI o3-mini-low"; source table rank 1; main-problem resolve rate 10.8% over 80 main problems; subproblem accuracy 33.3% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "OpenAI o3-mini-medium"; source table rank 3; main-problem resolve rate 9.2% over 80 main problems; subproblem accuracy 33% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred. Official SciCode repository README aggregate row; exact source model label "Qwen2-72B-Instruct"; source table rank not ranked; main-problem resolve rate 1.5% over 80 main problems; subproblem accuracy 17% over 338 subproblems; repository revision e3158ea011d4235245a547460d3688d7ccbf9900 dated 2025-10-07; result notice date not stated in the pinned README. No task-level material, prompts, answers, code, test cases, or cost is retained or inferred.

Method / source
SciCode · Benchmaxxer