MMMUBrowse 296

MMMU

Can the model combine visual evidence with college-level knowledge and deliberate reasoning?

Latest stablePro · no tools
Vision23 ranked models33 reported values5 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

23
RankModelBest scoreBest reported setting
1Gemini 3.5 FlashGoogle83.6%Reported configurationGoogle report1 value · 1 reportMay 19, 2026 · source
2GPT-5.6 SolOpenAI83.0%max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source
3Kimi K3Moonshot AI81.6%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
4Claude Fable 5Anthropic81.2%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
5GPT-5.5OpenAI81.2%xhigh effortMoonshot AI report3 values · 3 reportsJul 17, 2026 · source
6Gemini 3 FlashGoogle81.2%Reported configurationGoogle report1 value · 1 reportMay 19, 2026 · source
7GPT-5.4OpenAI81.2%Reported configurationMMMU-Pro official feed1 value · 1 reportMar 5, 2026 · source
8Gemini 3 ProGoogle81.0%high effortGoogle report1 value · 1 reportFeb 19, 2026 · source
9GPT-5.6 TerraOpenAI80.7%Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source
10Gemini 3.1 Pro PreviewGoogle80.5%Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source
11Gemini 3.1 ProGoogle80.5%Reported configurationGoogle report2 values · 2 reportsMay 19, 2026 · source
12Muse SparkMeta80.4%Reported configurationMMMU-Pro official feed1 value · 1 reportApr 8, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model combine visual evidence with college-level knowledge and deliberate reasoning?

Multidiscipline multimodal questions filtered for genuine visual dependence under a no-tools protocol. Higher is better. MMMU Team's pro.overall percentage under the official zero-shot evaluation.

Tool access, exact source setup label, model publication date, model size, author-provided marker, and the maintainer's selected column are material. Do not mix this row with validation, test, or provider-reported values.

MMMU Team exact source label "Claude Opus 4.6 w/o tools"; pro.overall printed value 73.9; source marker author; publication date 2026-02-05; model type proprietary; size -; publisher provenance https://www.anthropic.com/news/claude-opus-4-6. MMMU Team exact source label "Claude Sonnet 4.6 w/o tools"; pro.overall printed value 74.5; source marker author; publication date 2026-02-17; model type proprietary; size -; publisher provenance https://www.anthropic.com/news/claude-sonnet-4-6. MMMU Team exact source label "Gemini 3.1 Pro Thinking (High)"; pro.overall printed value 80.5; source marker author; publication date 2026-02-19; model type proprietary; size -; publisher provenance https://deepmind.google/models/gemini/pro/. MMMU Team exact source label "Gemma 4 26B A4B"; pro.overall printed value 73.8; source marker author; publication date 2026-04-02; model type open_source; size 26B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 31B"; pro.overall printed value 76.9; source marker author; publication date 2026-04-02; model type open_source; size 31B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 E2B"; pro.overall printed value 44.2; source marker author; publication date 2026-04-02; model type open_source; size 2B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 E4B"; pro.overall printed value 52.6; source marker author; publication date 2026-04-02; model type open_source; size 4B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "GPT-5.1 Thinking"; pro.overall printed value 79.0; source marker author; publication date 2025-11-13; model type proprietary; size -; publisher provenance https://openai.com/index/gpt-5-1-for-developers/. MMMU Team exact source label "GPT-5.2 Thinking w/o Python"; pro.overall printed value 80.4; source marker author; publication date 2025-12-11; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-2/. MMMU Team exact source label "GPT-5.2 Thinking w/o tools"; pro.overall printed value 79.5; source marker author; publication date 2025-12-11; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-2/. MMMU Team exact source label "GPT-5.4 Thinking w/o tools"; pro.overall printed value 81.2; source marker author; publication date 2026-03-05; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-4/. MMMU Team exact source label "Muse Spark Thinking"; pro.overall printed value 80.4; source marker author; publication date 2026-04-08; model type proprietary; size -; publisher provenance https://ai.meta.com/blog/introducing-muse-spark-msl/. No tools.

Method / source