MMMU
Can the model combine visual evidence with college-level knowledge and deliberate reasoning?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | Gemini 3.5 FlashGoogle | 83.6% | Reported configurationGoogle report1 value · 1 reportMay 19, 2026 · source |
| 2 | GPT-5.6 SolOpenAI | 83.0% | max effortMoonshot AI report2 values · 2 reportsJul 17, 2026 · source |
| 3 | Kimi K3Moonshot AI | 81.6% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 4 | Claude Fable 5Anthropic | 81.2% | max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source |
| 5 | GPT-5.5OpenAI | 81.2% | xhigh effortMoonshot AI report3 values · 3 reportsJul 17, 2026 · source |
| 6 | Gemini 3 FlashGoogle | 81.2% | Reported configurationGoogle report1 value · 1 reportMay 19, 2026 · source |
| 7 | GPT-5.4OpenAI | 81.2% | Reported configurationMMMU-Pro official feed1 value · 1 reportMar 5, 2026 · source |
| 8 | Gemini 3 ProGoogle | 81.0% | high effortGoogle report1 value · 1 reportFeb 19, 2026 · source |
| 9 | GPT-5.6 TerraOpenAI | 80.7% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 10 | Gemini 3.1 Pro PreviewGoogle | 80.5% | Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source |
| 11 | Gemini 3.1 ProGoogle | 80.5% | Reported configurationGoogle report2 values · 2 reportsMay 19, 2026 · source |
| 12 | Muse SparkMeta | 80.4% | Reported configurationMMMU-Pro official feed1 value · 1 reportApr 8, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
Can the model combine visual evidence with college-level knowledge and deliberate reasoning?
Multidiscipline multimodal questions filtered for genuine visual dependence under a no-tools protocol. Higher is better. MMMU Team's pro.overall percentage under the official zero-shot evaluation.
Tool access, exact source setup label, model publication date, model size, author-provided marker, and the maintainer's selected column are material. Do not mix this row with validation, test, or provider-reported values.
MMMU Team exact source label "Claude Opus 4.6 w/o tools"; pro.overall printed value 73.9; source marker author; publication date 2026-02-05; model type proprietary; size -; publisher provenance https://www.anthropic.com/news/claude-opus-4-6. MMMU Team exact source label "Claude Sonnet 4.6 w/o tools"; pro.overall printed value 74.5; source marker author; publication date 2026-02-17; model type proprietary; size -; publisher provenance https://www.anthropic.com/news/claude-sonnet-4-6. MMMU Team exact source label "Gemini 3.1 Pro Thinking (High)"; pro.overall printed value 80.5; source marker author; publication date 2026-02-19; model type proprietary; size -; publisher provenance https://deepmind.google/models/gemini/pro/. MMMU Team exact source label "Gemma 4 26B A4B"; pro.overall printed value 73.8; source marker author; publication date 2026-04-02; model type open_source; size 26B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 31B"; pro.overall printed value 76.9; source marker author; publication date 2026-04-02; model type open_source; size 31B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 E2B"; pro.overall printed value 44.2; source marker author; publication date 2026-04-02; model type open_source; size 2B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "Gemma 4 E4B"; pro.overall printed value 52.6; source marker author; publication date 2026-04-02; model type open_source; size 4B; publisher provenance https://deepmind.google/models/gemma/gemma-4/. MMMU Team exact source label "GPT-5.1 Thinking"; pro.overall printed value 79.0; source marker author; publication date 2025-11-13; model type proprietary; size -; publisher provenance https://openai.com/index/gpt-5-1-for-developers/. MMMU Team exact source label "GPT-5.2 Thinking w/o Python"; pro.overall printed value 80.4; source marker author; publication date 2025-12-11; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-2/. MMMU Team exact source label "GPT-5.2 Thinking w/o tools"; pro.overall printed value 79.5; source marker author; publication date 2025-12-11; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-2/. MMMU Team exact source label "GPT-5.4 Thinking w/o tools"; pro.overall printed value 81.2; source marker author; publication date 2026-03-05; model type proprietary; size -; publisher provenance https://openai.com/index/introducing-gpt-5-4/. MMMU Team exact source label "Muse Spark Thinking"; pro.overall printed value 80.4; source marker author; publication date 2026-04-08; model type proprietary; size -; publisher provenance https://ai.meta.com/blog/introducing-muse-spark-msl/. No tools.
Method / source