gdp.pdf
Can the model extract and reason over information in images or documents?
Top rankings
One row per model, using its best reported score across effort settings.
| Rank | Model | Best score | Best reported setting |
|---|---|---|---|
| 1 | GPT-5.6 SolOpenAI | 40.0% | Reported configurationGoogle DeepMind report2 values · 2 reportsSep 2, 2026 · source |
| 2 | Claude Opus 5Anthropic | 37.0% | Reported configurationGoogle DeepMind report1 value · 1 reportSep 2, 2026 · source |
| 3 | Gemini 3.8 FlashGoogle | 35.0% | Reported configurationGoogle DeepMind report1 value · 1 reportSep 2, 2026 · source |
| 4 | Gemini 3.7 FlashGoogle | 34.0% | Reported configurationGoogle DeepMind report2 values · 2 reportsSep 2, 2026 · source |
| 5 | Claude Fable 5Anthropic | 29.8% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 6 | GPT-5.6 TerraOpenAI | 29.0% | Reported configurationGoogle DeepMind report3 values · 3 reportsSep 2, 2026 · source |
| 7 | Claude Sonnet 5Anthropic | 28.0% | Reported configurationGoogle DeepMind report2 values · 2 reportsSep 2, 2026 · source |
| 8 | GPT-5.5OpenAI | 26.0% | Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source |
| 9 | GPT-5.6 LunaOpenAI | 22.7% | Reported configurationOpenAI report1 value · 1 reportJul 9, 2026 · source |
| 10 | Claude Opus 4.8Anthropic | 22.5% | Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source |
| 11 | Gemini 3.6 FlashGoogle | 22.0% | Reported configurationGoogle report1 value · 1 reportAug 13, 2026 · source |
| 12 | Gemini 3.1 Pro PreviewGoogle | 16.7% | Reported configurationOpenAI report2 values · 2 reportsJul 9, 2026 · source |
Effort curve
Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.
Definition and comparison boundarypublic methodology
Can the model extract and reason over information in images or documents?
Visual perception and multimodal reasoning under the reported protocol. Higher is better. The value is the percentage reported in this lab's table.
Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.
No tools. Gemini and Claude Sonnet 5 cells are Google-computed; GPT-5.6 Terra and Muse Spark 1.2 come from the official public leaderboard. Expert PDF document comprehension; the model card labels this metric All pass rate.
Method / source