Humanity’s Last ExamBrowse 296

Humanity’s Last Exam

Can the model answer deliberately difficult expert questions across academic fields?

Latest stableFinal 2,500
Reasoning63 ranked models102 reported values11 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

63
RankModelBest scoreBest reported setting
1Claude Mythos PreviewAnthropic56.80%Reported configurationAnthropic report1 value · 1 reportJun 9, 2026 · source
2GPT-6 AstraOpenAI54.80%Reported configurationHLE · full1 value · 1 reportSep 9, 2026 · source
3Claude Fable 5Anthropic53.30%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source
4Muse Spark 1.1Meta52.20%xhigh effortMeta report1 value · 1 reportJul 9, 2026 · source
5Claude Opus 4.8Anthropic49.80%max effortMoonshot AI report5 values · 5 reportsJul 17, 2026 · source
6Claude Opus 4.7Anthropic46.90%Reported configurationGoogle report2 values · 2 reportsMay 19, 2026 · source
7Claude Fable 5.1Anthropic46.50%xhigh effortHLE · full1 value · 1 reportSep 3, 2026 · source
8Gemini 3.1 Pro PreviewGoogle46.44%high effortHLE · full2 values · 2 reportsApr 10, 2026 · source
9Gemini 3.1 ProGoogle45.40%high effortMeta report4 values · 4 reportsJul 9, 2026 · source
10GPT-5.5OpenAI44.80%xhigh effortMeta report5 values · 5 reportsJul 9, 2026 · source
11Gemini 3.8 FlashGoogle44.52%Reported configurationHLE · full1 value · 1 reportSep 9, 2026 · source
12GPT-5.6 SolOpenAI44.50%max effortMoonshot AI report1 value · 1 reportJul 17, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundarypublic methodology

Can the model answer deliberately difficult expert questions across academic fields?

Broad expert-level reasoning and knowledge without external tools. Higher is better. The value is the percentage reported in this lab's table.

Harness, tools, prompts, attempt count, and benchmark version can materially change the result. Compare within one reporting context.

DeepSeek prints HLE 36.8 (39.1*); this row retains the main printed HLE value and the separate starred pure-text value is stored in hle-text-only. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude 3.5 Sonnet (October 2024)"; source version not reported; company anthropic; rank 43; score 4.08 on publisher max score 58.4422; confidenceInterval_upper 0.78; calibrationError 84; contamination message This model was used as an initial filter for the dataset.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude 3.7 Sonnet (Thinking)"; source version not reported; company anthropic; rank 31; score 8.04 on publisher max score 58.4422; confidenceInterval_upper 1.07; calibrationError 80; contamination message Thinking budget: 16,000 tokens. Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude Opus 4 "; source version not reported; company anthropic; rank 31; score 6.68 on publisher max score 58.4422; confidenceInterval_upper 0.98; calibrationError 74; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-23T15:37:38.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-1-20250805"; source version not reported; company anthropic; rank 31; score 7.92 on publisher max score 58.4422; confidenceInterval_upper 1.06; calibrationError 70; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-08T17:18:51.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-1-20250805-thinking"; source version not reported; company anthropic; rank 27; score 11.52 on publisher max score 58.4422; confidenceInterval_upper 1.25; calibrationError 71; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-08T17:18:22.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-5-20251101"; source version not reported; company anthropic; rank 23; score 14.16 on publisher max score 58.4422; confidenceInterval_upper 1.37; calibrationError 56; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-26T20:25:03.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-5-20251101-thinking"; source version not reported; company anthropic; rank 10; score 25.2 on publisher max score 58.4422; confidenceInterval_upper 1.7; calibrationError 55; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-26T20:23:24.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-6 (Non-Thinking)"; source version not reported; company anthropic; rank 16; score 19 on publisher max score 58.4422; confidenceInterval_upper 1.54; calibrationError 44; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-02-17T17:05:29.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-6-thinking-max"; source version not reported; company anthropic; rank 6; score 34.44 on publisher max score 58.4422; confidenceInterval_upper 1.86; calibrationError 46; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-02-17T17:04:32.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-opus-4-7"; source version not reported; company anthropic; rank 6; score 36.2 on publisher max score 58.4422; confidenceInterval_upper 1.88; calibrationError 47; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-22T20:13:01.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude Opus 4 (Thinking)"; source version not reported; company anthropic; rank 28; score 10.72 on publisher max score 58.4422; confidenceInterval_upper 1.21; calibrationError 73; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-24T06:35:10.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude Sonnet 4"; source version not reported; company anthropic; rank 42; score 5.52 on publisher max score 58.4422; confidenceInterval_upper 0.9; calibrationError 76; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-23T15:37:26.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-sonnet-4-5-20250929"; source version not reported; company anthropic; rank 31; score 7.52 on publisher max score 58.4422; confidenceInterval_upper 1.03; calibrationError 70; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-10-02T17:28:52.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "claude-sonnet-4-5-20250929-thinking"; source version not reported; company anthropic; rank 23; score 13.72 on publisher max score 58.4422; confidenceInterval_upper 1.35; calibrationError 65; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-10-02T17:27:45.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Claude Sonnet 4 (Thinking)\n"; source version not reported; company anthropic; rank 31; score 7.76 on publisher max score 58.4422; confidenceInterval_upper 1.05; calibrationError 75; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-23T15:37:20.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Fable 5.1 (xhigh)\n"; source version not reported; company anthropic; rank 2; score 46.5 on publisher max score 58.4422; confidenceInterval_upper 2; calibrationError 20; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=true; new=true; deprecated=false; entry timestamp 2026-09-03T18:22:21.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini-1.5-Pro-002"; source version not reported; company google; rank 43; score 4.6 on publisher max score 58.4422; confidenceInterval_upper 0.82; calibrationError 88; contamination message This model was used as an initial filter for the dataset.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 2.0 Flash Thinking (January 2025)"; source version not reported; company google; rank 41; score 6.56 on publisher max score 58.4422; confidenceInterval_upper 0.97; calibrationError 82; contamination message Sampled at temperature 0.7; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 2.5 Flash (April 2025)"; source version not reported; company google; rank 23; score 12.08 on publisher max score 58.4422; confidenceInterval_upper 1.28; calibrationError 80; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-17T19:55:59.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 2.5 Flash Preview (May 2025) "; source version not reported; company google; rank 28; score 10.96 on publisher max score 58.4422; confidenceInterval_upper 1.22; calibrationError 82; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-20T18:29:27.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 2.5 Pro Experimental (March 2025)"; source version not reported; company google; rank 20; score 18.16 on publisher max score 58.4422; confidenceInterval_upper 1.51; calibrationError 71; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at temperature = 1.0, top_p = 0.95.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:50.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gemini-2.5-pro-preview-06-05"; source version not reported; company google; rank 15; score 21.64 on publisher max score 58.4422; confidenceInterval_upper 1.61; calibrationError 72; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-06-05T16:27:37.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 2.5 Pro Preview (May 06 2025)"; source version not reported; company google; rank 20; score 17.8 on publisher max score 58.4422; confidenceInterval_upper 1.5; calibrationError 70; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-06T16:28:50.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gemini-3.1-flash-lite-preview"; source version not reported; company google; rank 30; score 8.64 on publisher max score 58.4422; confidenceInterval_upper 1.1; calibrationError 83; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-03-23T21:13:29.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gemini-3.1-pro-preview (thinking high)"; source version not reported; company google; rank 2; score 46.44 on publisher max score 58.4422; confidenceInterval_upper 1.96; calibrationError 51; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-10T15:51:06.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Gemini 3.8 Flash"; source version not reported; company google; rank 2; score 44.52 on publisher max score 58.4422; confidenceInterval_upper 1.96; calibrationError 51; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=true; new=true; deprecated=false; entry timestamp 2026-09-09T19:01:48.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gemini-3-pro-preview"; source version not reported; company google; rank 5; score 37.52 on publisher max score 58.4422; confidenceInterval_upper 1.9; calibrationError 57; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-19T23:50:49.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "glm-4p5"; source version not reported; company zai; rank 31; score 8.32 on publisher max score 58.4422; confidenceInterval_upper 1.08; calibrationError 79; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at 32K Tokens, temp = null (default temp).; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-13T20:52:24.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "glm-4p5-air"; source version not reported; company zai; rank 31; score 8.12 on publisher max score 58.4422; confidenceInterval_upper 1.07; calibrationError 77; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at 32K Tokens, temp = null (default temp) ; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-13T20:52:59.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "GPT-4.1"; source version not reported; company openai; rank 42; score 5.4 on publisher max score 58.4422; confidenceInterval_upper 0.89; calibrationError 89; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-14T18:07:35.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "GPT 4.5 Preview"; source version not reported; company openai; rank 42; score 5.44 on publisher max score 58.4422; confidenceInterval_upper 0.89; calibrationError 85; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "GPT-4o (November 2024)"; source version not reported; company openai; rank 45; score 2.72 on publisher max score 58.4422; confidenceInterval_upper 0.64; calibrationError 89; contamination message This model was used as an initial filter for the dataset.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5.1-instant"; source version not reported; company openai; rank 31; score 6.8 on publisher max score 58.4422; confidenceInterval_upper 0.99; calibrationError 69; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-26T20:25:52.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5.1-thinking"; source version not reported; company openai; rank 14; score 23.68 on publisher max score 58.4422; confidenceInterval_upper 1.67; calibrationError 55; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-26T20:24:19.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5.2-2025-12-11"; source version not reported; company openai; rank 10; score 27.8 on publisher max score 58.4422; confidenceInterval_upper 1.76; calibrationError 45; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-12-15T23:38:26.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5-2025-08-07"; source version not reported; company openai; rank 10; score 25.32 on publisher max score 58.4422; confidenceInterval_upper 1.7; calibrationError 50; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at reasoning_effort: 'high'.; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-07T21:16:47.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5.4-2026-03-05 (xhigh thinking)"; source version not reported; company openai; rank 6; score 36.24 on publisher max score 58.4422; confidenceInterval_upper 1.88; calibrationError 42; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-03-10T21:09:26.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5.4-pro-2026-03-05"; source version not reported; company openai; rank 2; score 44.32 on publisher max score 58.4422; confidenceInterval_upper 1.95; calibrationError 38; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-03-23T21:12:56.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5-mini-2025-08-07"; source version not reported; company openai; rank 16; score 19.44 on publisher max score 58.4422; confidenceInterval_upper 1.55; calibrationError 65; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-08-22T21:44:43.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "gpt-5-pro-2025-10-06"; source version not reported; company openai; rank 9; score 31.64 on publisher max score 58.4422; confidenceInterval_upper 1.82; calibrationError 49; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-11-06T22:48:20.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "GPT 6 Astra"; source version not reported; company openai; rank 1; score 54.8 on publisher max score 58.4422; confidenceInterval_upper 1.94; calibrationError 39; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=true; new=true; deprecated=false; entry timestamp 2026-09-09T18:59:21.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "kimi-k2.5"; source version not reported; company moonshot; rank 10; score 24.37 on publisher max score 58.4422; confidenceInterval_upper 1.81; calibrationError 67; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-02-13T14:32:29.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Llama 4 Maverick"; source version not reported; company meta; rank 42; score 5.68 on publisher max score 58.4422; confidenceInterval_upper 0.91; calibrationError 83; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Mistral Medium 3"; source version not reported; company mistral; rank 43; score 4.52 on publisher max score 58.4422; confidenceInterval_upper 0.81; calibrationError 77; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-05-13T17:44:37.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Muse Spark"; source version not reported; company meta; rank 4; score 40.56 on publisher max score 58.4422; confidenceInterval_upper 1.92; calibrationError 50; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2026-04-08T16:57:23.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Nova Lite"; source version not reported; company amazon; rank 44; score 3.64 on publisher max score 58.4422; confidenceInterval_upper 0.73; calibrationError 82; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "Nova Pro"; source version not reported; company amazon; rank 43; score 4.4 on publisher max score 58.4422; confidenceInterval_upper 0.8; calibrationError 80; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o1 (December 2024)"; source version not reported; company openai; rank 31; score 7.96 on publisher max score 58.4422; confidenceInterval_upper 1.06; calibrationError 83; contamination message This model was used as an initial filter for the dataset.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T19:24:55.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o1 Pro"; source version not reported; company openai; rank 31; score 8.12 on publisher max score 58.4422; confidenceInterval_upper 1.07; calibrationError 82; contamination message 9% (216 prompts) failed due to a post-training bug and were counted as failures. OpenAI has been informed and is working on a fix. --- Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-10T21:16:40.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o3 (high) (April 2025)"; source version not reported; company openai; rank 16; score 20.32 on publisher max score 58.4422; confidenceInterval_upper 1.58; calibrationError 34; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-16T23:13:17.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o3 (medium) (April 2025)"; source version not reported; company openai; rank 16; score 19.2 on publisher max score 58.4422; confidenceInterval_upper 1.54; calibrationError 39; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-16T17:02:05.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o4-mini (high) (April 2025)"; source version not reported; company openai; rank 20; score 18.08 on publisher max score 58.4422; confidenceInterval_upper 1.51; calibrationError 57; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-16T23:13:24.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Official Scale AI Humanity's Last Exam full aggregate row; exact source model label "o4-mini (medium) (April 2025)"; source version not reported; company openai; rank 23; score 14.28 on publisher max score 58.4422; confidenceInterval_upper 1.37; calibrationError 59; contamination message Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions.; isNew=false; new=false; deprecated=false; entry timestamp 2025-04-16T17:02:45.000Z. The public aggregate feed does not expose a reproducible agent configuration or cost series, so no tool setup, task-level result, or cost is inferred. Source marks this competitor value with an asterisk. No tools. No tools. Anthropic updated the Sonnet 4.6 value after changing the grader model. Full set without tools. Full-set value marked with an asterisk.

Method / source