reasoning
Humanity's Last Exam
- #8max of 58455.5%
Leader: Opus 5.5 · 61.4%
Anthropic · released 2026-06-09
Measured on 103 benchmarks across 5 reasoning efforts.
General frontier
68.0
#5 of 32 models · best at high
reasoning
Leader: Opus 5.5 · 61.4%
knowledge
Leader: Opus 5.5 · 46.4
reasoning
Leader: Opus 4.8 · 95.0%
coding
Leader: GPT-6 Astra · 74.1%
math
Leader: GPT-6.1 Sol · 100.0%
reasoning
Leader: Opus 5.5 · 88.4%
agentic
Leader: Opus 5.5 · 64.8%
agentic
Leader: Gemini 3.7 Flash · 60.0%
agentic
Leader: Sonnet 5.5 · 1823
agentic
Leader: Opus 5.5 · 14.3%
agentic
Leader: Gemini 4 · 77.5%
agentic
Leader: this model · 51.1%
agentic
Leader: Muse Spark 1.2 · 25.4%
agentic
Leader: Muse Spark 1.3 · 55.3%
agentic
Leader: this model · 86.0%
agentic
Leader: this model · 41.8%
agentic
Leader: Fable 5.1 · 49.2%
agentic
Leader: GPT-6 Astra · 87.4%
agentic
Leader: Opus 5.5 · 65.2%
agentic
Leader: GPT-6 Astra · 62.9%
agentic
Leader: GPT-6 Astra · $15,515
agentic
Leader: GPT-6 Astra · $12,400
agentic
Leader: this model · 48.5%
agentic
Leader: GPT-6 Astra · 42.2%
agentic
Leader: GLM 5.2 · 99.1%
agentic
Leader: Qwen 3.8 Max · 51.3%
coding
Leader: Opus 5.5 · 1749
coding
Leader: Opus 5.5 · 1813
coding
Leader: Sonnet 5.5 · 69.8%
coding
Leader: Opus 5.5 · 97.3%
coding
Leader: Opus 5.5 · 65.3%
coding
Leader: Opus 5.5 · 54.6%
coding
Leader: this model · 88.2%
coding
Leader: GPT-6 Astra · 65.5%
coding
Leader: Fable 5.1 · 90.5%
coding
Leader: Opus 5.5 · 65.0%
coding
Leader: Opus 5.5 · 18.5%
coding
Leader: Opus 5.5 · 87.0%
coding
Leader: Opus 5.5 · 66.9%
coding
Leader: Opus 5 · 63.2%
coding
Leader: Opus 5 · 97.0%
coding
Leader: Opus 5.5 · 87.6%
coding
Leader: GPT-5.6 Sol · 65.9%
coding
Leader: Fable 5.1 · 91.4%
coding
Leader: Sonnet 5.5 · 63.6%
coding
Leader: Sonnet 5.5 · 92.4%
coding
Leader: GPT-6 Astra Pro · 93.6%
composite
Leader: Opus 5.5 · 57.6
composite
Leader: Gemini 4 · 1525
composite
Leader: Gemini 4 · 1553
composite
Leader: Gemini 4 · 68.9%
composite
Leader: this model · 74.2%
composite
Leader: Opus 5.5 · 37.3%
composite
Leader: GPT-6 Sol · 22.8%
composite
Leader: Opus 5.5 · 35.1%
composite
Leader: Opus 5.5 · 46.5%
composite
Leader: Opus 5.5 · 53.0%
cost
Leader: GPT-6 Luna · $0.00
cost
Leader: GPT-6 Luna · $10.63
creative
Leader: GPT-6 Astra Pro · 2278
creative
Leader: GPT-6 Astra · 2654
economics
Leader: Opus 5 · 73.2%
economics
Leader: Fable 5.1 · 76.7%
economics
Leader: Gemini 4 · 65.4%
economics
Leader: GPT-6 Astra · 32.2%
economics
Leader: Opus 5.5 · 1866
economics
Leader: Opus 5 · 72.1%
economics
Leader: this model · 71.7%
economics
Leader: Muse Spark 1.2 · 80.4%
instruction
Leader: Grok 4.3 · 83.3%
knowledge
Leader: Fable 5.1 · 67.2%
knowledge
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Leader: GPT-5.6 Sol · 1257
knowledge
Leader: GPT-6 Astra · 96.3%
knowledge
Leader: Gemini 3.1 Pro · 95.5%
knowledge
Leader: Opus 5 · 63.6%
knowledge
Leader: Opus 5.5 · 91.4%
knowledge
Leader: Fable 5.1 · 92.4%
knowledge
Leader: Fable 5.1 · 90.6%
knowledge
Leader: Opus 5 · 76.9%
long-context
Leader: Kimi K3 · 88.7%
long-context
Leader: Opus 5 · 1516
long-context
Leader: Sonnet 5.5 · 75.0%
math
Leader: GPT-6 Astra · 2.9%
math
Leader: GPT-6.1 Sol · 93.7%
math
Leader: Sonnet 5.5 · 100.0%
other
Leader: this model · 1308
other
Leader: Opus 5.5 · 29.1%
reasoning
Leader: this model · 98.5%
reasoning
Leader: GPT-6 Astra · 95.0%
reasoning
Leader: Gemini 4 · 54.4%
reasoning
Leader: GPT-5.6 Sol · 32.3%
reasoning
Leader: Opus 5.5 · 83.3%
reasoning
Leader: GPT-6 Astra · 53.6%
reasoning
Leader: this model · 88.6%
reasoning
Leader: GPT-6 Astra · 91.2%
reasoning
Leader: Opus 4.7 · 56.1%
reliability
Leader: GPT-5.4 Pro · 0.0%
reliability
Leader: GPT-5 Pro · 100.0%
speed
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Leader: Celeris-1 · 1461.1
speed
Leader: Qwen 2.5 7B · 0.2s
speed
Leader: Llama 3.2 3b · 199.5