reasoning
Humanity's Last Exam
- #5high of 58457.1%
Leader: Opus 5.5 · 61.4%
Google · released 2026-09-30
Measured on 54 benchmarks.
General frontier
73.5
#2 of 133 models · best at high
includes estimates: 24% of the mix is measured for this model
reasoning
Leader: Opus 5.5 · 61.4%
knowledge
Leader: Opus 5.5 · 46.4
agentic
Leader: Sonnet 5.5 · 1823
agentic
Leader: Opus 5.5 · 14.3%
agentic
Leader: this model · 51.3%
agentic
Leader: this model · 77.5%
agentic
Leader: Sonnet 5.5 · 81.1%
agentic
Leader: GPT-6 Astra · 19.2%
agentic
Leader: Muse Spark 1.2 · 25.4%
agentic
Leader: Muse Spark 1.3 · 55.3%
agentic
Leader: Fable 5.1 · 49.2%
agentic
Leader: Opus 5.5 · 65.2%
agentic
Leader: GPT-6 Astra · 62.9%
agentic
Leader: Opus 5.5 · 91.3%
agentic
Leader: GPT-6 Astra · $15,515
coding
Leader: Opus 5.5 · 1813
coding
Leader: Sonnet 5.5 · 69.8%
coding
Leader: GPT-6 Astra · 65.5%
coding
Leader: this model · 100.0%
coding
Leader: Opus 5.5 · 65.0%
coding
Leader: Opus 5.5 · 18.5%
coding
Leader: Opus 5.5 · 87.0%
coding
Leader: Opus 5.5 · 66.9%
coding
Leader: Sonnet 5.5 · 63.6%
coding
Leader: Sonnet 5.5 · 92.4%
composite
Leader: Opus 5.5 · 57.6
composite
Leader: this model · 1525
composite
Leader: this model · 1553
composite
Leader: this model · 68.9%
composite
Leader: Opus 5.5 · 37.3%
composite
Leader: GPT-6 Sol · 22.8%
composite
Leader: Opus 5.5 · 35.1%
composite
Leader: Opus 5.5 · 46.5%
composite
Leader: Opus 5.5 · 53.0%
cost
Leader: GPT-6 Luna · $0.00
cost
Leader: GPT-6 Luna · $10.63
economics
Leader: Fable 5.1 · 76.7%
economics
Leader: this model · 65.4%
economics
Leader: GPT-6 Astra · 32.2%
economics
Leader: Opus 5.5 · 1866
knowledge
Leader: Fable 5.1 · 67.2%
knowledge
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Leader: Opus 5 · 63.6%
knowledge
Leader: Opus 5.5 · 91.4%
knowledge
Leader: Opus 5 · 76.9%
long-context
Leader: Kimi K3 · 88.7%
math
Leader: Sonnet 5.5 · 100.0%
reasoning
Leader: this model · 54.4%
reasoning
Leader: GPT-5.6 Sol · 32.3%
reasoning
Leader: Fable 5 · 88.6%
reasoning
Leader: GPT-6 Astra · 53.2%
reasoning
Leader: Opus 4.7 · 56.1%
safety
Leader: GPT-6 Sol · 78.0%
safety
Leader: GPT-6 Astra · 56.9%