priorsBuilder

Google · released 2026-09-30

Gemini 4

Measured on 54 benchmarks.

General frontier

73.5

#2 of 133 models · best at high

includes estimates: 24% of the mix is measured for this model

Your mix’s benchmarks

2 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #5high of 58457.1%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #9high of 52342.4

Leader: Opus 5.5 · 46.4

All 54 benchmarks

show 52 more

agentic

AA-Briefcase v1.1

222 results
  1. #35high of 2221488

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #7high of 499.3%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #1high of 11551.3%
  2. #2medium of 11550.1%

Leader: this model · 51.3%

agentic

AutomationBench-AA

217 results
  1. #1high of 21777.5%

Leader: this model · 77.5%

agentic

BioMysteryBench

25 results
  1. #6high of 2576.3%

Leader: Sonnet 5.5 · 81.1%

agentic

CUA-bench

8 results
  1. #7high of 84.8%

Leader: GPT-6 Astra · 19.2%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #5high of 7619.6%

Leader: Muse Spark 1.2 · 25.4%

agentic

Legal Research Bench

75 results
  1. #4high of 7554.8%

Leader: Muse Spark 1.3 · 55.3%

agentic

Tax Agent Bench

67 results
  1. #4high of 6744.6%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #5high of 4557.6%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #5high of 3944.3%

Leader: GPT-6 Astra · 62.9%

agentic

Time Horizon Index: KSP

13 results
  1. #4high of 1352.5%

Leader: Opus 5.5 · 91.3%

agentic

Vending-Bench 2

67 results
  1. #3max of 67$13,718

Leader: GPT-6 Astra · $15,515

coding

Arena WebDev

122 results
  1. #9high of 1221678

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #2high of 7568.2%

Leader: Sonnet 5.5 · 69.8%

coding

FrontierSWE v2

21 results
  1. #5max of 2155.0%

Leader: GPT-6 Astra · 65.5%

coding

IOI

42 results
  1. #1high of 42100.0%

Leader: this model · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #7high of 6241.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #7high of 622.5%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #4high of 6283.7%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #4high of 22461.8%

Leader: Opus 5.5 · 66.9%

coding

Terminal-Bench v4-AA

214 results
  1. #6high of 21457.1%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench v1.1

103 results
  1. #2high of 10391.9%

Leader: Sonnet 5.5 · 92.4%

composite

AA Intelligence Index v4.3.2

607 results
  1. #8high of 60752.6

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #1high of 2201525

Leader: this model · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #1high of 2201553

Leader: this model · 1553

composite

Vals Index

45 results
  1. #1high of 4568.9%

Leader: this model · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #4max of 2430.6%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #14max of 2413.9%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #5max of 2415.4%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #2max of 2444.7%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #4max of 2448.0%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #147high of 177$1.99

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #144high of 177$2,407

Leader: GPT-6 Luna · $10.63

economics

Excel Modeling Benchmark

72 results
  1. #4high of 7275.2%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #1high of 7665.4%

Leader: this model · 65.4%

economics

GDP.pdf

208 results
  1. #49high of 20821.8%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #21high of 2861624

Leader: Opus 5.5 · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #61high of 52349.9%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #5high of 52384.9%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

MedCode

94 results
  1. #3high of 9458.8%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #16high of 9687.4%

Leader: Opus 5.5 · 91.4%

knowledge

Public Benefits Bench v1.1

48 results
  1. #5high of 4869.8%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #83high of 52079.7%

Leader: Kimi K3 · 88.7%

math

ProofBench v1.1

50 results
  1. #7high of 5099.0%

Leader: Sonnet 5.5 · 100.0%

reasoning

Blueprint-Bench 2

28 results
  1. #1max of 2854.4%

Leader: this model · 54.4%

reasoning

CritPt

526 results
  1. #31high of 52627.1%

Leader: GPT-5.6 Sol · 32.3%

reasoning

LegalBench

134 results
  1. #3high of 13488.3%

Leader: Fable 5 · 88.6%

reasoning

MysteryMechanism

25 results
  1. #6high of 2545.5%

Leader: GPT-6 Astra · 53.2%

reasoning

SAGE

80 results
  1. #4high of 8053.6%

Leader: Opus 4.7 · 56.1%

safety

CyberBench v1.1

43 results
  1. #2high of 4377.9%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #3high of 3244.3%

Leader: GPT-6 Astra · 56.9%