priorsBuilder

Anthropic · released 2026-09-22

Opus 5.5

Measured on 89 benchmarks across 5 reasoning efforts.

General frontier

76.3

#1 of 32 models · best at max

Reasoning efforts in your mix

max76.3#1 of 128
xhigh74.3#2 of 128
high73.2#5 of 128
medium71.6#13 of 128
low64.0#26 of 128

Your mix’s benchmarks

6 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #1max of 58461.4%
  2. #4xhigh of 58457.5%
  3. #7high of 58455.6%
  4. #11medium of 58454.7%
  5. #31low of 58448.3%

Leader: this model · 61.4%

knowledge

AA-Omniscience

523 results
  1. #1max of 52346.4
  2. #7xhigh of 52342.6
  3. #15high of 52340.6
  4. #17medium of 52340.3
  5. #19low of 52338.9

Leader: this model · 46.4

reasoning

BullshitBench v2

192 results
  1. #40low of 19263.0%
  2. #43max of 19262.0%

Leader: Opus 4.8 · 95.0%

math

FrontierMath Tier 4 (v2)

70 results
  1. #6max of 7095.0%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #1max of 9788.4%

Leader: this model · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #1max of 3564.8%

Leader: this model · 64.8%

All 89 benchmarks

show 83 more

agentic

AA-AnalystAgent

39 results
  1. #4max of 3956.3%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #2max of 2221807
  2. #3xhigh of 2221766
  3. #5high of 2221690
  4. #13medium of 2221627
  5. #70low of 2221279

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #1high of 4914.3%

Leader: this model · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #4max of 11542.5%
  2. #9xhigh of 11535.8%
  3. #12high of 11533.0%
  4. #19medium of 11529.5%
  5. #28low of 11524.2%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #3max of 21769.5%
  2. #13xhigh of 21765.0%
  3. #19high of 21763.2%
  4. #27medium of 21761.2%
  5. #66low of 21752.9%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #4max of 2579.3%

Leader: Sonnet 5.5 · 81.1%

agentic

CUA-bench

8 results
  1. #2max of 814.0%

Leader: GPT-6 Astra · 19.2%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #38max of 763.8%

Leader: Muse Spark 1.2 · 25.4%

agentic

ITBench-AA

45 results
  1. #23max of 4538.2%

Leader: GPT-5.6 Sol · 56.2%

agentic

Legal Research Bench

75 results
  1. #5max of 7550.5%

Leader: Muse Spark 1.3 · 55.3%

agentic

Tax Agent Bench

67 results
  1. #2max of 6745.6%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #1max of 4565.2%

Leader: this model · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #3max of 3947.1%

Leader: GPT-6 Astra · 62.9%

agentic

Time Horizon Index: KSP

13 results
  1. #1max of 1391.3%

Leader: this model · 91.3%

agentic

Vending-Bench 2

67 results
  1. #9max of 67$9,235

Leader: GPT-6 Astra · $15,515

agentic

Vending-Bench Arena

12 results
  1. #2max of 12$8,100

Leader: GPT-6 Astra · $12,400

agentic

WeirdML v3

17 results
  1. #3xhigh of 1731.2%

Leader: GPT-6 Astra · 42.2%

coding

Arena Image-to-WebDev

58 results
  1. #1max of 581749

Leader: this model · 1749

coding

Arena WebDev

122 results
  1. #1max of 1221813

Leader: this model · 1813

coding

Code Migration

75 results
  1. #4max of 7566.6%

Leader: Sonnet 5.5 · 69.8%

coding

Drone-Bench

11 results
  1. #1max of 1197.3%

Leader: this model · 97.3%

coding

FrontierCode 1.1 Extended

130 results
  1. #1medium of 13065.3%
  2. #2high of 13065.2%
  3. #10max of 13063.6%
  4. #11xhigh of 13063.5%
  5. #27low of 13060.3%

Leader: this model · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #1medium of 13054.6%
  2. #2max of 13054.4%
  3. #3high of 13054.0%
  4. #10xhigh of 13051.4%
  5. #35low of 13047.3%

Leader: this model · 54.6%

coding

FrontierSWE v2

21 results
  1. #2max of 2162.3%

Leader: GPT-6 Astra · 65.5%

coding

IOI

42 results
  1. #4max of 4295.1%

Leader: Gemini 4 · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #1max of 6265.0%

Leader: this model · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #1max of 6218.5%

Leader: this model · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #1max of 6287.0%

Leader: this model · 87.0%

coding

SciCode

224 results
  1. #1max of 22466.9%
  2. #2xhigh of 22465.0%
  3. #9high of 22460.4%
  4. #13medium of 22459.3%
  5. #20low of 22458.6%

Leader: this model · 66.9%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #1high of 7487.6%

Leader: this model · 87.6%

coding

Terminal-Bench Science-AA

45 results
  1. #2xhigh of 4561.9%
  2. #3max of 4559.0%
  3. #7high of 4549.0%
  4. #9medium of 4543.3%
  5. #13low of 4524.3%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v4-AA

214 results
  1. #2max of 21459.6%
  2. #3xhigh of 21459.6%
  3. #8high of 21456.6%
  4. #13medium of 21452.5%
  5. #40low of 21431.3%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench 1-100

22 results
  1. #1xhigh of 2230.4%

Leader: this model · 30.4%

coding

Vibe Code Bench v1.1

103 results
  1. #5max of 10390.3%

Leader: Sonnet 5.5 · 92.4%

composite

AA Intelligence Index v4.3.2

607 results
  1. #1max of 60757.6
  2. #3xhigh of 60756.0
  3. #4high of 60753.6
  4. #12medium of 60751.2
  5. #47low of 60742.3

Leader: this model · 57.6

composite

Arena Text

220 results
  1. #2high of 2201507

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #10high of 2201498

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #3max of 4567.0%

Leader: Gemini 4 · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #1max of 2437.3%

Leader: this model · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #12max of 2414.5%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #1max of 2435.1%

Leader: this model · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #1max of 2446.5%

Leader: this model · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #1max of 2453.0%

Leader: this model · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #95low of 177$0.55
  2. #131medium of 177$1.34
  3. #145high of 177$1.82
  4. #163xhigh of 177$3.46
  5. #175max of 177$5.98

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #90low of 177$860
  2. #128medium of 177$1,627
  3. #140high of 177$2,172
  4. #160xhigh of 177$4,057
  5. #174max of 177$8,708

Leader: GPT-6 Luna · $10.63

creative

MineBench

72 results
  1. #3max of 722186

Leader: GPT-6 Astra Pro · 2278

creative

VoxelBench (text)

54 results
  1. #2max of 542593

Leader: GPT-6 Astra · 2654

economics

Excel Modeling Benchmark

72 results
  1. #2max of 7275.9%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #9max of 7658.6%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #10high of 20828.8%
  2. #18xhigh of 20826.6%
  3. #20max of 20826.2%
  4. #25medium of 20825.6%
  5. #26low of 20825.6%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #1max of 2861866
  2. #3xhigh of 2861836
  3. #10high of 2861705
  4. #35medium of 2861584
  5. #108low of 2861234

Leader: this model · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #2max of 52366.2%
  2. #4xhigh of 52365.4%
  3. #7high of 52364.6%
  4. #8medium of 52364.5%
  5. #9low of 52363.5%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #146max of 52341.4%
  2. #184xhigh of 52334.3%
  3. #198high of 52332.4%
  4. #199low of 52332.4%
  5. #204medium of 52331.6%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

Harvey LAB-AA v1.1

25 results
  1. #7max of 254.2%

Leader: Grok 4.7 · 9.4%

knowledge

MedCode

94 results
  1. #18max of 9449.8%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #1max of 9691.4%

Leader: this model · 91.4%

knowledge

MMMU-Pro

264 results
  1. #1max of 26487.7%
  2. #3xhigh of 26486.6%
  3. #7high of 26485.8%
  4. #8medium of 26485.7%
  5. #17low of 26484.7%

Leader: this model · 87.7%

knowledge

Public Benefits Bench v1.1

48 results
  1. #3max of 4870.6%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #5max of 52084.7%
  2. #6xhigh of 52084.7%
  3. #8medium of 52084.3%
  4. #34high of 52082.7%
  5. #62low of 52080.7%

Leader: Kimi K3 · 88.7%

long-context

MLCR-AA

93 results
  1. #3max of 9366.7%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Erdős

7 results
  1. #2max of 72.9%

Leader: GPT-6 Astra · 2.9%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #3max of 11591.2%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #2max of 50100.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #10high of 1161293

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #1max of 8529.1%

Leader: this model · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #5high of 21598.5%
  2. #16max of 21597.5%
  3. #17xhigh of 21597.5%
  4. #18medium of 21597.5%
  5. #77low of 21588.5%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #4high of 21693.3%
  2. #6xhigh of 21692.5%
  3. #9max of 21691.7%
  4. #23medium of 21687.5%
  5. #52low of 21670.1%

Leader: GPT-6 Astra · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #2max of 2851.2%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #2max of 52631.7%
  2. #3xhigh of 52631.7%
  3. #11high of 52630.9%
  4. #27medium of 52627.7%
  5. #71low of 52617.7%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #1max of 3183.3%

Leader: this model · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #3xhigh of 2544.6%

Leader: GPT-6 Astra · 53.6%

reasoning

MysteryMechanism

25 results
  1. #2max of 2549.5%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #4high of 10285.9%
  2. #8low of 10283.0%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #37max of 8045.8%

Leader: Opus 4.7 · 56.1%

reliability

Tool-call error rate

171 results
  1. #45max of 1710.6%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #85max of 18199.6%

Leader: GPT-5 Pro · 100.0%

safety

CWE-Bench

18 results
  1. #7max of 1858.3%

Leader: Grok 4.7 · 68.3%

safety

CyberBench v1.1

43 results
  1. #8max of 4374.6%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #4max of 3233.6%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #256low of 3558.7s
  2. #292medium of 35522s
  3. #312high of 35540s
  4. #342xhigh of 355128s
  5. #355max of 355704s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #156max of 35597.0
  2. #193xhigh of 35583.2
  3. #203high of 35579.9
  4. #207medium of 35579.2
  5. #210low of 35578.2

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #145max of 1783.1s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #53max of 17781.0

Leader: Llama 3.2 3b · 199.5