priorsBuilder

Anthropic · released 2026-07-24

Opus 5

Measured on 109 benchmarks across 5 reasoning efforts.

General frontier

65.3

#8 of 32 models · best at max

Reasoning efforts in your mix

max65.3#22 of 128
xhigh65.2#23 of 128
high63.7#28 of 128
medium61.2#32 of 128
low53.4#45 of 128

Your mix’s benchmarks

8 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #10max of 58454.9%
  2. #14xhigh of 58454.4%
  3. #18high of 58452.8%
  4. #22medium of 58451.3%
  5. #53low of 58443.4%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #22max of 52337.1
  2. #23xhigh of 52335.4
  3. #25high of 52333.7
  4. #29medium of 52331.0
  5. #37low of 52328.6

Leader: Opus 5.5 · 46.4

reasoning

ARC-AGI-3

57 results
  1. #13high of 5730.2%

Leader: GPT-6 Astra · 99.9%

reasoning

BullshitBench v2

192 results
  1. #21low of 19273.0%
  2. #27xhigh of 19270.0%

Leader: Opus 4.8 · 95.0%

coding

DeepSWE v1.1

70 results
  1. #3max of 7073.6%
  2. #6xhigh of 7073.2%
  3. #7high of 7072.8%
  4. #17medium of 7068.9%
  5. #36low of 7058.1%

Leader: GPT-6 Astra · 74.1%

math

FrontierMath Tier 4 (v2)

70 results
  1. #16max of 7073.2%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #9max of 9780.6%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #12xhigh of 3553.9%
  2. #13max of 3551.8%
  3. #15high of 3550.3%
  4. #17medium of 3544.9%
  5. #24low of 3534.9%

Leader: Opus 5.5 · 64.8%

All 109 benchmarks

show 101 more

agentic

AA-AnalystAgent

39 results
  1. #5max of 3953.8%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #7max of 2221660
  2. #11xhigh of 2221635
  3. #20high of 2221561
  4. #47medium of 2221435
  5. #83low of 2221209

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #9max of 498.1%
  2. #10high of 498.0%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #22max of 11526.9%
  2. #25xhigh of 11525.3%
  3. #29medium of 11523.9%
  4. #38high of 11520.6%
  5. #39low of 11520.4%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #48max of 21756.6%
  2. #57medium of 21754.3%
  3. #61high of 21753.6%
  4. #63xhigh of 21753.2%
  5. #70low of 21751.8%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #5max of 2579.3%

Leader: Sonnet 5.5 · 81.1%

agentic

CUA-bench

8 results
  1. #4max of 89.0%

Leader: GPT-6 Astra · 19.2%

agentic

EnterpriseOps-Gym-AA

50 results
  1. #7max of 5047.5%

Leader: Fable 5 · 51.1%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #29max of 766.7%

Leader: Muse Spark 1.2 · 25.4%

agentic

Legal Research Bench

75 results
  1. #2max of 7555.3%

Leader: Muse Spark 1.3 · 55.3%

agentic

OSWorld 2.0

6 results
  1. #1max of 631.4%
  2. #2xhigh of 630.2%
  3. #3high of 629.0%
  4. #5medium of 625.3%
  5. #6low of 622.3%

Leader: this model · 31.4%

agentic

OSWorld-Verified

13 results
  1. #2max of 1383.4%

Leader: Fable 5 · 86.0%

agentic

PostTrainBench v1.1

12 results
  1. #3max of 1235.0%

Leader: Fable 5 · 41.8%

agentic

SkillsBench

34 results
  1. #7max of 3460.4%

Leader: DeepSeek V4.1 Flash · 69.8%

agentic

Tax Agent Bench

67 results
  1. #3max of 6745.5%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #7max of 4553.5%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #8max of 3927.1%

Leader: GPT-6 Astra · 62.9%

agentic

Time Horizon Index: KSP

13 results
  1. #7max of 1318.8%

Leader: Opus 5.5 · 91.3%

agentic

Vending-Bench 2

67 results
  1. #4max of 67$11,182

Leader: GPT-6 Astra · $15,515

agentic

Vending-Bench Arena

12 results
  1. #6max of 12$7,000

Leader: GPT-6 Astra · $12,400

agentic

WeirdML v3

17 results
  1. #7xhigh of 1719.0%

Leader: GPT-6 Astra · 42.2%

agentic

τ³-Banking

206 results
  1. #16high of 20644.7%
  2. #19xhigh of 20643.3%
  3. #23max of 20642.1%
  4. #36medium of 20638.6%
  5. #64low of 20630.3%

Leader: Qwen 3.8 Max · 51.3%

coding

Arena Image-to-WebDev

58 results
  1. #7max of 581661

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #7max of 1221691
  2. #11high of 1221657

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #6max of 7557.5%

Leader: Sonnet 5.5 · 69.8%

coding

Drone-Bench

11 results
  1. #4max of 1186.9%

Leader: Opus 5.5 · 97.3%

coding

FrontierCode 1.1 Extended

130 results
  1. #8medium of 13063.6%
  2. #40max of 13058.9%
  3. #42high of 13058.5%
  4. #50xhigh of 13056.9%
  5. #60low of 13055.8%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #5medium of 13053.4%
  2. #26max of 13048.0%
  3. #29high of 13048.0%
  4. #50xhigh of 13043.6%
  5. #60low of 13041.9%

Leader: Opus 5.5 · 54.6%

coding

FrontierSWE v2

21 results
  1. #6max of 2152.0%

Leader: GPT-6 Astra · 65.5%

coding

IOI

42 results
  1. #8max of 4284.3%

Leader: Gemini 4 · 100.0%

coding

LiveCodeBench (Vals)

127 results
  1. #4max of 12789.0%

Leader: Fable 5.1 · 90.5%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #6max of 6241.5%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #5max of 623.0%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #6max of 6282.3%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #35max of 22456.4%
  2. #44xhigh of 22455.7%
  3. #47high of 22455.4%
  4. #90medium of 22451.5%
  5. #112low of 22449.2%

Leader: Opus 5.5 · 66.9%

coding

SWE Atlas-QnA

22 results
  1. #1xhigh of 2263.2%

Leader: this model · 63.2%

coding

SWE-bench Verified (Vals)

85 results
  1. #1max of 8597.0%

Leader: this model · 97.0%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #5high of 7484.6%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Science-AA

45 results
  1. #12max of 4528.6%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v2.1-AA

238 results
  1. #8max of 23889.1%
  2. #12xhigh of 23888.0%
  3. #18high of 23887.6%
  4. #21medium of 23886.1%
  5. #59low of 23876.4%

Leader: Fable 5.1 · 91.4%

coding

Terminal-Bench v4-AA

214 results
  1. #18max of 21449.0%
  2. #20xhigh of 21446.5%
  3. #21high of 21446.0%
  4. #34medium of 21434.3%
  5. #48low of 21426.3%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench 1-100

22 results
  1. #2max of 2228.5%

Leader: Opus 5.5 · 30.4%

coding

Vibe Code Bench v1.1

103 results
  1. #9max of 10388.4%

Leader: Sonnet 5.5 · 92.4%

coding

WeirdML v2

127 results
  1. #7xhigh of 12791.8%
  2. #8high of 12791.6%
  3. #13max of 12786.3%

Leader: GPT-6 Astra Pro · 93.6%

composite

AA Intelligence Index v4.3.2

607 results
  1. #16max of 60750.8
  2. #18xhigh of 60749.7
  3. #22high of 60748.1
  4. #35medium of 60744.8
  5. #65low of 60739.4

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #13high of 2201490
  2. #14max of 2201489

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #25max of 2201486
  2. #31high of 2201484

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #5max of 4563.7%

Leader: Gemini 4 · 68.9%

composite

Vals Multimodal Index v1.2

33 results
  1. #2max of 3373.9%

Leader: Fable 5 · 74.2%

composite

Vals RSI Index v1.1

24 results
  1. #3max of 2433.0%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #10max of 2414.8%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #2max of 2431.8%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #6max of 2442.6%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #6max of 2442.8%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #121low of 177$1.10
  2. #153medium of 177$2.19
  3. #164high of 177$3.61
  4. #169xhigh of 177$4.88
  5. #173max of 177$5.86

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #124low of 177$1,561
  2. #150medium of 177$2,732
  3. #162high of 177$4,332
  4. #169xhigh of 177$5,868
  5. #173max of 177$7,275

Leader: GPT-6 Luna · $10.63

creative

MineBench

72 results
  1. #5max of 722051

Leader: GPT-6 Astra Pro · 2278

creative

VoxelBench (text)

54 results
  1. #5max of 542176

Leader: GPT-6 Astra · 2654

economics

CorpFin (Vals)

117 results
  1. #1max of 11773.2%

Leader: this model · 73.2%

economics

Excel Modeling Benchmark

72 results
  1. #6max of 7273.6%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #8max of 7658.6%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #50max of 20821.6%
  2. #53xhigh of 20821.0%
  3. #62medium of 20820.0%
  4. #64high of 20819.6%
  5. #79low of 20817.2%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #7max of 2861722
  2. #11xhigh of 2861692
  3. #31high of 2861594
  4. #49medium of 2861490
  5. #97low of 2861301

Leader: Opus 5.5 · 1866

economics

MortgageTax (Vals)

84 results
  1. #1max of 8472.1%

Leader: this model · 72.1%

economics

TaxEval (Vals)

125 results
  1. #23max of 12575.1%

Leader: Muse Spark 1.2 · 80.4%

knowledge

AA-Omniscience Accuracy

523 results
  1. #15max of 52360.9%
  2. #22xhigh of 52359.5%
  3. #24high of 52358.9%
  4. #31medium of 52357.1%
  5. #34low of 52356.0%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #151xhigh of 52340.5%
  2. #155medium of 52339.3%
  3. #156max of 52339.2%
  4. #160high of 52338.8%
  5. #167low of 52337.8%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

GPQA Diamond

538 results
  1. #12xhigh of 53893.7%
  2. #13high of 53893.7%
  3. #22max of 53893.2%
  4. #44medium of 53891.9%
  5. #84low of 53888.9%

Leader: GPT-6 Astra · 96.3%

knowledge

GPQA Diamond (Vals)

122 results
  1. #8max of 12293.4%

Leader: Gemini 3.1 Pro · 95.5%

knowledge

MedCode

94 results
  1. #1max of 9463.6%

Leader: this model · 63.6%

knowledge

MedScribe

96 results
  1. #4max of 9691.0%

Leader: Opus 5.5 · 91.4%

knowledge

MMLU Pro (Vals)

122 results
  1. #2max of 12291.6%

Leader: Fable 5.1 · 92.4%

knowledge

MMMU Pro (Vals)

81 results
  1. #2max of 8189.9%

Leader: Fable 5.1 · 90.6%

knowledge

MMMU-Pro

264 results
  1. #15max of 26484.7%
  2. #22xhigh of 26484.0%
  3. #32high of 26482.4%
  4. #37medium of 26481.6%
  5. #53low of 26479.8%

Leader: Opus 5.5 · 87.7%

knowledge

Public Benefits Bench v1.1

48 results
  1. #1max of 4876.9%

Leader: this model · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #43medium of 52082.0%
  2. #54low of 52081.3%
  3. #66xhigh of 52080.3%
  4. #92max of 52079.3%
  5. #100high of 52079.0%

Leader: Kimi K3 · 88.7%

long-context

Arena Document

43 results
  1. #1high of 431516

Leader: this model · 1516

long-context

MLCR-AA

93 results
  1. #6high of 9359.4%
  2. #7xhigh of 9358.3%
  3. #8medium of 9356.1%
  4. #9max of 9355.6%
  5. #11low of 9353.9%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #11max of 11585.6%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #8max of 5099.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #14high of 1161287

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #4max of 859.1%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #11high of 21597.5%
  2. #12max of 21597.5%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #12max of 21690.4%
  2. #21high of 21688.3%

Leader: GPT-6 Astra · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #18max of 2830.4%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #19max of 52629.1%
  2. #25high of 52628.3%
  3. #28xhigh of 52627.7%
  4. #34medium of 52626.9%
  5. #48low of 52623.1%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #6max of 3160.8%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #14max of 2530.5%

Leader: GPT-6 Astra · 53.6%

reasoning

LegalBench

134 results
  1. #8max of 13487.0%

Leader: Fable 5 · 88.6%

reasoning

MysteryMechanism

25 results
  1. #7max of 2537.4%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #27high of 10274.2%
  2. #37low of 10271.5%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #21max of 8049.4%

Leader: Opus 4.7 · 56.1%

reasoning

SimpleQA Verified

47 results
  1. #10max of 4763.3%

Leader: Gemini 3.1 Pro · 77.5%

reliability

Tool-call error rate

171 results
  1. #46max of 1710.6%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #45max of 18199.9%

Leader: GPT-5 Pro · 100.0%

safety

CyberBench v1.1

43 results
  1. #29max of 4365.4%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #10max of 3212.2%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #202low of 3553.1s
  2. #236medium of 3555.2s
  3. #289high of 35520s
  4. #306xhigh of 35532s
  5. #322max of 35557s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #286high of 35554.0
  2. #288xhigh of 35553.9
  3. #291low of 35553.2
  4. #292max of 35553.1
  5. #294medium of 35552.9

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #144max of 1783.1s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #73max of 17771.0

Leader: Llama 3.2 3b · 199.5