priorsBuilder

Anthropic · released 2026-06-09

Fable 5

Measured on 103 benchmarks across 5 reasoning efforts.

General frontier

68.0

#5 of 32 models · best at high

Reasoning efforts in your mix

high68.0#17 of 128
max65.2#24 of 128
medium64.7#25 of 128
xhigh63.5#29 of 128
low56.1#41 of 128

Your mix’s benchmarks

7 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #8max of 58455.5%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #6max of 52343.3

Leader: Opus 5.5 · 46.4

reasoning

BullshitBench v2

192 results
  1. #59xhigh of 19256.0%
  2. #83low of 19245.0%

Leader: Opus 4.8 · 95.0%

coding

DeepSWE v1.1

70 results
  1. #12xhigh of 7069.9%
  2. #13max of 7069.7%
  3. #18high of 7068.6%
  4. #26medium of 7065.4%
  5. #34low of 7059.6%

Leader: GPT-6 Astra · 74.1%

math

FrontierMath Tier 4 (v2)

70 results
  1. #7max of 7090.2%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #7max of 9781.9%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #18max of 3544.5%

Leader: Opus 5.5 · 64.8%

All 103 benchmarks

show 96 more

agentic

AA-AnalystAgent

39 results
  1. #9max of 3948.8%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #24max of 2221540

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #8high of 498.9%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench-AA

217 results
  1. #58max of 21754.1%

Leader: Gemini 4 · 77.5%

agentic

EnterpriseOps-Gym-AA

50 results
  1. #1max of 5051.1%

Leader: this model · 51.1%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #13max of 7611.3%

Leader: Muse Spark 1.2 · 25.4%

agentic

Legal Research Bench

75 results
  1. #6max of 7549.5%

Leader: Muse Spark 1.3 · 55.3%

agentic

OSWorld-Verified

13 results
  1. #1max of 1386.0%

Leader: this model · 86.0%

agentic

PostTrainBench v1.1

12 results
  1. #1max of 1241.8%

Leader: this model · 41.8%

agentic

Tax Agent Bench

67 results
  1. #8max of 6738.3%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 2.1

18 results
  1. #6xhigh of 1883.8%
  2. #8high of 1880.5%

Leader: GPT-6 Astra · 87.4%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #9max of 4541.4%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #10max of 3915.7%

Leader: GPT-6 Astra · 62.9%

agentic

Vending-Bench 2

67 results
  1. #23high of 67$5,680
  2. #32low of 67$5,019
  3. #34max of 67$4,967
  4. #36none of 67$4,530
  5. #38medium of 67$4,340

Leader: GPT-6 Astra · $15,515

agentic

Vending-Bench Arena

12 results
  1. #10max of 12$4,200

Leader: GPT-6 Astra · $12,400

agentic

Web Search Index

4 results
  1. #1max of 448.5%

Leader: this model · 48.5%

agentic

WeirdML v3

17 results
  1. #8xhigh of 1715.6%

Leader: GPT-6 Astra · 42.2%

agentic

τ²-Bench Telecom

392 results
  1. #4max of 39298.5%

Leader: GLM 5.2 · 99.1%

agentic

τ³-Banking

206 results
  1. #38max of 20638.1%

Leader: Qwen 3.8 Max · 51.3%

coding

Arena Image-to-WebDev

58 results
  1. #10high of 581622

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #20high of 1221625

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #8max of 7555.1%

Leader: Sonnet 5.5 · 69.8%

coding

Drone-Bench

11 results
  1. #5max of 1184.3%

Leader: Opus 5.5 · 97.3%

coding

FrontierCode 1.1 Extended

130 results
  1. #3xhigh of 13064.9%
  2. #6high of 13064.3%
  3. #7max of 13063.6%
  4. #13medium of 13062.8%
  5. #22low of 13060.8%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #4xhigh of 13053.5%
  2. #7high of 13052.7%
  3. #9max of 13051.6%
  4. #18medium of 13049.8%
  5. #27low of 13048.0%

Leader: Opus 5.5 · 54.6%

coding

FrontierSWE v1

17 results
  1. #1max of 1788.2%

Leader: this model · 88.2%

coding

FrontierSWE v2

21 results
  1. #7max of 2147.0%

Leader: GPT-6 Astra · 65.5%

coding

LiveCodeBench (Vals)

127 results
  1. #2max of 12789.8%

Leader: Fable 5.1 · 90.5%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #8max of 6233.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #9max of 622.0%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #9max of 6276.8%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #6max of 22461.0%

Leader: Opus 5.5 · 66.9%

coding

SWE Atlas-QnA

22 results
  1. #12xhigh of 2239.0%

Leader: Opus 5 · 63.2%

coding

SWE-bench Verified (Vals)

85 results
  1. #7max of 8595.0%

Leader: Opus 5 · 97.0%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #10max of 7480.5%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Hard-AA

386 results
  1. #2max of 38662.9%

Leader: GPT-5.6 Sol · 65.9%

coding

Terminal-Bench v2.1-AA

238 results
  1. #28max of 23884.6%

Leader: Fable 5.1 · 91.4%

coding

Terminal-Bench v4-AA

214 results
  1. #25max of 21442.4%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench v1.1

103 results
  1. #4max of 10390.4%

Leader: Sonnet 5.5 · 92.4%

coding

WeirdML v2

127 results
  1. #6xhigh of 12791.9%
  2. #11max of 12787.8%

Leader: GPT-6 Astra Pro · 93.6%

composite

AA Intelligence Index v4.3.2

607 results
  1. #19max of 60749.6

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #4high of 2201504

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #3high of 2201516

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #7max of 4561.4%

Leader: Gemini 4 · 68.9%

composite

Vals Multimodal Index v1.2

33 results
  1. #1max of 3374.2%

Leader: this model · 74.2%

composite

Vals RSI Index v1.1

24 results
  1. #9max of 2424.3%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #7max of 2418.1%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #22max of 240.0%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #9max of 2438.3%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #10max of 2440.8%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #177max of 177$8.75

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #176max of 177$11,161

Leader: GPT-6 Luna · $10.63

creative

MineBench

72 results
  1. #10max of 721906

Leader: GPT-6 Astra Pro · 2278

creative

VoxelBench (text)

54 results
  1. #7max of 542125

Leader: GPT-6 Astra · 2654

economics

CorpFin (Vals)

117 results
  1. #2max of 11771.8%

Leader: Opus 5 · 73.2%

economics

Excel Modeling Benchmark

72 results
  1. #5max of 7273.7%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #15max of 7656.3%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #35max of 20824.0%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #27max of 2861610

Leader: Opus 5.5 · 1866

economics

MortgageTax (Vals)

84 results
  1. #8max of 8468.9%

Leader: Opus 5 · 72.1%

economics

Public Benefits Bench v1 (Vals)

13 results
  1. #1max of 1371.7%

Leader: this model · 71.7%

economics

TaxEval (Vals)

125 results
  1. #5max of 12576.9%

Leader: Muse Spark 1.2 · 80.4%

instruction

IFBench

398 results
  1. #112max of 39863.5%

Leader: Grok 4.3 · 83.3%

knowledge

AA-Omniscience Accuracy

523 results
  1. #5max of 52365.3%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #173max of 52336.4%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

Arena Search

28 results
  1. #5high of 281230

Leader: GPT-5.6 Sol · 1257

knowledge

GPQA Diamond

538 results
  1. #32max of 53892.6%

Leader: GPT-6 Astra · 96.3%

knowledge

GPQA Diamond (Vals)

122 results
  1. #10max of 12293.2%

Leader: Gemini 3.1 Pro · 95.5%

knowledge

MedCode

94 results
  1. #4max of 9456.1%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #11max of 9688.5%

Leader: Opus 5.5 · 91.4%

knowledge

MMLU Pro (Vals)

122 results
  1. #3max of 12291.5%

Leader: Fable 5.1 · 92.4%

knowledge

MMMU Pro (Vals)

81 results
  1. #3max of 8189.3%

Leader: Fable 5.1 · 90.6%

knowledge

Public Benefits Bench v1.1

48 results
  1. #4max of 4870.4%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #38max of 52082.3%

Leader: Kimi K3 · 88.7%

long-context

Arena Document

43 results
  1. #5high of 431496

Leader: Opus 5 · 1516

long-context

MLCR-AA

93 results
  1. #4max of 9364.4%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Erdős

7 results
  1. #6max of 70.0%

Leader: GPT-6 Astra · 2.9%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #9max of 11587.0%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #9max of 5095.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #1high of 1161308

Leader: this model · 1308

other

T3 Code usage share

85 results
  1. #14max of 850.5%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #1max of 21598.5%
  2. #2xhigh of 21598.5%
  3. #29high of 21595.5%
  4. #51medium of 21592.5%
  5. #66low of 21590.5%

Leader: this model · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #17max of 21689.2%
  2. #20xhigh of 21688.3%
  3. #22high of 21687.5%
  4. #37medium of 21682.5%
  5. #42low of 21676.8%

Leader: GPT-6 Astra · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #6max of 2838.6%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #24max of 52628.6%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #18max of 3135.8%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #8max of 2534.5%

Leader: GPT-6 Astra · 53.6%

reasoning

LegalBench

134 results
  1. #1max of 13488.6%

Leader: this model · 88.6%

reasoning

Roboflow Visual Reasoning

102 results
  1. #45low of 10266.2%
  2. #46high of 10266.2%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #9max of 8051.9%

Leader: Opus 4.7 · 56.1%

reliability

Tool-call error rate

171 results
  1. #41max of 1710.5%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #105max of 18199.2%

Leader: GPT-5 Pro · 100.0%

speed

Latency (Time to First Token)

355 results
  1. #330max of 35581s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #245max of 35564.4

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #155max of 1783.6s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #64max of 17774.0

Leader: Llama 3.2 3b · 199.5