priorsBuilder

OpenAI · released 2026-07-09

GPT-5.6 Sol

Measured on 113 benchmarks across 6 reasoning efforts.

General frontier

58.3

#10 of 32 models · best at max

Reasoning efforts in your mix

max58.3#35 of 128
xhigh56.4#37 of 128
high55.5#42 of 128
medium51.7#46 of 128
low45.1#68 of 128
none36.2#104 of 128

Your mix’s benchmarks

8 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #25max of 58449.5%
  2. #37xhigh of 58447.3%
  3. #43high of 58446.0%
  4. #67medium of 58442.2%
  5. #86low of 58439.4%
  6. #230none of 58416.7%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #56max of 52322.0
  2. #59xhigh of 52321.0
  3. #64high of 52320.4
  4. #67medium of 52319.4
  5. #69low of 52318.9
  6. #119none of 5231.1

Leader: Opus 5.5 · 46.4

reasoning

ARC-AGI-3

57 results
  1. #20max of 577.8%
  2. #22xhigh of 577.0%
  3. #25high of 572.1%
  4. #30medium of 571.1%
  5. #41low of 570.3%

Leader: GPT-6 Astra · 99.9%

reasoning

BullshitBench v2

192 results
  1. #74low of 19248.0%
  2. #77max of 19247.0%

Leader: Opus 4.8 · 95.0%

coding

DeepSWE v1.1

70 results
  1. #9max of 7072.7%
  2. #11xhigh of 7070.7%
  3. #15high of 7069.4%
  4. #32medium of 7061.1%
  5. #54low of 7045.4%

Leader: GPT-6 Astra · 74.1%

math

FrontierMath Tier 4 (v2)

70 results
  1. #12max of 7082.9%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #26xhigh of 9764.8%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #22max of 3537.3%

Leader: Opus 5.5 · 64.8%

All 113 benchmarks

show 105 more

agentic

AA-AnalystAgent

39 results
  1. #10max of 3947.5%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #39max of 2221480
  2. #46xhigh of 2221438
  3. #57high of 2221370
  4. #80medium of 2221242
  5. #109low of 2221044
  6. #115none of 2221008

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #12xhigh of 496.2%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #21max of 11528.8%
  2. #24xhigh of 11526.3%
  3. #27high of 11524.8%
  4. #40medium of 11519.6%
  5. #67low of 11511.7%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #31max of 21760.1%
  2. #53high of 21755.3%
  3. #54xhigh of 21755.3%
  4. #71medium of 21751.3%
  5. #89low of 21741.0%
  6. #125none of 21722.7%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #10max of 2571.1%

Leader: Sonnet 5.5 · 81.1%

agentic

CUA-bench

8 results
  1. #5max of 88.3%

Leader: GPT-6 Astra · 19.2%

agentic

EnterpriseOps-Gym-AA

50 results
  1. #20max of 5042.9%

Leader: Fable 5 · 51.1%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #43max of 762.5%

Leader: Muse Spark 1.2 · 25.4%

agentic

ITBench-AA

45 results
  1. #1max of 4556.2%

Leader: this model · 56.2%

agentic

Legal Research Bench

75 results
  1. #9max of 7548.1%

Leader: Muse Spark 1.3 · 55.3%

agentic

OSWorld 2.0

6 results
  1. #4max of 627.3%

Leader: Opus 5 · 31.4%

agentic

PostTrainBench v1.1

12 results
  1. #2max of 1236.2%

Leader: Fable 5 · 41.8%

agentic

SkillsBench

34 results
  1. #14max of 3454.1%

Leader: DeepSeek V4.1 Flash · 69.8%

agentic

Tax Agent Bench

67 results
  1. #20max of 6731.1%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #11max of 4537.9%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #9max of 3920.0%

Leader: GPT-6 Astra · 62.9%

agentic

Time Horizon Index: KSP

13 results
  1. #5max of 1323.8%

Leader: Opus 5.5 · 91.3%

agentic

Vending-Bench 2

67 results
  1. #8max of 67$9,619

Leader: GPT-6 Astra · $15,515

agentic

Vending-Bench Arena

12 results
  1. #5max of 12$7,400

Leader: GPT-6 Astra · $12,400

agentic

Web Search Index

4 results
  1. #2max of 445.2%

Leader: Fable 5 · 48.5%

agentic

WeirdML v3

17 results
  1. #9xhigh of 1715.0%

Leader: GPT-6 Astra · 42.2%

agentic

τ²-Bench Telecom

392 results
  1. #95max of 39285.1%
  2. #98xhigh of 39284.8%
  3. #110high of 39283.3%
  4. #119medium of 39281.0%
  5. #135low of 39276.0%

Leader: GLM 5.2 · 99.1%

agentic

τ³-Banking

206 results
  1. #17max of 20644.3%
  2. #39xhigh of 20638.1%
  3. #42high of 20636.7%
  4. #44medium of 20636.5%
  5. #70low of 20629.1%
  6. #101none of 20619.6%

Leader: Qwen 3.8 Max · 51.3%

coding

Arena Image-to-WebDev

58 results
  1. #12xhigh of 581608

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #23xhigh of 1221618

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #10max of 7552.9%

Leader: Sonnet 5.5 · 69.8%

coding

FrontierCode 1.1 Extended

130 results
  1. #24max of 13060.6%
  2. #30xhigh of 13060.0%
  3. #41high of 13058.7%
  4. #68medium of 13054.7%
  5. #91low of 13050.0%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #34max of 13047.5%
  2. #37xhigh of 13046.8%
  3. #48high of 13045.1%
  4. #71medium of 13039.9%
  5. #88low of 13035.4%

Leader: Opus 5.5 · 54.6%

coding

FrontierSWE v2

21 results
  1. #9max of 2132.2%

Leader: GPT-6 Astra · 65.5%

coding

IOI

42 results
  1. #5max of 4291.2%

Leader: Gemini 4 · 100.0%

coding

LiveCodeBench (Vals)

127 results
  1. #52max of 12782.6%

Leader: Fable 5.1 · 90.5%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #10max of 6223.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #12max of 621.5%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #8max of 6277.6%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #22high of 22457.8%
  2. #26medium of 22457.4%
  3. #29max of 22457.1%
  4. #30xhigh of 22457.1%
  5. #37low of 22456.4%
  6. #122none of 22447.7%

Leader: Opus 5.5 · 66.9%

coding

SWE Atlas-QnA

22 results
  1. #7xhigh of 2246.0%

Leader: Opus 5 · 63.2%

coding

SWE-bench Verified (Vals)

85 results
  1. #3max of 8596.2%

Leader: Opus 5 · 97.0%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #3max of 7485.8%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Hard-AA

386 results
  1. #1max of 38665.9%
  2. #3medium of 38662.9%
  3. #5high of 38662.1%
  4. #6xhigh of 38661.4%
  5. #8low of 38660.6%

Leader: this model · 65.9%

coding

Terminal-Bench Science-AA

45 results
  1. #14max of 4522.4%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v2.1-AA

238 results
  1. #6xhigh of 23889.5%
  2. #14max of 23888.0%
  3. #20high of 23887.3%
  4. #23medium of 23886.1%
  5. #58low of 23876.8%
  6. #66none of 23874.2%

Leader: Fable 5.1 · 91.4%

coding

Terminal-Bench v4-AA

214 results
  1. #29max of 21439.9%
  2. #52xhigh of 21424.7%
  3. #57high of 21420.7%
  4. #66medium of 21414.6%
  5. #122low of 2141.0%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench 1-100

22 results
  1. #8max of 2220.0%

Leader: Opus 5.5 · 30.4%

coding

Vibe Code Bench v1.1

103 results
  1. #22max of 10380.5%

Leader: Sonnet 5.5 · 92.4%

coding

WeirdML v2

127 results
  1. #10max of 12788.8%
  2. #12xhigh of 12787.0%

Leader: GPT-6 Astra Pro · 93.6%

composite

AA Intelligence Index v4.3.2

607 results
  1. #26max of 60747.0
  2. #40xhigh of 60744.0
  3. #46high of 60742.3
  4. #66medium of 60739.2
  5. #100low of 60733.5
  6. #129none of 60728.3

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #19xhigh of 2201485

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #20xhigh of 2201488

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #10max of 4558.0%

Leader: Gemini 4 · 68.9%

composite

Vals Multimodal Index v1.2

33 results
  1. #4max of 3372.6%

Leader: Fable 5 · 74.2%

composite

Vals RSI Index v1.1

24 results
  1. #11max of 2423.9%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #6max of 2418.5%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #20max of 240.0%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #8max of 2441.0%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #18max of 2436.1%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #66low of 177$0.26
  2. #90medium of 177$0.50
  3. #105high of 177$0.81
  4. #127xhigh of 177$1.18
  5. #146max of 177$1.99

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #81low of 177$637
  2. #96medium of 177$997
  3. #117high of 177$1,487
  4. #136xhigh of 177$2,082
  5. #155max of 177$3,465

Leader: GPT-6 Luna · $10.63

creative

VoxelBench (text)

54 results
  1. #6max of 542166

Leader: GPT-6 Astra · 2654

economics

CorpFin (Vals)

117 results
  1. #39max of 11764.4%

Leader: Opus 5 · 73.2%

economics

Excel Modeling Benchmark

72 results
  1. #7max of 7272.3%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #27max of 7653.8%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #12high of 20827.8%
  2. #13xhigh of 20827.6%
  3. #14max of 20827.2%
  4. #23medium of 20826.2%
  5. #55low of 20821.0%
  6. #91none of 20815.2%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #28max of 2861609
  2. #37xhigh of 2861569
  3. #48high of 2861504
  4. #66medium of 2861421
  5. #96low of 2861304
  6. #107none of 2861242

Leader: Opus 5.5 · 1866

economics

MortgageTax (Vals)

84 results
  1. #27max of 8467.3%

Leader: Opus 5 · 72.1%

economics

TaxEval (Vals)

125 results
  1. #30max of 12574.8%

Leader: Muse Spark 1.2 · 80.4%

instruction

IFBench

398 results
  1. #46max of 39872.7%
  2. #61xhigh of 39871.0%
  3. #72medium of 39869.6%
  4. #73high of 39869.2%
  5. #93low of 39866.5%

Leader: Grok 4.3 · 83.3%

knowledge

AA-Omniscience Accuracy

523 results
  1. #23max of 52359.4%
  2. #26xhigh of 52358.8%
  3. #27high of 52358.4%
  4. #29medium of 52357.8%
  5. #30low of 52357.2%
  6. #68none of 52348.7%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #388low of 52310.6%
  2. #410medium of 5239.2%
  3. #423high of 5238.8%
  4. #442xhigh of 5238.1%
  5. #449max of 5237.8%
  6. #461none of 5237.2%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

Arena Search

28 results
  1. #1xhigh of 281257

Leader: this model · 1257

knowledge

GPQA Diamond

538 results
  1. #7max of 53894.1%
  2. #25xhigh of 53893.1%
  3. #29high of 53892.8%
  4. #33medium of 53892.6%
  5. #68low of 53889.8%
  6. #200none of 53879.0%

Leader: GPT-6 Astra · 96.3%

knowledge

GPQA Diamond (Vals)

122 results
  1. #2max of 12295.2%

Leader: Gemini 3.1 Pro · 95.5%

knowledge

MedCode

94 results
  1. #40max of 9444.0%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #29max of 9685.2%

Leader: Opus 5.5 · 91.4%

knowledge

MMLU Pro (Vals)

122 results
  1. #16max of 12289.1%

Leader: Fable 5.1 · 92.4%

knowledge

MMMU Pro (Vals)

81 results
  1. #6max of 8188.8%

Leader: Fable 5.1 · 90.6%

knowledge

MMMU-Pro

264 results
  1. #25max of 26483.4%
  2. #30xhigh of 26482.7%
  3. #36high of 26481.8%
  4. #38medium of 26481.4%
  5. #41low of 26481.0%
  6. #130none of 26471.9%

Leader: Opus 5.5 · 87.7%

knowledge

Public Benefits Bench v1.1

48 results
  1. #16max of 4866.5%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #11max of 52084.0%
  2. #40xhigh of 52082.3%
  3. #48high of 52081.7%
  4. #70medium of 52080.3%
  5. #116low of 52078.0%
  6. #258none of 52062.3%

Leader: Kimi K3 · 88.7%

long-context

Arena Document

43 results
  1. #10xhigh of 431483

Leader: Opus 5 · 1516

long-context

MLCR-AA

93 results
  1. #22max of 9326.1%
  2. #34xhigh of 9318.3%
  3. #38high of 9317.2%
  4. #47medium of 9315.0%
  5. #54low of 9312.8%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Erdős

7 results
  1. #7max of 70.0%

Leader: GPT-6 Astra · 2.9%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #6max of 11589.1%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #14max of 5083.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #15xhigh of 1161287

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #6max of 855.4%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #10xhigh of 21597.5%
  2. #20high of 21597.0%
  3. #22max of 21596.5%
  4. #50medium of 21592.5%
  5. #117low of 21574.5%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #5max of 21692.5%
  2. #13xhigh of 21690.0%
  3. #26high of 21685.4%
  4. #61medium of 21667.1%
  5. #101low of 21642.5%

Leader: GPT-6 Astra · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #11max of 2833.6%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #1max of 52632.3%
  2. #23xhigh of 52628.6%
  3. #38high of 52625.7%
  4. #49medium of 52622.9%
  5. #89low of 52614.9%
  6. #148none of 5265.1%

Leader: this model · 32.3%

reasoning

Furniture Assembly

31 results
  1. #8max of 3156.7%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #15max of 2530.0%

Leader: GPT-6 Astra · 53.6%

reasoning

LegalBench

134 results
  1. #9max of 13487.0%

Leader: Fable 5 · 88.6%

reasoning

MysteryMechanism

25 results
  1. #11max of 2533.3%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #36high of 10271.7%
  2. #47low of 10266.0%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #6max of 8052.6%

Leader: Opus 4.7 · 56.1%

reasoning

SimpleQA Verified

47 results
  1. #8max of 4769.2%

Leader: Gemini 3.1 Pro · 77.5%

reliability

Tool-call error rate

171 results
  1. #67max of 1711.0%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #112max of 18199.0%

Leader: GPT-5 Pro · 100.0%

safety

CyberBench v1.1

43 results
  1. #3max of 4376.3%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #6max of 3230.5%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #55none of 3551.0s
  2. #152low of 3552.2s
  3. #234medium of 3555.0s
  4. #302high of 35527s
  5. #313xhigh of 35541s
  6. #337max of 355103s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #178xhigh of 35587.4
  2. #200max of 35580.8
  3. #206none of 35579.7
  4. #214high of 35576.7
  5. #222medium of 35574.4
  6. #226low of 35572.2

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #151max of 1783.4s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #74max of 17771.0

Leader: Llama 3.2 3b · 199.5