priorsBuilder

OpenAI · released 2026-09-22

GPT-6 Sol

Measured on 82 benchmarks across 6 reasoning efforts.

General frontier

62.0

#9 of 32 models · best at max

Reasoning efforts in your mix

max62.0#31 of 128
xhigh58.1#36 of 128
high56.2#39 of 128
medium54.6#43 of 128
low49.9#49 of 128
none34.1#111 of 128

Your mix’s benchmarks

7 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #32max of 58447.9%
  2. #41xhigh of 58446.3%
  3. #49high of 58444.1%
  4. #75medium of 58441.0%
  5. #117low of 58434.9%
  6. #221none of 58418.4%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #42max of 52327.1
  2. #43medium of 52327.0
  3. #44high of 52326.8
  4. #45xhigh of 52326.7
  5. #46low of 52326.5
  6. #136none of 523-0.8

Leader: Opus 5.5 · 46.4

reasoning

ARC-AGI-3

57 results
  1. #15max of 5723.0%
  2. #18xhigh of 579.5%
  3. #23high of 575.9%
  4. #24medium of 574.9%
  5. #27low of 571.8%
  6. #29none of 571.2%

Leader: GPT-6 Astra · 99.9%

reasoning

BullshitBench v2

192 results
  1. #50none of 19259.0%
  2. #63max of 19254.0%

Leader: Opus 4.8 · 95.0%

math

FrontierMath Tier 4 (v2)

70 results
  1. #8max of 7090.0%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #17max of 9773.1%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #16max of 3549.4%

Leader: Opus 5.5 · 64.8%

All 82 benchmarks

show 75 more

agentic

AA-Briefcase v1.1

222 results
  1. #40max of 2221480
  2. #60xhigh of 2221357
  3. #75high of 2221267
  4. #92medium of 2221149
  5. #99none of 2221103
  6. #135low of 222886

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #6max of 499.9%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #11xhigh of 11533.2%
  2. #13max of 11532.0%
  3. #14high of 11531.2%
  4. #23medium of 11526.9%
  5. #35low of 11521.2%
  6. #77none of 1159.1%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #25xhigh of 21761.7%
  2. #26max of 21761.6%
  3. #30high of 21760.1%
  4. #40medium of 21758.0%
  5. #60low of 21753.9%
  6. #105none of 21734.2%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #7max of 2574.8%

Leader: Sonnet 5.5 · 81.1%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #47max of 761.7%

Leader: Muse Spark 1.2 · 25.4%

agentic

ITBench-AA

45 results
  1. #7max of 4549.4%

Leader: GPT-5.6 Sol · 56.2%

agentic

Legal Research Bench

75 results
  1. #45max of 7528.8%

Leader: Muse Spark 1.3 · 55.3%

agentic

Tax Agent Bench

67 results
  1. #49max of 6715.0%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #8max of 4544.4%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #7max of 3930.0%

Leader: GPT-6 Astra · 62.9%

agentic

Vending-Bench 2

67 results
  1. #2max of 67$14,428

Leader: GPT-6 Astra · $15,515

agentic

WeirdML v3

17 results
  1. #6xhigh of 1719.7%

Leader: GPT-6 Astra · 42.2%

coding

Arena Image-to-WebDev

58 results
  1. #6max of 581661

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #8max of 1221688

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #7max of 7557.2%

Leader: Sonnet 5.5 · 69.8%

coding

FrontierCode 1.1 Extended

130 results
  1. #23max of 13060.7%
  2. #32high of 13059.6%
  3. #38xhigh of 13059.1%
  4. #48medium of 13057.1%
  5. #87low of 13050.5%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #22max of 13049.3%
  2. #25xhigh of 13048.4%
  3. #31high of 13047.7%
  4. #42medium of 13045.9%
  5. #79low of 13037.3%

Leader: Opus 5.5 · 54.6%

coding

IOI

42 results
  1. #10max of 4282.6%

Leader: Gemini 4 · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #9max of 6232.5%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #11max of 622.0%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #7max of 6281.9%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #23max of 22457.6%
  2. #48xhigh of 22455.1%
  3. #54high of 22454.9%
  4. #69medium of 22453.8%
  5. #103low of 22450.2%
  6. #124none of 22447.3%

Leader: Opus 5.5 · 66.9%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #6max of 7483.1%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Science-AA

45 results
  1. #11max of 4530.0%
  2. #17xhigh of 4517.6%
  3. #18high of 4517.1%
  4. #28medium of 458.6%
  5. #34none of 455.7%
  6. #37low of 453.3%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v4-AA

214 results
  1. #23max of 21443.9%
  2. #42xhigh of 21430.3%
  3. #47high of 21426.3%
  4. #61medium of 21418.7%
  5. #73none of 21413.1%
  6. #87low of 2149.1%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench v1.1

103 results
  1. #10max of 10387.8%

Leader: Sonnet 5.5 · 92.4%

composite

AA Intelligence Index v4.3.2

607 results
  1. #25max of 60747.6
  2. #38xhigh of 60744.2
  3. #45high of 60742.4
  4. #60medium of 60739.8
  5. #91low of 60734.2
  6. #126none of 60728.5

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #59max of 2201456

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #55max of 2201467

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #11max of 4557.5%

Leader: Gemini 4 · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #5max of 2428.1%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #1max of 2422.8%

Leader: this model · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #4max of 2416.3%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #7max of 2442.3%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #20max of 2431.1%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #47low of 177$0.13
  2. #64medium of 177$0.25
  3. #75none of 177$0.33
  4. #81high of 177$0.37
  5. #93xhigh of 177$0.52
  6. #119max of 177$1.04

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #50low of 177$269
  2. #62medium of 177$416
  3. #66none of 177$449
  4. #78high of 177$605
  5. #89xhigh of 177$855
  6. #121max of 177$1,536

Leader: GPT-6 Luna · $10.63

creative

VoxelBench (text)

54 results
  1. #4max of 542397

Leader: GPT-6 Astra · 2654

economics

Excel Modeling Benchmark

72 results
  1. #10max of 7271.5%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #44max of 7649.0%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #27max of 20825.2%
  2. #30xhigh of 20824.6%
  3. #32high of 20824.4%
  4. #34low of 20824.2%
  5. #38medium of 20823.8%
  6. #83none of 20817.0%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #47max of 2861508
  2. #55xhigh of 2861455
  3. #71high of 2861394
  4. #85medium of 2861345
  5. #103none of 2861249
  6. #114low of 2861200

Leader: Opus 5.5 · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #40max of 52354.5%
  2. #43xhigh of 52353.8%
  3. #44high of 52353.7%
  4. #46medium of 52353.4%
  5. #57low of 52351.2%
  6. #89none of 52345.2%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #112low of 52349.3%
  2. #141medium of 52343.2%
  3. #145high of 52341.9%
  4. #148xhigh of 52341.1%
  5. #153max of 52339.9%
  6. #318none of 52316.0%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

Harvey LAB-AA v1.1

25 results
  1. #8max of 253.6%

Leader: Grok 4.7 · 9.4%

knowledge

MedCode

94 results
  1. #34max of 9447.1%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #47max of 9682.0%

Leader: Opus 5.5 · 91.4%

knowledge

MMMU-Pro

264 results
  1. #28max of 26482.9%
  2. #31xhigh of 26482.5%
  3. #35high of 26482.0%
  4. #46medium of 26480.4%
  5. #48low of 26480.2%
  6. #140none of 26470.3%

Leader: Opus 5.5 · 87.7%

knowledge

Public Benefits Bench v1.1

48 results
  1. #38max of 4856.6%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #16max of 52083.7%
  2. #17high of 52083.7%
  3. #41medium of 52082.3%
  4. #52xhigh of 52081.3%
  5. #94low of 52079.3%
  6. #248none of 52064.0%

Leader: Kimi K3 · 88.7%

long-context

MLCR-AA

93 results
  1. #41max of 9316.1%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #5max of 11589.8%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #13max of 5083.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #36max of 1161270

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #7max of 852.6%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #31max of 21595.5%
  2. #47xhigh of 21592.7%
  3. #62high of 21591.0%
  4. #102medium of 21583.7%
  5. #121low of 21572.2%
  6. #178none of 21529.3%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #16max of 21689.6%
  2. #39xhigh of 21678.1%
  3. #54high of 21668.9%
  4. #86medium of 21657.8%
  5. #114low of 21631.5%
  6. #181none of 2161.7%

Leader: GPT-6 Astra · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #8max of 2836.9%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #12max of 52630.9%
  2. #26xhigh of 52628.0%
  3. #40high of 52625.4%
  4. #45medium of 52624.6%
  5. #82low of 52616.3%
  6. #162none of 5264.0%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #7max of 3158.3%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #12max of 2531.2%

Leader: GPT-6 Astra · 53.6%

reasoning

MysteryMechanism

25 results
  1. #13max of 2530.2%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #16high of 10277.7%
  2. #32low of 10272.6%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #43max of 8044.8%

Leader: Opus 4.7 · 56.1%

reasoning

SimpleQA Verified

47 results
  1. #9max of 4764.4%

Leader: Gemini 3.1 Pro · 77.5%

reliability

Tool-call error rate

171 results
  1. #16max of 1710.1%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #58max of 18199.9%

Leader: GPT-5 Pro · 100.0%

safety

CWE-Bench

18 results
  1. #3max of 1864.7%

Leader: Grok 4.7 · 68.3%

safety

CyberBench v1.1

43 results
  1. #1max of 4378.0%

Leader: this model · 78.0%

safety

SRE Bench

32 results
  1. #7max of 3230.5%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #50none of 3550.9s
  2. #119low of 3551.7s
  3. #300high of 35526s
  4. #321xhigh of 35551s
  5. #343max of 355140s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #150xhigh of 35598.2
  2. #151max of 35598.1
  3. #169low of 35591.5
  4. #172high of 35589.1
  5. #177none of 35587.6

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #140max of 1782.8s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #95max of 17757.0

Leader: Llama 3.2 3b · 199.5