priorsBuilder

OpenAI · released 2026-09-29

GPT-6.1 Sol

Measured on 78 benchmarks across 5 reasoning efforts.

General frontier

72.8

#3 of 32 models · best at max

Reasoning efforts in your mix

max72.8#7 of 128
xhigh72.3#10 of 128
high72.3#11 of 128
medium70.2#15 of 128
low65.9#21 of 128

Your mix’s benchmarks

7 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #17max of 58452.9%
  2. #20xhigh of 58452.6%
  3. #21high of 58451.4%
  4. #24medium of 58449.9%
  5. #36low of 58447.4%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #11max of 52341.5
  2. #12high of 52341.5
  3. #13xhigh of 52340.9
  4. #18medium of 52340.0
  5. #20low of 52337.6

Leader: Opus 5.5 · 46.4

reasoning

ARC-AGI-3

57 results
  1. #7xhigh of 5796.4%
  2. #8max of 5796.2%
  3. #9high of 5795.0%
  4. #10medium of 5791.0%
  5. #11low of 5782.8%

Leader: GPT-6 Astra · 99.9%

reasoning

BullshitBench v2

192 results
  1. #34low of 19265.0%
  2. #47max of 19259.0%

Leader: Opus 4.8 · 95.0%

math

FrontierMath Tier 4 (v2)

70 results
  1. #1max of 70100.0%

Leader: this model · 100.0%

reasoning

SimpleBench

97 results
  1. #5max of 9782.9%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #4max of 3558.2%

Leader: Opus 5.5 · 64.8%

All 78 benchmarks

show 71 more

agentic

AA-AnalystAgent

39 results
  1. #7max of 3950.0%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #21max of 2221557
  2. #31xhigh of 2221503
  3. #41high of 2221465
  4. #58medium of 2221361
  5. #96low of 2221116

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #5max of 4911.7%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench-AA

217 results
  1. #10xhigh of 21766.6%
  2. #14max of 21764.9%
  3. #16high of 21764.5%
  4. #21medium of 21762.6%
  5. #67low of 21752.6%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #2max of 2579.6%

Leader: Sonnet 5.5 · 81.1%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #32max of 765.4%

Leader: Muse Spark 1.2 · 25.4%

agentic

Legal Research Bench

75 results
  1. #32max of 7538.5%

Leader: Muse Spark 1.3 · 55.3%

agentic

Tax Agent Bench

67 results
  1. #39max of 6721.8%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #6max of 4555.1%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #2max of 3952.9%

Leader: GPT-6 Astra · 62.9%

agentic

WeirdML v3

17 results
  1. #2xhigh of 1738.3%

Leader: GPT-6 Astra · 42.2%

coding

Arena Image-to-WebDev

58 results
  1. #5max of 581703

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #4max of 1221755

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #5max of 7565.1%

Leader: Sonnet 5.5 · 69.8%

coding

FrontierCode 1.1 Extended

130 results
  1. #25medium of 13060.4%
  2. #26xhigh of 13060.4%
  3. #31max of 13059.8%
  4. #37high of 13059.3%
  5. #46low of 13058.1%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #16medium of 13050.2%
  2. #21xhigh of 13049.3%
  3. #30high of 13048.0%
  4. #33max of 13047.6%
  5. #46low of 13045.5%

Leader: Opus 5.5 · 54.6%

coding

IOI

42 results
  1. #3max of 4296.9%

Leader: Gemini 4 · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #5max of 6246.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #6max of 623.0%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #3max of 6285.3%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #40high of 22455.8%
  2. #43xhigh of 22455.7%
  3. #61max of 22454.2%
  4. #73medium of 22453.2%
  5. #75low of 22453.2%

Leader: Opus 5.5 · 66.9%

coding

Terminal-Bench Science-AA

45 results
  1. #4max of 4558.1%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v4-AA

214 results
  1. #9max of 21456.1%
  2. #11xhigh of 21454.0%
  3. #16high of 21451.5%
  4. #19medium of 21448.0%
  5. #41low of 21430.8%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench v1.1

103 results
  1. #8max of 10388.9%

Leader: Sonnet 5.5 · 92.4%

composite

AA Intelligence Index v4.3.2

607 results
  1. #11max of 60751.8
  2. #14xhigh of 60751.0
  3. #17high of 60750.2
  4. #24medium of 60747.8
  5. #49low of 60742.1

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #20max of 2201484

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #24max of 2201487

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #8max of 4561.2%

Leader: Gemini 4 · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #7max of 2426.6%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #4max of 2420.1%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #13max of 240.0%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #4max of 2443.5%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #5max of 2442.9%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #46low of 177$0.13
  2. #61medium of 177$0.21
  3. #72high of 177$0.32
  4. #83xhigh of 177$0.39
  5. #102max of 177$0.72

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #48low of 177$250
  2. #58medium of 177$361
  3. #73high of 177$521
  4. #84xhigh of 177$662
  5. #100max of 177$1,082

Leader: GPT-6 Luna · $10.63

creative

VoxelBench (image)

16 results
  1. #1max of 162177

Leader: this model · 2177

creative

VoxelBench (text)

54 results
  1. #3max of 542552

Leader: GPT-6 Astra · 2654

economics

Excel Modeling Benchmark

72 results
  1. #12max of 7270.8%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #33max of 7652.0%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #2high of 20832.0%
  2. #3xhigh of 20831.8%
  3. #5max of 20831.0%
  4. #9medium of 20830.0%
  5. #15low of 20827.0%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #32max of 2861592
  2. #42xhigh of 2861538
  3. #46high of 2861510
  4. #59medium of 2861449
  5. #94low of 2861317

Leader: Opus 5.5 · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #12max of 52362.1%
  2. #16xhigh of 52360.8%
  3. #17high of 52360.8%
  4. #19medium of 52360.4%
  5. #25low of 52358.9%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #107high of 52350.6%
  2. #114xhigh of 52349.1%
  3. #120low of 52348.4%
  4. #121medium of 52348.4%
  5. #135max of 52345.7%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

Harvey LAB-AA v1.1

25 results
  1. #4max of 256.9%

Leader: Grok 4.7 · 9.4%

knowledge

MedCode

94 results
  1. #27max of 9448.8%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #22max of 9686.5%

Leader: Opus 5.5 · 91.4%

knowledge

MMMU-Pro

264 results
  1. #6max of 26486.0%
  2. #11xhigh of 26485.1%
  3. #13high of 26484.9%
  4. #23medium of 26483.9%
  5. #27low of 26483.1%

Leader: Opus 5.5 · 87.7%

knowledge

Public Benefits Bench v1.1

48 results
  1. #33max of 4859.3%

Leader: Opus 5 · 76.9%

long-context

AA-LCR v1.1

520 results
  1. #12low of 52084.0%
  2. #19medium of 52083.3%
  3. #25max of 52083.0%
  4. #37high of 52082.3%
  5. #85xhigh of 52079.7%

Leader: Kimi K3 · 88.7%

long-context

MLCR-AA

93 results
  1. #18max of 9333.9%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #1max of 11593.7%

Leader: this model · 93.7%

math

ProofBench v1.1

50 results
  1. #5max of 5099.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #8max of 1161294

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #5max of 855.5%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #7xhigh of 21598.5%
  2. #8high of 21598.5%
  3. #26max of 21596.5%
  4. #32medium of 21595.5%
  5. #44low of 21593.5%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #2max of 21694.2%
  2. #10xhigh of 21691.7%
  3. #11high of 21691.7%
  4. #24medium of 21686.7%
  5. #43low of 21676.7%

Leader: GPT-6 Astra · 95.0%

reasoning

CritPt

526 results
  1. #5max of 52631.7%
  2. #6xhigh of 52631.7%
  3. #15high of 52630.0%
  4. #29medium of 52627.7%
  5. #43low of 52624.9%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #2max of 3180.0%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #2max of 2546.6%

Leader: GPT-6 Astra · 53.6%

reasoning

MysteryMechanism

25 results
  1. #5max of 2546.4%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #2high of 10288.7%
  2. #7low of 10283.7%

Leader: GPT-6 Astra · 91.2%

reasoning

SAGE

80 results
  1. #35max of 8046.5%

Leader: Opus 4.7 · 56.1%

reasoning

SimpleQA Verified

47 results
  1. #4max of 4773.8%

Leader: Gemini 3.1 Pro · 77.5%

reliability

Tool-call error rate

171 results
  1. #26max of 1710.2%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #94max of 18199.5%

Leader: GPT-5 Pro · 100.0%

safety

CyberBench v1.1

43 results
  1. #41max of 4339.3%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #2max of 3250.8%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #200low of 3553.0s
  2. #242medium of 3555.9s
  3. #327high of 35562s
  4. #347xhigh of 355174s
  5. #352max of 355309s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #262max of 35559.3
  2. #282xhigh of 35554.5
  3. #293high of 35553.0
  4. #299medium of 35552.4
  5. #304low of 35551.3

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #161max of 1784.1s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #142max of 17741.3

Leader: Llama 3.2 3b · 199.5