priorsBuilder

OpenAI · released 2026-09-03

GPT-6 Astra

Measured on 97 benchmarks across 6 reasoning efforts.

General frontier

73.5

#2 of 32 models · best at max

Reasoning efforts in your mix

max73.5#3 of 128
xhigh73.5#4 of 128
high73.0#6 of 128
medium71.7#12 of 128
low67.1#19 of 128
none60.4#33 of 128

Your mix’s benchmarks

8 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #12max of 58454.7%
  2. #13xhigh of 58454.6%
  3. #16high of 58453.1%
  4. #19medium of 58452.7%
  5. #27low of 58449.2%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #2high of 52343.7
  2. #4xhigh of 52343.4
  3. #5max of 52343.4
  4. #10medium of 52342.2
  5. #16low of 52340.5

Leader: Opus 5.5 · 46.4

reasoning

ARC-AGI-3

57 results
  1. #1high of 5799.9%
  2. #2max of 5798.6%
  3. #3xhigh of 5798.4%
  4. #4medium of 5798.4%
  5. #5low of 5798.0%
  6. #6none of 5796.7%

Leader: this model · 99.9%

reasoning

BullshitBench v2

192 results
  1. #28max of 19269.0%
  2. #37low of 19264.0%

Leader: Opus 4.8 · 95.0%

coding

DeepSWE v1.1

70 results
  1. #1xhigh of 7074.1%
  2. #4high of 7073.2%
  3. #5max of 7073.2%
  4. #8medium of 7072.8%
  5. #23low of 7067.0%

Leader: this model · 74.1%

math

FrontierMath Tier 4 (v2)

70 results
  1. #2high of 7097.6%
  2. #3xhigh of 7097.6%
  3. #4max of 7097.6%
  4. #5medium of 7097.6%
  5. #10low of 7087.8%
  6. #11none of 7082.9%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #4max of 9783.6%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #3max of 3558.2%
  2. #6xhigh of 3557.9%
  3. #7high of 3557.9%
  4. #10medium of 3554.2%
  5. #14low of 3550.6%

Leader: Opus 5.5 · 64.8%

All 97 benchmarks

show 89 more

agentic

AA-AnalystAgent

39 results
  1. #6max of 3951.2%

Leader: Gemini 3.7 Flash · 60.0%

agentic

AA-Briefcase v1.1

222 results
  1. #19max of 2221570
  2. #22xhigh of 2221546
  3. #30high of 2221506
  4. #42medium of 2221461
  5. #77low of 2221257

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #2max of 4913.1%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #5max of 11541.4%
  2. #6xhigh of 11539.0%
  3. #7high of 11537.1%
  4. #10medium of 11534.1%
  5. #16low of 11530.3%
  6. #26none of 11525.3%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #5max of 21768.5%
  2. #6xhigh of 21767.2%
  3. #9high of 21766.6%
  4. #15medium of 21764.6%
  5. #37low of 21759.1%

Leader: Gemini 4 · 77.5%

agentic

BioMysteryBench

25 results
  1. #3max of 2579.3%

Leader: Sonnet 5.5 · 81.1%

agentic

CUA-bench

8 results
  1. #1max of 819.2%

Leader: this model · 19.2%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #31max of 765.4%

Leader: Muse Spark 1.2 · 25.4%

agentic

ITBench-AA

45 results
  1. #8max of 4548.6%

Leader: GPT-5.6 Sol · 56.2%

agentic

Legal Research Bench

75 results
  1. #28max of 7539.4%

Leader: Muse Spark 1.3 · 55.3%

agentic

Tax Agent Bench

67 results
  1. #41max of 6720.7%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 2.1

18 results
  1. #1high of 1887.4%
  2. #2medium of 1887.0%
  3. #3low of 1886.7%
  4. #4max of 1886.7%
  5. #5xhigh of 1885.8%

Leader: this model · 87.4%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #3max of 4559.6%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #1max of 3962.9%

Leader: this model · 62.9%

agentic

Time Horizon Index: KSP

13 results
  1. #2max of 1390.5%

Leader: Opus 5.5 · 91.3%

agentic

Vending-Bench 2

67 results
  1. #1max of 67$15,515

Leader: this model · $15,515

agentic

Vending-Bench Arena

12 results
  1. #1max of 12$12,400

Leader: this model · $12,400

agentic

WeirdML v3

17 results
  1. #1xhigh of 1742.2%

Leader: this model · 42.2%

agentic

τ³-Banking

206 results
  1. #21xhigh of 20643.1%
  2. #26max of 20641.4%
  3. #30high of 20640.0%
  4. #45medium of 20635.5%
  5. #59low of 20632.0%

Leader: Qwen 3.8 Max · 51.3%

coding

Arena Image-to-WebDev

58 results
  1. #3max of 581731

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #2max of 1221786

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #3max of 7567.7%

Leader: Sonnet 5.5 · 69.8%

coding

Drone-Bench

11 results
  1. #2max of 1195.1%

Leader: Opus 5.5 · 97.3%

coding

FrontierCode 1.1 Extended

130 results
  1. #4max of 13064.5%
  2. #12high of 13063.1%
  3. #16xhigh of 13062.1%
  4. #28medium of 13060.3%
  5. #47low of 13057.4%

Leader: Opus 5.5 · 65.3%

coding

FrontierCode 1.1 Main

130 results
  1. #6max of 13053.3%
  2. #11high of 13050.9%
  3. #13xhigh of 13050.6%
  4. #23medium of 13048.8%
  5. #47low of 13045.3%

Leader: Opus 5.5 · 54.6%

coding

FrontierSWE v2

21 results
  1. #1max of 2165.5%

Leader: this model · 65.5%

coding

IOI

42 results
  1. #2max of 42100.0%

Leader: Gemini 4 · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #3max of 6250.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #4max of 625.5%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #2max of 6285.4%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #33max of 22456.5%
  2. #42xhigh of 22455.7%
  3. #46high of 22455.4%
  4. #62medium of 22454.2%
  5. #64low of 22454.1%

Leader: Opus 5.5 · 66.9%

coding

SWE Atlas-QnA

22 results
  1. #3xhigh of 2259.1%

Leader: Opus 5 · 63.2%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #2max of 7487.3%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Science-AA

45 results
  1. #1max of 4563.3%

Leader: this model · 63.3%

coding

Terminal-Bench v2.1-AA

238 results
  1. #4high of 23889.9%
  2. #5medium of 23889.5%
  3. #7xhigh of 23889.1%
  4. #10max of 23888.4%
  5. #15low of 23888.0%

Leader: Fable 5.1 · 91.4%

coding

Terminal-Bench v4-AA

214 results
  1. #4xhigh of 21459.6%
  2. #5max of 21459.1%
  3. #12high of 21454.0%
  4. #17medium of 21449.5%
  5. #26low of 21441.9%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench 1-100

22 results
  1. #4max of 2227.6%

Leader: Opus 5.5 · 30.4%

coding

Vibe Code Bench v1.1

103 results
  1. #7max of 10389.6%

Leader: Sonnet 5.5 · 92.4%

coding

WeirdML v2

127 results
  1. #2xhigh of 12793.3%
  2. #4high of 12792.9%

Leader: GPT-6 Astra Pro · 93.6%

composite

AA Intelligence Index v4.3.2

607 results
  1. #7max of 60752.7
  2. #9xhigh of 60752.4
  3. #15high of 60750.9
  4. #20medium of 60749.6
  5. #32low of 60745.8

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #34max of 2201475

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #28max of 2201484

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #6max of 4563.1%

Leader: Gemini 4 · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #6max of 2427.1%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #18max of 249.5%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #6max of 247.8%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #5max of 2442.6%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #3max of 2448.4%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #106low of 177$0.82
  2. #137medium of 177$1.54
  3. #142high of 177$1.73
  4. #154xhigh of 177$2.31
  5. #162max of 177$3.26

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #122low of 177$1,537
  2. #145medium of 177$2,434
  3. #151high of 177$2,925
  4. #157xhigh of 177$3,803
  5. #168max of 177$5,324

Leader: GPT-6 Luna · $10.63

creative

VoxelBench (text)

54 results
  1. #1max of 542654

Leader: this model · 2654

economics

Excel Modeling Benchmark

72 results
  1. #9max of 7271.7%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #29max of 7653.5%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #1xhigh of 20832.2%
  2. #4max of 20831.0%
  3. #6high of 20831.0%
  4. #7medium of 20830.4%
  5. #8low of 20830.4%

Leader: this model · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #36max of 2861574
  2. #38xhigh of 2861554
  3. #44high of 2861524
  4. #50medium of 2861485
  5. #73low of 2861389

Leader: Opus 5.5 · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #11max of 52362.6%
  2. #13xhigh of 52361.9%
  3. #14high of 52361.1%
  4. #18medium of 52360.6%
  5. #21low of 52359.5%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #90high of 52355.2%
  2. #97medium of 52353.5%
  3. #98low of 52353.1%
  4. #101xhigh of 52351.7%
  5. #117max of 52348.7%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

GPQA Diamond

538 results
  1. #1xhigh of 53896.3%
  2. #2max of 53896.1%
  3. #4high of 53894.9%
  4. #10medium of 53893.9%
  5. #24low of 53893.1%

Leader: this model · 96.3%

knowledge

Harvey LAB-AA v1.1

25 results
  1. #3max of 258.6%

Leader: Grok 4.7 · 9.4%

knowledge

MedCode

94 results
  1. #28max of 9448.5%

Leader: Opus 5 · 63.6%

knowledge

MedScribe

96 results
  1. #15max of 9687.9%

Leader: Opus 5.5 · 91.4%

knowledge

MMMU-Pro

264 results
  1. #2max of 26486.9%
  2. #4high of 26486.4%
  3. #5xhigh of 26486.2%
  4. #12medium of 26485.1%
  5. #18low of 26484.6%

Leader: Opus 5.5 · 87.7%

long-context

AA-LCR v1.1

520 results
  1. #61max of 52080.7%
  2. #73xhigh of 52080.0%
  3. #74high of 52080.0%
  4. #75low of 52080.0%
  5. #86medium of 52079.7%

Leader: Kimi K3 · 88.7%

long-context

Arena Document

43 results
  1. #16max of 431468

Leader: Opus 5 · 1516

long-context

MLCR-AA

93 results
  1. #17max of 9335.0%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Erdős

7 results
  1. #1max of 72.9%

Leader: this model · 2.9%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #2max of 11593.7%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #6max of 5099.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #18max of 1161286

Leader: Fable 5 · 1308

other

T3 Code usage share

85 results
  1. #2max of 8517.0%

Leader: Opus 5.5 · 29.1%

reasoning

ARC-AGI-1

215 results
  1. #3xhigh of 21598.5%
  2. #4high of 21598.5%
  3. #14max of 21597.5%
  4. #15medium of 21597.5%
  5. #25low of 21596.5%
  6. #95none of 21586.0%

Leader: Fable 5 · 98.5%

reasoning

ARC-AGI-2

216 results
  1. #1max of 21695.0%
  2. #3xhigh of 21693.3%
  3. #7high of 21692.1%
  4. #8medium of 21692.1%
  5. #27low of 21685.4%
  6. #81none of 21659.6%

Leader: this model · 95.0%

reasoning

Blueprint-Bench 2

28 results
  1. #3max of 2849.7%

Leader: Gemini 4 · 54.4%

reasoning

CritPt

526 results
  1. #4max of 52631.7%
  2. #8xhigh of 52631.4%
  3. #20medium of 52629.1%
  4. #22high of 52628.9%
  5. #36low of 52626.3%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Furniture Assembly

31 results
  1. #3max of 3180.0%

Leader: Opus 5.5 · 83.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #1max of 2553.6%

Leader: this model · 53.6%

reasoning

MysteryMechanism

25 results
  1. #1max of 2553.2%

Leader: this model · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #1high of 10291.2%
  2. #3low of 10287.2%

Leader: this model · 91.2%

reasoning

SAGE

80 results
  1. #36max of 8046.4%

Leader: Opus 4.7 · 56.1%

reasoning

SimpleQA Verified

47 results
  1. #2max of 4775.8%

Leader: Gemini 3.1 Pro · 77.5%

reliability

Tool-call error rate

171 results
  1. #13max of 1710.1%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #93max of 18199.5%

Leader: GPT-5 Pro · 100.0%

safety

CWE-Bench

18 results
  1. #4max of 1863.3%

Leader: Grok 4.7 · 68.3%

safety

CyberBench v1.1

43 results
  1. #40max of 4341.1%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #1max of 3256.9%

Leader: this model · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #177low of 3552.6s
  2. #250medium of 3558.3s
  3. #331high of 35588s
  4. #348xhigh of 355207s
  5. #353max of 355388s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #303high of 35551.4
  2. #310max of 35547.7
  3. #313xhigh of 35546.1
  4. #320low of 35542.8
  5. #323medium of 35541.9

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #168max of 1784.9s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #131max of 17746.0

Leader: Llama 3.2 3b · 199.5