- #40max of 2221480
- #60xhigh of 2221357
- #75high of 2221267
- #92medium of 2221149
- #99none of 2221103
- #135low of 222886
Leader: Sonnet 5.5 · 1823
- #6max of 499.9%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #11xhigh of 11533.2%
- #13max of 11532.0%
- #14high of 11531.2%
- #23medium of 11526.9%
- #35low of 11521.2%
- #77none of 1159.1%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #25xhigh of 21761.7%
- #26max of 21761.6%
- #30high of 21760.1%
- #40medium of 21758.0%
- #60low of 21753.9%
- #105none of 21734.2%
Leader: Gemini 4 · 77.5%
- #7max of 2574.8%
Leader: Sonnet 5.5 · 81.1%
agentic
Harvey's Legal Agent Benchmark
76 results- #47max of 761.7%
Leader: Muse Spark 1.2 · 25.4%
- #7max of 4549.4%
Leader: GPT-5.6 Sol · 56.2%
agentic
Legal Research Bench
75 results- #45max of 7528.8%
Leader: Muse Spark 1.3 · 55.3%
- #49max of 6715.0%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #8max of 4544.4%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #7max of 3930.0%
Leader: GPT-6 Astra · 62.9%
- #2max of 67$14,428
Leader: GPT-6 Astra · $15,515
- #6xhigh of 1719.7%
Leader: GPT-6 Astra · 42.2%
coding
Arena Image-to-WebDev
58 results- #6max of 581661
Leader: Opus 5.5 · 1749
- #8max of 1221688
Leader: Opus 5.5 · 1813
- #7max of 7557.2%
Leader: Sonnet 5.5 · 69.8%
coding
FrontierCode 1.1 Extended
130 results- #23max of 13060.7%
- #32high of 13059.6%
- #38xhigh of 13059.1%
- #48medium of 13057.1%
- #87low of 13050.5%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #22max of 13049.3%
- #25xhigh of 13048.4%
- #31high of 13047.7%
- #42medium of 13045.9%
- #79low of 13037.3%
Leader: Opus 5.5 · 54.6%
- #10max of 4282.6%
Leader: Gemini 4 · 100.0%
coding
ProgramBench v1 Almost Resolved
62 results- #9max of 6232.5%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #11max of 622.0%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #7max of 6281.9%
Leader: Opus 5.5 · 87.0%
- #23max of 22457.6%
- #48xhigh of 22455.1%
- #54high of 22454.9%
- #69medium of 22453.8%
- #103low of 22450.2%
- #124none of 22447.3%
Leader: Opus 5.5 · 66.9%
coding
Terminal-Bench 2.1 (Vals)
74 results- #6max of 7483.1%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Science-AA
45 results- #11max of 4530.0%
- #17xhigh of 4517.6%
- #18high of 4517.1%
- #28medium of 458.6%
- #34none of 455.7%
- #37low of 453.3%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v4-AA
214 results- #23max of 21443.9%
- #42xhigh of 21430.3%
- #47high of 21426.3%
- #61medium of 21418.7%
- #73none of 21413.1%
- #87low of 2149.1%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench v1.1
103 results- #10max of 10387.8%
Leader: Sonnet 5.5 · 92.4%
composite
AA Intelligence Index v4.3.2
607 results- #25max of 60747.6
- #38xhigh of 60744.2
- #45high of 60742.4
- #60medium of 60739.8
- #91low of 60734.2
- #126none of 60728.5
Leader: Opus 5.5 · 57.6
- #59max of 2201456
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #55max of 2201467
Leader: Gemini 4 · 1553
- #11max of 4557.5%
Leader: Gemini 4 · 68.9%
composite
Vals RSI Index v1.1
24 results- #5max of 2428.1%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #1max of 2422.8%
Leader: this model · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #4max of 2416.3%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #7max of 2442.3%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #20max of 2431.1%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #47low of 177$0.13
- #64medium of 177$0.25
- #75none of 177$0.33
- #81high of 177$0.37
- #93xhigh of 177$0.52
- #119max of 177$1.04
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #50low of 177$269
- #62medium of 177$416
- #66none of 177$449
- #78high of 177$605
- #89xhigh of 177$855
- #121max of 177$1,536
Leader: GPT-6 Luna · $10.63
creative
VoxelBench (text)
54 results- #4max of 542397
Leader: GPT-6 Astra · 2654
economics
Excel Modeling Benchmark
72 results- #10max of 7271.5%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #44max of 7649.0%
Leader: Gemini 4 · 65.4%
- #27max of 20825.2%
- #30xhigh of 20824.6%
- #32high of 20824.4%
- #34low of 20824.2%
- #38medium of 20823.8%
- #83none of 20817.0%
Leader: GPT-6 Astra · 32.2%
- #47max of 2861508
- #55xhigh of 2861455
- #71high of 2861394
- #85medium of 2861345
- #103none of 2861249
- #114low of 2861200
Leader: Opus 5.5 · 1866
knowledge
AA-Omniscience Accuracy
523 results- #40max of 52354.5%
- #43xhigh of 52353.8%
- #44high of 52353.7%
- #46medium of 52353.4%
- #57low of 52351.2%
- #89none of 52345.2%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #112low of 52349.3%
- #141medium of 52343.2%
- #145high of 52341.9%
- #148xhigh of 52341.1%
- #153max of 52339.9%
- #318none of 52316.0%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Harvey LAB-AA v1.1
25 results- #8max of 253.6%
Leader: Grok 4.7 · 9.4%
- #34max of 9447.1%
Leader: Opus 5 · 63.6%
- #47max of 9682.0%
Leader: Opus 5.5 · 91.4%
- #28max of 26482.9%
- #31xhigh of 26482.5%
- #35high of 26482.0%
- #46medium of 26480.4%
- #48low of 26480.2%
- #140none of 26470.3%
Leader: Opus 5.5 · 87.7%
knowledge
Public Benefits Bench v1.1
48 results- #38max of 4856.6%
Leader: Opus 5 · 76.9%
- #16max of 52083.7%
- #17high of 52083.7%
- #41medium of 52082.3%
- #52xhigh of 52081.3%
- #94low of 52079.3%
- #248none of 52064.0%
Leader: Kimi K3 · 88.7%
- #41max of 9316.1%
Leader: Sonnet 5.5 · 75.0%
math
FrontierMath Tiers 1–3 (v2)
115 results- #5max of 11589.8%
Leader: GPT-6.1 Sol · 93.7%
- #13max of 5083.0%
Leader: Sonnet 5.5 · 100.0%
- #36max of 1161270
Leader: Fable 5 · 1308
- #7max of 852.6%
Leader: Opus 5.5 · 29.1%
- #31max of 21595.5%
- #47xhigh of 21592.7%
- #62high of 21591.0%
- #102medium of 21583.7%
- #121low of 21572.2%
- #178none of 21529.3%
Leader: Fable 5 · 98.5%
- #16max of 21689.6%
- #39xhigh of 21678.1%
- #54high of 21668.9%
- #86medium of 21657.8%
- #114low of 21631.5%
- #181none of 2161.7%
Leader: GPT-6 Astra · 95.0%
reasoning
Blueprint-Bench 2
28 results- #8max of 2836.9%
Leader: Gemini 4 · 54.4%
- #12max of 52630.9%
- #26xhigh of 52628.0%
- #40high of 52625.4%
- #45medium of 52624.6%
- #82low of 52616.3%
- #162none of 5264.0%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #7max of 3158.3%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #12max of 2531.2%
Leader: GPT-6 Astra · 53.6%
reasoning
MysteryMechanism
25 results- #13max of 2530.2%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #16high of 10277.7%
- #32low of 10272.6%
Leader: GPT-6 Astra · 91.2%
- #43max of 8044.8%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #9max of 4764.4%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #16max of 1710.1%
Leader: GPT-5.4 Pro · 0.0%
- #58max of 18199.9%
Leader: GPT-5 Pro · 100.0%
- #3max of 1864.7%
Leader: Grok 4.7 · 68.3%
- #1max of 4378.0%
Leader: this model · 78.0%
- #7max of 3230.5%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #50none of 3550.9s
- #119low of 3551.7s
- #300high of 35526s
- #321xhigh of 35551s
- #343max of 355140s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #150xhigh of 35598.2
- #151max of 35598.1
- #169low of 35591.5
- #172high of 35589.1
- #177none of 35587.6
Leader: Celeris-1 · 1461.1
- #140max of 1782.8s
Leader: Qwen 2.5 7B · 0.2s
- #95max of 17757.0
Leader: Llama 3.2 3b · 199.5