- #7max of 3950.0%
Leader: Gemini 3.7 Flash · 60.0%
- #21max of 2221557
- #31xhigh of 2221503
- #41high of 2221465
- #58medium of 2221361
- #96low of 2221116
Leader: Sonnet 5.5 · 1823
- #5max of 4911.7%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench-AA
217 results- #10xhigh of 21766.6%
- #14max of 21764.9%
- #16high of 21764.5%
- #21medium of 21762.6%
- #67low of 21752.6%
Leader: Gemini 4 · 77.5%
- #2max of 2579.6%
Leader: Sonnet 5.5 · 81.1%
agentic
Harvey's Legal Agent Benchmark
76 results- #32max of 765.4%
Leader: Muse Spark 1.2 · 25.4%
agentic
Legal Research Bench
75 results- #32max of 7538.5%
Leader: Muse Spark 1.3 · 55.3%
- #39max of 6721.8%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #6max of 4555.1%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #2max of 3952.9%
Leader: GPT-6 Astra · 62.9%
- #2xhigh of 1738.3%
Leader: GPT-6 Astra · 42.2%
coding
Arena Image-to-WebDev
58 results- #5max of 581703
Leader: Opus 5.5 · 1749
- #4max of 1221755
Leader: Opus 5.5 · 1813
- #5max of 7565.1%
Leader: Sonnet 5.5 · 69.8%
coding
FrontierCode 1.1 Extended
130 results- #25medium of 13060.4%
- #26xhigh of 13060.4%
- #31max of 13059.8%
- #37high of 13059.3%
- #46low of 13058.1%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #16medium of 13050.2%
- #21xhigh of 13049.3%
- #30high of 13048.0%
- #33max of 13047.6%
- #46low of 13045.5%
Leader: Opus 5.5 · 54.6%
- #3max of 4296.9%
Leader: Gemini 4 · 100.0%
coding
ProgramBench v1 Almost Resolved
62 results- #5max of 6246.0%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #6max of 623.0%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #3max of 6285.3%
Leader: Opus 5.5 · 87.0%
- #40high of 22455.8%
- #43xhigh of 22455.7%
- #61max of 22454.2%
- #73medium of 22453.2%
- #75low of 22453.2%
Leader: Opus 5.5 · 66.9%
coding
Terminal-Bench Science-AA
45 results- #4max of 4558.1%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v4-AA
214 results- #9max of 21456.1%
- #11xhigh of 21454.0%
- #16high of 21451.5%
- #19medium of 21448.0%
- #41low of 21430.8%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench v1.1
103 results- #8max of 10388.9%
Leader: Sonnet 5.5 · 92.4%
composite
AA Intelligence Index v4.3.2
607 results- #11max of 60751.8
- #14xhigh of 60751.0
- #17high of 60750.2
- #24medium of 60747.8
- #49low of 60742.1
Leader: Opus 5.5 · 57.6
- #20max of 2201484
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #24max of 2201487
Leader: Gemini 4 · 1553
- #8max of 4561.2%
Leader: Gemini 4 · 68.9%
composite
Vals RSI Index v1.1
24 results- #7max of 2426.6%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #4max of 2420.1%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #13max of 240.0%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #4max of 2443.5%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #5max of 2442.9%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #46low of 177$0.13
- #61medium of 177$0.21
- #72high of 177$0.32
- #83xhigh of 177$0.39
- #102max of 177$0.72
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #48low of 177$250
- #58medium of 177$361
- #73high of 177$521
- #84xhigh of 177$662
- #100max of 177$1,082
Leader: GPT-6 Luna · $10.63
creative
VoxelBench (image)
16 results- #1max of 162177
Leader: this model · 2177
creative
VoxelBench (text)
54 results- #3max of 542552
Leader: GPT-6 Astra · 2654
economics
Excel Modeling Benchmark
72 results- #12max of 7270.8%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #33max of 7652.0%
Leader: Gemini 4 · 65.4%
- #2high of 20832.0%
- #3xhigh of 20831.8%
- #5max of 20831.0%
- #9medium of 20830.0%
- #15low of 20827.0%
Leader: GPT-6 Astra · 32.2%
- #32max of 2861592
- #42xhigh of 2861538
- #46high of 2861510
- #59medium of 2861449
- #94low of 2861317
Leader: Opus 5.5 · 1866
knowledge
AA-Omniscience Accuracy
523 results- #12max of 52362.1%
- #16xhigh of 52360.8%
- #17high of 52360.8%
- #19medium of 52360.4%
- #25low of 52358.9%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #107high of 52350.6%
- #114xhigh of 52349.1%
- #120low of 52348.4%
- #121medium of 52348.4%
- #135max of 52345.7%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Harvey LAB-AA v1.1
25 results- #4max of 256.9%
Leader: Grok 4.7 · 9.4%
- #27max of 9448.8%
Leader: Opus 5 · 63.6%
- #22max of 9686.5%
Leader: Opus 5.5 · 91.4%
- #6max of 26486.0%
- #11xhigh of 26485.1%
- #13high of 26484.9%
- #23medium of 26483.9%
- #27low of 26483.1%
Leader: Opus 5.5 · 87.7%
knowledge
Public Benefits Bench v1.1
48 results- #33max of 4859.3%
Leader: Opus 5 · 76.9%
- #12low of 52084.0%
- #19medium of 52083.3%
- #25max of 52083.0%
- #37high of 52082.3%
- #85xhigh of 52079.7%
Leader: Kimi K3 · 88.7%
- #18max of 9333.9%
Leader: Sonnet 5.5 · 75.0%
math
FrontierMath Tiers 1–3 (v2)
115 results- #1max of 11593.7%
Leader: this model · 93.7%
- #5max of 5099.0%
Leader: Sonnet 5.5 · 100.0%
- #8max of 1161294
Leader: Fable 5 · 1308
- #5max of 855.5%
Leader: Opus 5.5 · 29.1%
- #7xhigh of 21598.5%
- #8high of 21598.5%
- #26max of 21596.5%
- #32medium of 21595.5%
- #44low of 21593.5%
Leader: Fable 5 · 98.5%
- #2max of 21694.2%
- #10xhigh of 21691.7%
- #11high of 21691.7%
- #24medium of 21686.7%
- #43low of 21676.7%
Leader: GPT-6 Astra · 95.0%
- #5max of 52631.7%
- #6xhigh of 52631.7%
- #15high of 52630.0%
- #29medium of 52627.7%
- #43low of 52624.9%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #2max of 3180.0%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #2max of 2546.6%
Leader: GPT-6 Astra · 53.6%
reasoning
MysteryMechanism
25 results- #5max of 2546.4%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #2high of 10288.7%
- #7low of 10283.7%
Leader: GPT-6 Astra · 91.2%
- #35max of 8046.5%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #4max of 4773.8%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #26max of 1710.2%
Leader: GPT-5.4 Pro · 0.0%
- #94max of 18199.5%
Leader: GPT-5 Pro · 100.0%
- #41max of 4339.3%
Leader: GPT-6 Sol · 78.0%
- #2max of 3250.8%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #200low of 3553.0s
- #242medium of 3555.9s
- #327high of 35562s
- #347xhigh of 355174s
- #352max of 355309s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #262max of 35559.3
- #282xhigh of 35554.5
- #293high of 35553.0
- #299medium of 35552.4
- #304low of 35551.3
Leader: Celeris-1 · 1461.1
- #161max of 1784.1s
Leader: Qwen 2.5 7B · 0.2s
- #142max of 17741.3
Leader: Llama 3.2 3b · 199.5