- #10max of 3947.5%
Leader: Gemini 3.7 Flash · 60.0%
- #39max of 2221480
- #46xhigh of 2221438
- #57high of 2221370
- #80medium of 2221242
- #109low of 2221044
- #115none of 2221008
Leader: Sonnet 5.5 · 1823
- #12xhigh of 496.2%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #21max of 11528.8%
- #24xhigh of 11526.3%
- #27high of 11524.8%
- #40medium of 11519.6%
- #67low of 11511.7%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #31max of 21760.1%
- #53high of 21755.3%
- #54xhigh of 21755.3%
- #71medium of 21751.3%
- #89low of 21741.0%
- #125none of 21722.7%
Leader: Gemini 4 · 77.5%
- #10max of 2571.1%
Leader: Sonnet 5.5 · 81.1%
- #5max of 88.3%
Leader: GPT-6 Astra · 19.2%
agentic
EnterpriseOps-Gym-AA
50 results- #20max of 5042.9%
Leader: Fable 5 · 51.1%
agentic
Harvey's Legal Agent Benchmark
76 results- #43max of 762.5%
Leader: Muse Spark 1.2 · 25.4%
- #1max of 4556.2%
Leader: this model · 56.2%
agentic
Legal Research Bench
75 results- #9max of 7548.1%
Leader: Muse Spark 1.3 · 55.3%
- #4max of 627.3%
Leader: Opus 5 · 31.4%
agentic
PostTrainBench v1.1
12 results- #2max of 1236.2%
Leader: Fable 5 · 41.8%
- #14max of 3454.1%
Leader: DeepSeek V4.1 Flash · 69.8%
- #20max of 6731.1%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #11max of 4537.9%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #9max of 3920.0%
Leader: GPT-6 Astra · 62.9%
agentic
Time Horizon Index: KSP
13 results- #5max of 1323.8%
Leader: Opus 5.5 · 91.3%
- #8max of 67$9,619
Leader: GPT-6 Astra · $15,515
agentic
Vending-Bench Arena
12 results- #5max of 12$7,400
Leader: GPT-6 Astra · $12,400
- #2max of 445.2%
Leader: Fable 5 · 48.5%
- #9xhigh of 1715.0%
Leader: GPT-6 Astra · 42.2%
- #95max of 39285.1%
- #98xhigh of 39284.8%
- #110high of 39283.3%
- #119medium of 39281.0%
- #135low of 39276.0%
Leader: GLM 5.2 · 99.1%
- #17max of 20644.3%
- #39xhigh of 20638.1%
- #42high of 20636.7%
- #44medium of 20636.5%
- #70low of 20629.1%
- #101none of 20619.6%
Leader: Qwen 3.8 Max · 51.3%
coding
Arena Image-to-WebDev
58 results- #12xhigh of 581608
Leader: Opus 5.5 · 1749
- #23xhigh of 1221618
Leader: Opus 5.5 · 1813
- #10max of 7552.9%
Leader: Sonnet 5.5 · 69.8%
coding
FrontierCode 1.1 Extended
130 results- #24max of 13060.6%
- #30xhigh of 13060.0%
- #41high of 13058.7%
- #68medium of 13054.7%
- #91low of 13050.0%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #34max of 13047.5%
- #37xhigh of 13046.8%
- #48high of 13045.1%
- #71medium of 13039.9%
- #88low of 13035.4%
Leader: Opus 5.5 · 54.6%
- #9max of 2132.2%
Leader: GPT-6 Astra · 65.5%
- #5max of 4291.2%
Leader: Gemini 4 · 100.0%
coding
LiveCodeBench (Vals)
127 results- #52max of 12782.6%
Leader: Fable 5.1 · 90.5%
coding
ProgramBench v1 Almost Resolved
62 results- #10max of 6223.0%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #12max of 621.5%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #8max of 6277.6%
Leader: Opus 5.5 · 87.0%
- #22high of 22457.8%
- #26medium of 22457.4%
- #29max of 22457.1%
- #30xhigh of 22457.1%
- #37low of 22456.4%
- #122none of 22447.7%
Leader: Opus 5.5 · 66.9%
- #7xhigh of 2246.0%
Leader: Opus 5 · 63.2%
coding
SWE-bench Verified (Vals)
85 results- #3max of 8596.2%
Leader: Opus 5 · 97.0%
coding
Terminal-Bench 2.1 (Vals)
74 results- #3max of 7485.8%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Hard-AA
386 results- #1max of 38665.9%
- #3medium of 38662.9%
- #5high of 38662.1%
- #6xhigh of 38661.4%
- #8low of 38660.6%
Leader: this model · 65.9%
coding
Terminal-Bench Science-AA
45 results- #14max of 4522.4%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v2.1-AA
238 results- #6xhigh of 23889.5%
- #14max of 23888.0%
- #20high of 23887.3%
- #23medium of 23886.1%
- #58low of 23876.8%
- #66none of 23874.2%
Leader: Fable 5.1 · 91.4%
coding
Terminal-Bench v4-AA
214 results- #29max of 21439.9%
- #52xhigh of 21424.7%
- #57high of 21420.7%
- #66medium of 21414.6%
- #122low of 2141.0%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench 1-100
22 results- #8max of 2220.0%
Leader: Opus 5.5 · 30.4%
coding
Vibe Code Bench v1.1
103 results- #22max of 10380.5%
Leader: Sonnet 5.5 · 92.4%
- #10max of 12788.8%
- #12xhigh of 12787.0%
Leader: GPT-6 Astra Pro · 93.6%
composite
AA Intelligence Index v4.3.2
607 results- #26max of 60747.0
- #40xhigh of 60744.0
- #46high of 60742.3
- #66medium of 60739.2
- #100low of 60733.5
- #129none of 60728.3
Leader: Opus 5.5 · 57.6
- #19xhigh of 2201485
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #20xhigh of 2201488
Leader: Gemini 4 · 1553
- #10max of 4558.0%
Leader: Gemini 4 · 68.9%
composite
Vals Multimodal Index v1.2
33 results- #4max of 3372.6%
Leader: Fable 5 · 74.2%
composite
Vals RSI Index v1.1
24 results- #11max of 2423.9%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #6max of 2418.5%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #20max of 240.0%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #8max of 2441.0%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #18max of 2436.1%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #66low of 177$0.26
- #90medium of 177$0.50
- #105high of 177$0.81
- #127xhigh of 177$1.18
- #146max of 177$1.99
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #81low of 177$637
- #96medium of 177$997
- #117high of 177$1,487
- #136xhigh of 177$2,082
- #155max of 177$3,465
Leader: GPT-6 Luna · $10.63
creative
VoxelBench (text)
54 results- #6max of 542166
Leader: GPT-6 Astra · 2654
- #39max of 11764.4%
Leader: Opus 5 · 73.2%
economics
Excel Modeling Benchmark
72 results- #7max of 7272.3%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #27max of 7653.8%
Leader: Gemini 4 · 65.4%
- #12high of 20827.8%
- #13xhigh of 20827.6%
- #14max of 20827.2%
- #23medium of 20826.2%
- #55low of 20821.0%
- #91none of 20815.2%
Leader: GPT-6 Astra · 32.2%
- #28max of 2861609
- #37xhigh of 2861569
- #48high of 2861504
- #66medium of 2861421
- #96low of 2861304
- #107none of 2861242
Leader: Opus 5.5 · 1866
economics
MortgageTax (Vals)
84 results- #27max of 8467.3%
Leader: Opus 5 · 72.1%
- #30max of 12574.8%
Leader: Muse Spark 1.2 · 80.4%
- #46max of 39872.7%
- #61xhigh of 39871.0%
- #72medium of 39869.6%
- #73high of 39869.2%
- #93low of 39866.5%
Leader: Grok 4.3 · 83.3%
knowledge
AA-Omniscience Accuracy
523 results- #23max of 52359.4%
- #26xhigh of 52358.8%
- #27high of 52358.4%
- #29medium of 52357.8%
- #30low of 52357.2%
- #68none of 52348.7%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #388low of 52310.6%
- #410medium of 5239.2%
- #423high of 5238.8%
- #442xhigh of 5238.1%
- #449max of 5237.8%
- #461none of 5237.2%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
- #1xhigh of 281257
Leader: this model · 1257
- #7max of 53894.1%
- #25xhigh of 53893.1%
- #29high of 53892.8%
- #33medium of 53892.6%
- #68low of 53889.8%
- #200none of 53879.0%
Leader: GPT-6 Astra · 96.3%
knowledge
GPQA Diamond (Vals)
122 results- #2max of 12295.2%
Leader: Gemini 3.1 Pro · 95.5%
- #40max of 9444.0%
Leader: Opus 5 · 63.6%
- #29max of 9685.2%
Leader: Opus 5.5 · 91.4%
- #16max of 12289.1%
Leader: Fable 5.1 · 92.4%
- #6max of 8188.8%
Leader: Fable 5.1 · 90.6%
- #25max of 26483.4%
- #30xhigh of 26482.7%
- #36high of 26481.8%
- #38medium of 26481.4%
- #41low of 26481.0%
- #130none of 26471.9%
Leader: Opus 5.5 · 87.7%
knowledge
Public Benefits Bench v1.1
48 results- #16max of 4866.5%
Leader: Opus 5 · 76.9%
- #11max of 52084.0%
- #40xhigh of 52082.3%
- #48high of 52081.7%
- #70medium of 52080.3%
- #116low of 52078.0%
- #258none of 52062.3%
Leader: Kimi K3 · 88.7%
long-context
Arena Document
43 results- #10xhigh of 431483
Leader: Opus 5 · 1516
- #22max of 9326.1%
- #34xhigh of 9318.3%
- #38high of 9317.2%
- #47medium of 9315.0%
- #54low of 9312.8%
Leader: Sonnet 5.5 · 75.0%
- #7max of 70.0%
Leader: GPT-6 Astra · 2.9%
math
FrontierMath Tiers 1–3 (v2)
115 results- #6max of 11589.1%
Leader: GPT-6.1 Sol · 93.7%
- #14max of 5083.0%
Leader: Sonnet 5.5 · 100.0%
- #15xhigh of 1161287
Leader: Fable 5 · 1308
- #6max of 855.4%
Leader: Opus 5.5 · 29.1%
- #10xhigh of 21597.5%
- #20high of 21597.0%
- #22max of 21596.5%
- #50medium of 21592.5%
- #117low of 21574.5%
Leader: Fable 5 · 98.5%
- #5max of 21692.5%
- #13xhigh of 21690.0%
- #26high of 21685.4%
- #61medium of 21667.1%
- #101low of 21642.5%
Leader: GPT-6 Astra · 95.0%
reasoning
Blueprint-Bench 2
28 results- #11max of 2833.6%
Leader: Gemini 4 · 54.4%
- #1max of 52632.3%
- #23xhigh of 52628.6%
- #38high of 52625.7%
- #49medium of 52622.9%
- #89low of 52614.9%
- #148none of 5265.1%
Leader: this model · 32.3%
reasoning
Furniture Assembly
31 results- #8max of 3156.7%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #15max of 2530.0%
Leader: GPT-6 Astra · 53.6%
- #9max of 13487.0%
Leader: Fable 5 · 88.6%
reasoning
MysteryMechanism
25 results- #11max of 2533.3%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #36high of 10271.7%
- #47low of 10266.0%
Leader: GPT-6 Astra · 91.2%
- #6max of 8052.6%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #8max of 4769.2%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #67max of 1711.0%
Leader: GPT-5.4 Pro · 0.0%
- #112max of 18199.0%
Leader: GPT-5 Pro · 100.0%
- #3max of 4376.3%
Leader: GPT-6 Sol · 78.0%
- #6max of 3230.5%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #55none of 3551.0s
- #152low of 3552.2s
- #234medium of 3555.0s
- #302high of 35527s
- #313xhigh of 35541s
- #337max of 355103s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #178xhigh of 35587.4
- #200max of 35580.8
- #206none of 35579.7
- #214high of 35576.7
- #222medium of 35574.4
- #226low of 35572.2
Leader: Celeris-1 · 1461.1
- #151max of 1783.4s
Leader: Qwen 2.5 7B · 0.2s
- #74max of 17771.0
Leader: Llama 3.2 3b · 199.5