- #5max of 3953.8%
Leader: Gemini 3.7 Flash · 60.0%
- #7max of 2221660
- #11xhigh of 2221635
- #20high of 2221561
- #47medium of 2221435
- #83low of 2221209
Leader: Sonnet 5.5 · 1823
- #9max of 498.1%
- #10high of 498.0%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #22max of 11526.9%
- #25xhigh of 11525.3%
- #29medium of 11523.9%
- #38high of 11520.6%
- #39low of 11520.4%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #48max of 21756.6%
- #57medium of 21754.3%
- #61high of 21753.6%
- #63xhigh of 21753.2%
- #70low of 21751.8%
Leader: Gemini 4 · 77.5%
- #5max of 2579.3%
Leader: Sonnet 5.5 · 81.1%
- #4max of 89.0%
Leader: GPT-6 Astra · 19.2%
agentic
EnterpriseOps-Gym-AA
50 results- #7max of 5047.5%
Leader: Fable 5 · 51.1%
agentic
Harvey's Legal Agent Benchmark
76 results- #29max of 766.7%
Leader: Muse Spark 1.2 · 25.4%
agentic
Legal Research Bench
75 results- #2max of 7555.3%
Leader: Muse Spark 1.3 · 55.3%
- #1max of 631.4%
- #2xhigh of 630.2%
- #3high of 629.0%
- #5medium of 625.3%
- #6low of 622.3%
Leader: this model · 31.4%
- #2max of 1383.4%
Leader: Fable 5 · 86.0%
agentic
PostTrainBench v1.1
12 results- #3max of 1235.0%
Leader: Fable 5 · 41.8%
- #7max of 3460.4%
Leader: DeepSeek V4.1 Flash · 69.8%
- #3max of 6745.5%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #7max of 4553.5%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #8max of 3927.1%
Leader: GPT-6 Astra · 62.9%
agentic
Time Horizon Index: KSP
13 results- #7max of 1318.8%
Leader: Opus 5.5 · 91.3%
- #4max of 67$11,182
Leader: GPT-6 Astra · $15,515
agentic
Vending-Bench Arena
12 results- #6max of 12$7,000
Leader: GPT-6 Astra · $12,400
- #7xhigh of 1719.0%
Leader: GPT-6 Astra · 42.2%
- #16high of 20644.7%
- #19xhigh of 20643.3%
- #23max of 20642.1%
- #36medium of 20638.6%
- #64low of 20630.3%
Leader: Qwen 3.8 Max · 51.3%
coding
Arena Image-to-WebDev
58 results- #7max of 581661
Leader: Opus 5.5 · 1749
- #7max of 1221691
- #11high of 1221657
Leader: Opus 5.5 · 1813
- #6max of 7557.5%
Leader: Sonnet 5.5 · 69.8%
- #4max of 1186.9%
Leader: Opus 5.5 · 97.3%
coding
FrontierCode 1.1 Extended
130 results- #8medium of 13063.6%
- #40max of 13058.9%
- #42high of 13058.5%
- #50xhigh of 13056.9%
- #60low of 13055.8%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #5medium of 13053.4%
- #26max of 13048.0%
- #29high of 13048.0%
- #50xhigh of 13043.6%
- #60low of 13041.9%
Leader: Opus 5.5 · 54.6%
- #6max of 2152.0%
Leader: GPT-6 Astra · 65.5%
- #8max of 4284.3%
Leader: Gemini 4 · 100.0%
coding
LiveCodeBench (Vals)
127 results- #4max of 12789.0%
Leader: Fable 5.1 · 90.5%
coding
ProgramBench v1 Almost Resolved
62 results- #6max of 6241.5%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #5max of 623.0%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #6max of 6282.3%
Leader: Opus 5.5 · 87.0%
- #35max of 22456.4%
- #44xhigh of 22455.7%
- #47high of 22455.4%
- #90medium of 22451.5%
- #112low of 22449.2%
Leader: Opus 5.5 · 66.9%
- #1xhigh of 2263.2%
Leader: this model · 63.2%
coding
SWE-bench Verified (Vals)
85 results- #1max of 8597.0%
Leader: this model · 97.0%
coding
Terminal-Bench 2.1 (Vals)
74 results- #5high of 7484.6%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Science-AA
45 results- #12max of 4528.6%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v2.1-AA
238 results- #8max of 23889.1%
- #12xhigh of 23888.0%
- #18high of 23887.6%
- #21medium of 23886.1%
- #59low of 23876.4%
Leader: Fable 5.1 · 91.4%
coding
Terminal-Bench v4-AA
214 results- #18max of 21449.0%
- #20xhigh of 21446.5%
- #21high of 21446.0%
- #34medium of 21434.3%
- #48low of 21426.3%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench 1-100
22 results- #2max of 2228.5%
Leader: Opus 5.5 · 30.4%
coding
Vibe Code Bench v1.1
103 results- #9max of 10388.4%
Leader: Sonnet 5.5 · 92.4%
- #7xhigh of 12791.8%
- #8high of 12791.6%
- #13max of 12786.3%
Leader: GPT-6 Astra Pro · 93.6%
composite
AA Intelligence Index v4.3.2
607 results- #16max of 60750.8
- #18xhigh of 60749.7
- #22high of 60748.1
- #35medium of 60744.8
- #65low of 60739.4
Leader: Opus 5.5 · 57.6
- #13high of 2201490
- #14max of 2201489
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #25max of 2201486
- #31high of 2201484
Leader: Gemini 4 · 1553
- #5max of 4563.7%
Leader: Gemini 4 · 68.9%
composite
Vals Multimodal Index v1.2
33 results- #2max of 3373.9%
Leader: Fable 5 · 74.2%
composite
Vals RSI Index v1.1
24 results- #3max of 2433.0%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #10max of 2414.8%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #2max of 2431.8%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #6max of 2442.6%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #6max of 2442.8%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #121low of 177$1.10
- #153medium of 177$2.19
- #164high of 177$3.61
- #169xhigh of 177$4.88
- #173max of 177$5.86
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #124low of 177$1,561
- #150medium of 177$2,732
- #162high of 177$4,332
- #169xhigh of 177$5,868
- #173max of 177$7,275
Leader: GPT-6 Luna · $10.63
- #5max of 722051
Leader: GPT-6 Astra Pro · 2278
creative
VoxelBench (text)
54 results- #5max of 542176
Leader: GPT-6 Astra · 2654
- #1max of 11773.2%
Leader: this model · 73.2%
economics
Excel Modeling Benchmark
72 results- #6max of 7273.6%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #8max of 7658.6%
Leader: Gemini 4 · 65.4%
- #50max of 20821.6%
- #53xhigh of 20821.0%
- #62medium of 20820.0%
- #64high of 20819.6%
- #79low of 20817.2%
Leader: GPT-6 Astra · 32.2%
- #7max of 2861722
- #11xhigh of 2861692
- #31high of 2861594
- #49medium of 2861490
- #97low of 2861301
Leader: Opus 5.5 · 1866
economics
MortgageTax (Vals)
84 results- #1max of 8472.1%
Leader: this model · 72.1%
- #23max of 12575.1%
Leader: Muse Spark 1.2 · 80.4%
knowledge
AA-Omniscience Accuracy
523 results- #15max of 52360.9%
- #22xhigh of 52359.5%
- #24high of 52358.9%
- #31medium of 52357.1%
- #34low of 52356.0%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #151xhigh of 52340.5%
- #155medium of 52339.3%
- #156max of 52339.2%
- #160high of 52338.8%
- #167low of 52337.8%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
- #12xhigh of 53893.7%
- #13high of 53893.7%
- #22max of 53893.2%
- #44medium of 53891.9%
- #84low of 53888.9%
Leader: GPT-6 Astra · 96.3%
knowledge
GPQA Diamond (Vals)
122 results- #8max of 12293.4%
Leader: Gemini 3.1 Pro · 95.5%
- #1max of 9463.6%
Leader: this model · 63.6%
- #4max of 9691.0%
Leader: Opus 5.5 · 91.4%
- #2max of 12291.6%
Leader: Fable 5.1 · 92.4%
- #2max of 8189.9%
Leader: Fable 5.1 · 90.6%
- #15max of 26484.7%
- #22xhigh of 26484.0%
- #32high of 26482.4%
- #37medium of 26481.6%
- #53low of 26479.8%
Leader: Opus 5.5 · 87.7%
knowledge
Public Benefits Bench v1.1
48 results- #1max of 4876.9%
Leader: this model · 76.9%
- #43medium of 52082.0%
- #54low of 52081.3%
- #66xhigh of 52080.3%
- #92max of 52079.3%
- #100high of 52079.0%
Leader: Kimi K3 · 88.7%
long-context
Arena Document
43 results- #1high of 431516
Leader: this model · 1516
- #6high of 9359.4%
- #7xhigh of 9358.3%
- #8medium of 9356.1%
- #9max of 9355.6%
- #11low of 9353.9%
Leader: Sonnet 5.5 · 75.0%
math
FrontierMath Tiers 1–3 (v2)
115 results- #11max of 11585.6%
Leader: GPT-6.1 Sol · 93.7%
- #8max of 5099.0%
Leader: Sonnet 5.5 · 100.0%
- #14high of 1161287
Leader: Fable 5 · 1308
- #4max of 859.1%
Leader: Opus 5.5 · 29.1%
- #11high of 21597.5%
- #12max of 21597.5%
Leader: Fable 5 · 98.5%
- #12max of 21690.4%
- #21high of 21688.3%
Leader: GPT-6 Astra · 95.0%
reasoning
Blueprint-Bench 2
28 results- #18max of 2830.4%
Leader: Gemini 4 · 54.4%
- #19max of 52629.1%
- #25high of 52628.3%
- #28xhigh of 52627.7%
- #34medium of 52626.9%
- #48low of 52623.1%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #6max of 3160.8%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #14max of 2530.5%
Leader: GPT-6 Astra · 53.6%
- #8max of 13487.0%
Leader: Fable 5 · 88.6%
reasoning
MysteryMechanism
25 results- #7max of 2537.4%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #27high of 10274.2%
- #37low of 10271.5%
Leader: GPT-6 Astra · 91.2%
- #21max of 8049.4%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #10max of 4763.3%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #46max of 1710.6%
Leader: GPT-5.4 Pro · 0.0%
- #45max of 18199.9%
Leader: GPT-5 Pro · 100.0%
- #29max of 4365.4%
Leader: GPT-6 Sol · 78.0%
- #10max of 3212.2%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #202low of 3553.1s
- #236medium of 3555.2s
- #289high of 35520s
- #306xhigh of 35532s
- #322max of 35557s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #286high of 35554.0
- #288xhigh of 35553.9
- #291low of 35553.2
- #292max of 35553.1
- #294medium of 35552.9
Leader: Celeris-1 · 1461.1
- #144max of 1783.1s
Leader: Qwen 2.5 7B · 0.2s
- #73max of 17771.0
Leader: Llama 3.2 3b · 199.5