- #2max of 3957.5%
Leader: Gemini 3.7 Flash · 60.0%
- #1max of 2221823
- #4xhigh of 2221752
- #10high of 2221638
- #45medium of 2221442
- #73low of 2221271
Leader: this model · 1823
- #4max of 4912.0%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #3max of 11544.8%
- #8xhigh of 11536.8%
- #20high of 11529.4%
- #31medium of 11523.3%
- #41low of 11518.3%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #2max of 21771.8%
- #12xhigh of 21765.5%
- #35high of 21759.4%
- #55medium of 21754.9%
- #74low of 21749.4%
Leader: Gemini 4 · 77.5%
- #1max of 2581.1%
Leader: this model · 81.1%
agentic
Harvey's Legal Agent Benchmark
76 results- #41max of 762.9%
Leader: Muse Spark 1.2 · 25.4%
agentic
Legal Research Bench
75 results- #10max of 7548.1%
Leader: Muse Spark 1.3 · 55.3%
- #6max of 6742.0%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #2max of 4564.1%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #4max of 3945.7%
Leader: GPT-6 Astra · 62.9%
- #5max of 67$11,165
Leader: GPT-6 Astra · $15,515
- #5xhigh of 1719.9%
Leader: GPT-6 Astra · 42.2%
coding
Arena Image-to-WebDev
58 results- #2xhigh of 581740
Leader: Opus 5.5 · 1749
- #3xhigh of 1221774
- #6high of 1221715
Leader: Opus 5.5 · 1813
- #1max of 7569.8%
Leader: this model · 69.8%
coding
FrontierCode 1.1 Extended
130 results- #5xhigh of 13064.4%
- #19high of 13061.5%
- #39max of 13059.1%
- #89medium of 13050.5%
- #105low of 13043.4%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #8xhigh of 13052.1%
- #20high of 13049.4%
- #40max of 13046.2%
- #85medium of 13036.5%
- #100low of 13029.3%
Leader: Opus 5.5 · 54.6%
- #3max of 2161.9%
Leader: GPT-6 Astra · 65.5%
- #9max of 4283.1%
Leader: Gemini 4 · 100.0%
coding
ProgramBench v1 Almost Resolved
62 results- #4max of 6249.5%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #3max of 626.5%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #10max of 6276.3%
Leader: Opus 5.5 · 87.0%
- #5max of 22461.0%
- #27xhigh of 22457.3%
- #70high of 22453.7%
- #77medium of 22452.9%
- #114low of 22449.1%
Leader: Opus 5.5 · 66.9%
coding
Terminal-Bench 2.1 (Vals)
74 results- #7high of 7483.1%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Science-AA
45 results- #5max of 4553.3%
- #6xhigh of 4552.4%
- #10high of 4534.3%
- #19medium of 4516.2%
- #23low of 4510.5%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v4-AA
214 results- #1max of 21463.6%
- #7xhigh of 21457.1%
- #24high of 21443.9%
- #43medium of 21429.8%
- #58low of 21420.7%
Leader: this model · 63.6%
coding
Vibe Code Bench v1.1
103 results- #1max of 10392.4%
Leader: this model · 92.4%
composite
AA Intelligence Index v4.3.2
607 results- #2max of 60756.0
- #10xhigh of 60751.9
- #28high of 60746.8
- #56medium of 60740.8
- #81low of 60735.9
Leader: Opus 5.5 · 57.6
- #31xhigh of 2201476
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #29xhigh of 2201484
Leader: Gemini 4 · 1553
- #2max of 4567.0%
Leader: Gemini 4 · 68.9%
cost
AA Intelligence Index Cost per Task
177 results- #77low of 177$0.35
- #88medium of 177$0.48
- #109high of 177$0.88
- #150xhigh of 177$2.01
- #172max of 177$5.46
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #71low of 177$482
- #80medium of 177$622
- #98high of 177$1,027
- #141xhigh of 177$2,177
- #172max of 177$7,260
Leader: GPT-6 Luna · $10.63
- #6max of 722039
Leader: GPT-6 Astra Pro · 2278
creative
VoxelBench (text)
54 results- #54max of 541000
Leader: GPT-6 Astra · 2654
economics
Excel Modeling Benchmark
72 results- #3max of 7275.7%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #10max of 7658.1%
Leader: Gemini 4 · 65.4%
- #24max of 20825.8%
- #28high of 20825.2%
- #29xhigh of 20824.6%
- #59medium of 20820.2%
- #85low of 20816.0%
Leader: GPT-6 Astra · 32.2%
- #2max of 2861838
- #6xhigh of 2861730
- #39high of 2861550
- #92medium of 2861323
- #118low of 2861175
Leader: Opus 5.5 · 1866
knowledge
AA-Omniscience Accuracy
523 results- #42max of 52353.9%
- #48xhigh of 52353.0%
- #53high of 52352.0%
- #75medium of 52347.2%
- #80low of 52346.3%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #100max of 52353.0%
- #110low of 52349.8%
- #116medium of 52348.8%
- #170xhigh of 52337.1%
- #180high of 52335.4%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Harvey LAB-AA v1.1
25 results- #11max of 252.8%
Leader: Grok 4.7 · 9.4%
- #12max of 9452.9%
Leader: Opus 5 · 63.6%
- #3max of 9691.1%
Leader: Opus 5.5 · 91.4%
knowledge
Public Benefits Bench v1.1
48 results- #13max of 4867.2%
Leader: Opus 5 · 76.9%
- #33max of 52082.7%
- #84xhigh of 52079.7%
- #115high of 52078.0%
- #139medium of 52076.3%
- #142low of 52076.0%
Leader: Kimi K3 · 88.7%
- #1max of 9375.0%
Leader: this model · 75.0%
- #3max of 72.9%
Leader: GPT-6 Astra · 2.9%
math
FrontierMath Tiers 1–3 (v2)
115 results- #7max of 11588.8%
Leader: GPT-6.1 Sol · 93.7%
- #1max of 50100.0%
Leader: this model · 100.0%
- #38xhigh of 1161266
Leader: Fable 5 · 1308
- #11max of 851.1%
Leader: Opus 5.5 · 29.1%
reasoning
Blueprint-Bench 2
28 results- #5max of 2839.3%
Leader: Gemini 4 · 54.4%
- #7max of 52631.4%
- #10xhigh of 52631.1%
- #44high of 52624.6%
- #77medium of 52616.9%
- #104low of 52611.4%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #4max of 3175.0%
Leader: Opus 5.5 · 83.3%
reasoning
MysteryMechanism
25 results- #3max of 2549.1%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #6high of 10283.9%
- #19low of 10276.4%
Leader: GPT-6 Astra · 91.2%
- #11max of 8051.8%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #18max of 4744.7%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #98max of 1712.0%
Leader: GPT-5.4 Pro · 0.0%
- #47max of 18199.9%
Leader: GPT-5 Pro · 100.0%
- #33max of 4359.6%
Leader: GPT-6 Sol · 78.0%
- #8max of 3230.2%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #34low of 3550.9s
- #135medium of 3551.9s
- #266high of 35512s
- #307xhigh of 35533s
- #354max of 355497s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #89max of 355141.8
- #132medium of 355108.7
- #133high of 355108.7
- #134xhigh of 355108.5
- #137low of 355108.2
Leader: Celeris-1 · 1461.1
- #97max of 1781.4s
Leader: Qwen 2.5 7B · 0.2s
- #24max of 177103.5
Leader: Llama 3.2 3b · 199.5