- #3max of 3957.5%
Leader: Gemini 3.7 Flash · 60.0%
- #6max of 2221675
- #8xhigh of 2221656
- #17high of 2221581
- #27medium of 2221528
- #38low of 2221482
Leader: Sonnet 5.5 · 1823
- #3max of 4912.7%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #33max of 11522.4%
- #34xhigh of 11521.3%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #36max of 21759.4%
- #43xhigh of 21757.8%
- #52high of 21755.3%
- #56medium of 21754.7%
- #68low of 21752.2%
Leader: Gemini 4 · 77.5%
- #3max of 813.2%
Leader: GPT-6 Astra · 19.2%
agentic
Harvey's Legal Agent Benchmark
76 results- #30max of 766.7%
Leader: Muse Spark 1.2 · 25.4%
- #6max of 4549.5%
Leader: GPT-5.6 Sol · 56.2%
agentic
Legal Research Bench
75 results- #3max of 7555.3%
Leader: Muse Spark 1.3 · 55.3%
- #5max of 3461.6%
Leader: DeepSeek V4.1 Flash · 69.8%
- #1max of 6749.2%
Leader: this model · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #4max of 4558.1%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #6max of 3940.0%
Leader: GPT-6 Astra · 62.9%
agentic
Time Horizon Index: KSP
13 results- #3max of 1363.3%
Leader: Opus 5.5 · 91.3%
- #26max of 67$5,422
Leader: GPT-6 Astra · $15,515
- #4xhigh of 1726.0%
Leader: GPT-6 Astra · 42.2%
- #8max of 20647.2%
- #12xhigh of 20645.8%
- #22high of 20643.1%
- #27medium of 20641.0%
- #34low of 20639.0%
Leader: Qwen 3.8 Max · 51.3%
coding
Arena Image-to-WebDev
58 results- #4max of 581720
Leader: Opus 5.5 · 1749
- #5max of 1221744
Leader: Opus 5.5 · 1813
- #9max of 7554.6%
Leader: Sonnet 5.5 · 69.8%
- #3max of 1191.2%
Leader: Opus 5.5 · 97.3%
coding
FrontierCode 1.1 Extended
130 results- #9medium of 13063.6%
- #14high of 13062.7%
- #17max of 13062.0%
- #18low of 13061.6%
- #20xhigh of 13061.4%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #12medium of 13050.9%
- #14high of 13050.3%
- #15max of 13050.3%
- #19low of 13049.8%
- #24xhigh of 13048.7%
Leader: Opus 5.5 · 54.6%
- #4max of 2156.3%
Leader: GPT-6 Astra · 65.5%
- #6max of 4290.8%
Leader: Gemini 4 · 100.0%
coding
LiveCodeBench (Vals)
127 results- #1max of 12790.5%
Leader: this model · 90.5%
coding
ProgramBench v1 Almost Resolved
62 results- #2max of 6253.5%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #2max of 627.0%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #5max of 6282.7%
Leader: Opus 5.5 · 87.0%
- #3max of 22463.1%
- #7xhigh of 22460.9%
- #18high of 22458.7%
- #31low of 22456.7%
- #36medium of 22456.4%
Leader: Opus 5.5 · 66.9%
- #2xhigh of 2260.0%
Leader: Opus 5 · 63.2%
coding
Terminal-Bench 2.1 (Vals)
74 results- #4high of 7485.0%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Science-AA
45 results- #8max of 4543.3%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v2.1-AA
238 results- #1max of 23891.4%
- #2xhigh of 23891.0%
- #3high of 23889.9%
- #13medium of 23888.0%
- #26low of 23885.0%
Leader: this model · 91.4%
coding
Terminal-Bench v4-AA
214 results- #10xhigh of 21455.1%
- #14max of 21452.0%
- #15high of 21452.0%
- #22medium of 21444.9%
- #28low of 21440.4%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench 1-100
22 results- #3max of 2228.0%
Leader: Opus 5.5 · 30.4%
coding
Vibe Code Bench v1.1
103 results- #6max of 10390.3%
Leader: Sonnet 5.5 · 92.4%
- #3xhigh of 12792.9%
- #5high of 12792.3%
Leader: GPT-6 Astra Pro · 93.6%
composite
AA Intelligence Index v4.3.2
607 results- #5max of 60753.4
- #6xhigh of 60753.2
- #13high of 60751.2
- #21medium of 60748.9
- #27low of 60746.8
Leader: Opus 5.5 · 57.6
- #6max of 2201501
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #27max of 2201485
Leader: Gemini 4 · 1553
- #4max of 4565.8%
Leader: Gemini 4 · 68.9%
composite
Vals RSI Index v1.1
24 results- #2max of 2436.1%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #2max of 2422.6%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #3max of 2428.0%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #3max of 2444.3%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #2max of 2449.5%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #155low of 177$2.37
- #161medium of 177$2.98
- #167high of 177$3.91
- #174xhigh of 177$5.98
- #176max of 177$7.63
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #152low of 177$3,158
- #159medium of 177$3,983
- #166high of 177$5,242
- #175xhigh of 177$9,063
- #177max of 177$13,129
Leader: GPT-6 Luna · $10.63
- #8max of 721967
Leader: GPT-6 Astra Pro · 2278
economics
Excel Modeling Benchmark
72 results- #1max of 7276.7%
Leader: this model · 76.7%
economics
Finance Agent v2
76 results- #7max of 7658.9%
Leader: Gemini 4 · 65.4%
- #11low of 20828.0%
- #16high of 20826.8%
- #17medium of 20826.8%
- #21max of 20826.2%
- #22xhigh of 20826.2%
Leader: GPT-6 Astra · 32.2%
- #4max of 2861756
- #5xhigh of 2861733
- #18high of 2861633
- #41medium of 2861545
- #52low of 2861468
Leader: Opus 5.5 · 1866
economics
MortgageTax (Vals)
84 results- #2max of 8470.8%
Leader: Opus 5 · 72.1%
- #9max of 12576.0%
Leader: Muse Spark 1.2 · 80.4%
knowledge
AA-Omniscience Accuracy
523 results- #1max of 52367.2%
- #3xhigh of 52366.2%
- #6high of 52364.9%
- #10medium of 52363.1%
- #20low of 52360.2%
Leader: this model · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #183low of 52334.4%
- #207high of 52331.2%
- #210medium of 52330.9%
- #216xhigh of 52329.5%
- #220max of 52327.4%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
- #11max of 53893.7%
- #21xhigh of 53893.4%
- #58high of 53890.6%
- #88medium of 53888.6%
- #92low of 53888.1%
Leader: GPT-6 Astra · 96.3%
knowledge
GPQA Diamond (Vals)
122 results- #9max of 12293.4%
Leader: Gemini 3.1 Pro · 95.5%
knowledge
Harvey LAB-AA v1.1
25 results- #5max of 256.4%
Leader: Grok 4.7 · 9.4%
- #8max of 9453.5%
Leader: Opus 5 · 63.6%
- #2max of 9691.3%
Leader: Opus 5.5 · 91.4%
- #1max of 12292.4%
Leader: this model · 92.4%
- #1max of 8190.6%
Leader: this model · 90.6%
knowledge
Public Benefits Bench v1.1
48 results- #2max of 4874.9%
Leader: Opus 5 · 76.9%
- #4max of 52085.3%
- #7medium of 52084.7%
- #15high of 52083.7%
- #24xhigh of 52083.0%
- #39low of 52082.3%
Leader: Kimi K3 · 88.7%
long-context
Arena Document
43 results- #2max of 431513
Leader: Opus 5 · 1516
- #2max of 9371.1%
Leader: Sonnet 5.5 · 75.0%
- #4max of 70.0%
Leader: GPT-6 Astra · 2.9%
math
FrontierMath Tiers 1–3 (v2)
115 results- #4max of 11590.2%
Leader: GPT-6.1 Sol · 93.7%
- #3max of 50100.0%
Leader: Sonnet 5.5 · 100.0%
- #16max of 1161286
Leader: Fable 5 · 1308
- #3max of 859.3%
Leader: Opus 5.5 · 29.1%
- #13max of 21597.5%
- #24xhigh of 21596.5%
- #28high of 21596.0%
- #38medium of 21594.5%
- #72low of 21590.0%
Leader: Fable 5 · 98.5%
- #14max of 21690.0%
- #15xhigh of 21690.0%
- #19high of 21688.8%
- #25medium of 21686.3%
- #38low of 21678.3%
Leader: GPT-6 Astra · 95.0%
reasoning
Blueprint-Bench 2
28 results- #4max of 2841.9%
Leader: Gemini 4 · 54.4%
- #9xhigh of 52631.1%
- #14high of 52630.3%
- #18max of 52629.7%
- #21medium of 52629.1%
- #30low of 52627.7%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #5max of 3170.0%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #5max of 2540.8%
Leader: GPT-6 Astra · 53.6%
- #2max of 13488.5%
Leader: Fable 5 · 88.6%
reasoning
MysteryMechanism
25 results- #4max of 2547.7%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #30high of 10273.1%
- #35low of 10272.0%
Leader: GPT-6 Astra · 91.2%
- #27max of 8048.5%
Leader: Opus 4.7 · 56.1%
reliability
Tool-call error rate
171 results- #55max of 1710.8%
Leader: GPT-5.4 Pro · 0.0%
- #74max of 18199.7%
Leader: GPT-5 Pro · 100.0%
- #13max of 1850.8%
Leader: Grok 4.7 · 68.3%
- #19max of 4370.4%
Leader: GPT-6 Sol · 78.0%
- #9max of 3222.9%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #195low of 3552.9s
- #247medium of 3557.4s
- #290high of 35521s
- #339xhigh of 355116s
- #350max of 355290s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #231max of 35570.2
- #247xhigh of 35563.4
- #268high of 35556.9
- #274medium of 35555.8
- #275low of 35555.8
Leader: Celeris-1 · 1461.1
- #163max of 1784.2s
Leader: Qwen 2.5 7B · 0.2s
- #122max of 17748.0
Leader: Llama 3.2 3b · 199.5