- #6max of 3951.2%
Leader: Gemini 3.7 Flash · 60.0%
- #19max of 2221570
- #22xhigh of 2221546
- #30high of 2221506
- #42medium of 2221461
- #77low of 2221257
Leader: Sonnet 5.5 · 1823
- #2max of 4913.1%
Leader: Opus 5.5 · 14.3%
agentic
AutomationBench (Zapier)
115 results- #5max of 11541.4%
- #6xhigh of 11539.0%
- #7high of 11537.1%
- #10medium of 11534.1%
- #16low of 11530.3%
- #26none of 11525.3%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #5max of 21768.5%
- #6xhigh of 21767.2%
- #9high of 21766.6%
- #15medium of 21764.6%
- #37low of 21759.1%
Leader: Gemini 4 · 77.5%
- #3max of 2579.3%
Leader: Sonnet 5.5 · 81.1%
- #1max of 819.2%
Leader: this model · 19.2%
agentic
Harvey's Legal Agent Benchmark
76 results- #31max of 765.4%
Leader: Muse Spark 1.2 · 25.4%
- #8max of 4548.6%
Leader: GPT-5.6 Sol · 56.2%
agentic
Legal Research Bench
75 results- #28max of 7539.4%
Leader: Muse Spark 1.3 · 55.3%
- #41max of 6720.7%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 2.1
18 results- #1high of 1887.4%
- #2medium of 1887.0%
- #3low of 1886.7%
- #4max of 1886.7%
- #5xhigh of 1885.8%
Leader: this model · 87.4%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #3max of 4559.6%
Leader: Opus 5.5 · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #1max of 3962.9%
Leader: this model · 62.9%
agentic
Time Horizon Index: KSP
13 results- #2max of 1390.5%
Leader: Opus 5.5 · 91.3%
- #1max of 67$15,515
Leader: this model · $15,515
agentic
Vending-Bench Arena
12 results- #1max of 12$12,400
Leader: this model · $12,400
- #1xhigh of 1742.2%
Leader: this model · 42.2%
- #21xhigh of 20643.1%
- #26max of 20641.4%
- #30high of 20640.0%
- #45medium of 20635.5%
- #59low of 20632.0%
Leader: Qwen 3.8 Max · 51.3%
coding
Arena Image-to-WebDev
58 results- #3max of 581731
Leader: Opus 5.5 · 1749
- #2max of 1221786
Leader: Opus 5.5 · 1813
- #3max of 7567.7%
Leader: Sonnet 5.5 · 69.8%
- #2max of 1195.1%
Leader: Opus 5.5 · 97.3%
coding
FrontierCode 1.1 Extended
130 results- #4max of 13064.5%
- #12high of 13063.1%
- #16xhigh of 13062.1%
- #28medium of 13060.3%
- #47low of 13057.4%
Leader: Opus 5.5 · 65.3%
coding
FrontierCode 1.1 Main
130 results- #6max of 13053.3%
- #11high of 13050.9%
- #13xhigh of 13050.6%
- #23medium of 13048.8%
- #47low of 13045.3%
Leader: Opus 5.5 · 54.6%
- #1max of 2165.5%
Leader: this model · 65.5%
- #2max of 42100.0%
Leader: Gemini 4 · 100.0%
coding
ProgramBench v1 Almost Resolved
62 results- #3max of 6250.0%
Leader: Opus 5.5 · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #4max of 625.5%
Leader: Opus 5.5 · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #2max of 6285.4%
Leader: Opus 5.5 · 87.0%
- #33max of 22456.5%
- #42xhigh of 22455.7%
- #46high of 22455.4%
- #62medium of 22454.2%
- #64low of 22454.1%
Leader: Opus 5.5 · 66.9%
- #3xhigh of 2259.1%
Leader: Opus 5 · 63.2%
coding
Terminal-Bench 2.1 (Vals)
74 results- #2max of 7487.3%
Leader: Opus 5.5 · 87.6%
coding
Terminal-Bench Science-AA
45 results- #1max of 4563.3%
Leader: this model · 63.3%
coding
Terminal-Bench v2.1-AA
238 results- #4high of 23889.9%
- #5medium of 23889.5%
- #7xhigh of 23889.1%
- #10max of 23888.4%
- #15low of 23888.0%
Leader: Fable 5.1 · 91.4%
coding
Terminal-Bench v4-AA
214 results- #4xhigh of 21459.6%
- #5max of 21459.1%
- #12high of 21454.0%
- #17medium of 21449.5%
- #26low of 21441.9%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench 1-100
22 results- #4max of 2227.6%
Leader: Opus 5.5 · 30.4%
coding
Vibe Code Bench v1.1
103 results- #7max of 10389.6%
Leader: Sonnet 5.5 · 92.4%
- #2xhigh of 12793.3%
- #4high of 12792.9%
Leader: GPT-6 Astra Pro · 93.6%
composite
AA Intelligence Index v4.3.2
607 results- #7max of 60752.7
- #9xhigh of 60752.4
- #15high of 60750.9
- #20medium of 60749.6
- #32low of 60745.8
Leader: Opus 5.5 · 57.6
- #34max of 2201475
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #28max of 2201484
Leader: Gemini 4 · 1553
- #6max of 4563.1%
Leader: Gemini 4 · 68.9%
composite
Vals RSI Index v1.1
24 results- #6max of 2427.1%
Leader: Opus 5.5 · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #18max of 249.5%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #6max of 247.8%
Leader: Opus 5.5 · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #5max of 2442.6%
Leader: Opus 5.5 · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #3max of 2448.4%
Leader: Opus 5.5 · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #106low of 177$0.82
- #137medium of 177$1.54
- #142high of 177$1.73
- #154xhigh of 177$2.31
- #162max of 177$3.26
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #122low of 177$1,537
- #145medium of 177$2,434
- #151high of 177$2,925
- #157xhigh of 177$3,803
- #168max of 177$5,324
Leader: GPT-6 Luna · $10.63
creative
VoxelBench (text)
54 results- #1max of 542654
Leader: this model · 2654
economics
Excel Modeling Benchmark
72 results- #9max of 7271.7%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #29max of 7653.5%
Leader: Gemini 4 · 65.4%
- #1xhigh of 20832.2%
- #4max of 20831.0%
- #6high of 20831.0%
- #7medium of 20830.4%
- #8low of 20830.4%
Leader: this model · 32.2%
- #36max of 2861574
- #38xhigh of 2861554
- #44high of 2861524
- #50medium of 2861485
- #73low of 2861389
Leader: Opus 5.5 · 1866
knowledge
AA-Omniscience Accuracy
523 results- #11max of 52362.6%
- #13xhigh of 52361.9%
- #14high of 52361.1%
- #18medium of 52360.6%
- #21low of 52359.5%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #90high of 52355.2%
- #97medium of 52353.5%
- #98low of 52353.1%
- #101xhigh of 52351.7%
- #117max of 52348.7%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
- #1xhigh of 53896.3%
- #2max of 53896.1%
- #4high of 53894.9%
- #10medium of 53893.9%
- #24low of 53893.1%
Leader: this model · 96.3%
knowledge
Harvey LAB-AA v1.1
25 results- #3max of 258.6%
Leader: Grok 4.7 · 9.4%
- #28max of 9448.5%
Leader: Opus 5 · 63.6%
- #15max of 9687.9%
Leader: Opus 5.5 · 91.4%
- #2max of 26486.9%
- #4high of 26486.4%
- #5xhigh of 26486.2%
- #12medium of 26485.1%
- #18low of 26484.6%
Leader: Opus 5.5 · 87.7%
- #61max of 52080.7%
- #73xhigh of 52080.0%
- #74high of 52080.0%
- #75low of 52080.0%
- #86medium of 52079.7%
Leader: Kimi K3 · 88.7%
long-context
Arena Document
43 results- #16max of 431468
Leader: Opus 5 · 1516
- #17max of 9335.0%
Leader: Sonnet 5.5 · 75.0%
- #1max of 72.9%
Leader: this model · 2.9%
math
FrontierMath Tiers 1–3 (v2)
115 results- #2max of 11593.7%
Leader: GPT-6.1 Sol · 93.7%
- #6max of 5099.0%
Leader: Sonnet 5.5 · 100.0%
- #18max of 1161286
Leader: Fable 5 · 1308
- #2max of 8517.0%
Leader: Opus 5.5 · 29.1%
- #3xhigh of 21598.5%
- #4high of 21598.5%
- #14max of 21597.5%
- #15medium of 21597.5%
- #25low of 21596.5%
- #95none of 21586.0%
Leader: Fable 5 · 98.5%
- #1max of 21695.0%
- #3xhigh of 21693.3%
- #7high of 21692.1%
- #8medium of 21692.1%
- #27low of 21685.4%
- #81none of 21659.6%
Leader: this model · 95.0%
reasoning
Blueprint-Bench 2
28 results- #3max of 2849.7%
Leader: Gemini 4 · 54.4%
- #4max of 52631.7%
- #8xhigh of 52631.4%
- #20medium of 52629.1%
- #22high of 52628.9%
- #36low of 52626.3%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #3max of 3180.0%
Leader: Opus 5.5 · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #1max of 2553.6%
Leader: this model · 53.6%
reasoning
MysteryMechanism
25 results- #1max of 2553.2%
Leader: this model · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #1high of 10291.2%
- #3low of 10287.2%
Leader: this model · 91.2%
- #36max of 8046.4%
Leader: Opus 4.7 · 56.1%
reasoning
SimpleQA Verified
47 results- #2max of 4775.8%
Leader: Gemini 3.1 Pro · 77.5%
reliability
Tool-call error rate
171 results- #13max of 1710.1%
Leader: GPT-5.4 Pro · 0.0%
- #93max of 18199.5%
Leader: GPT-5 Pro · 100.0%
- #4max of 1863.3%
Leader: Grok 4.7 · 68.3%
- #40max of 4341.1%
Leader: GPT-6 Sol · 78.0%
- #1max of 3256.9%
Leader: this model · 56.9%
speed
Latency (Time to First Token)
355 results- #177low of 3552.6s
- #250medium of 3558.3s
- #331high of 35588s
- #348xhigh of 355207s
- #353max of 355388s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #303high of 35551.4
- #310max of 35547.7
- #313xhigh of 35546.1
- #320low of 35542.8
- #323medium of 35541.9
Leader: Celeris-1 · 1461.1
- #168max of 1784.9s
Leader: Qwen 2.5 7B · 0.2s
- #131max of 17746.0
Leader: Llama 3.2 3b · 199.5