- #4max of 3956.3%
Leader: Gemini 3.7 Flash · 60.0%
- #2max of 2221807
- #3xhigh of 2221766
- #5high of 2221690
- #13medium of 2221627
- #70low of 2221279
Leader: Sonnet 5.5 · 1823
- #1high of 4914.3%
Leader: this model · 14.3%
agentic
AutomationBench (Zapier)
115 results- #4max of 11542.5%
- #9xhigh of 11535.8%
- #12high of 11533.0%
- #19medium of 11529.5%
- #28low of 11524.2%
Leader: Gemini 4 · 51.3%
agentic
AutomationBench-AA
217 results- #3max of 21769.5%
- #13xhigh of 21765.0%
- #19high of 21763.2%
- #27medium of 21761.2%
- #66low of 21752.9%
Leader: Gemini 4 · 77.5%
- #4max of 2579.3%
Leader: Sonnet 5.5 · 81.1%
- #2max of 814.0%
Leader: GPT-6 Astra · 19.2%
agentic
Harvey's Legal Agent Benchmark
76 results- #38max of 763.8%
Leader: Muse Spark 1.2 · 25.4%
- #23max of 4538.2%
Leader: GPT-5.6 Sol · 56.2%
agentic
Legal Research Bench
75 results- #5max of 7550.5%
Leader: Muse Spark 1.3 · 55.3%
- #2max of 6745.6%
Leader: Fable 5.1 · 49.2%
agentic
Terminal-Bench 4.0 (Vals)
45 results- #1max of 4565.2%
Leader: this model · 65.2%
agentic
Terminal-Bench Science (Vals)
39 results- #3max of 3947.1%
Leader: GPT-6 Astra · 62.9%
agentic
Time Horizon Index: KSP
13 results- #1max of 1391.3%
Leader: this model · 91.3%
- #9max of 67$9,235
Leader: GPT-6 Astra · $15,515
agentic
Vending-Bench Arena
12 results- #2max of 12$8,100
Leader: GPT-6 Astra · $12,400
- #3xhigh of 1731.2%
Leader: GPT-6 Astra · 42.2%
coding
Arena Image-to-WebDev
58 results- #1max of 581749
Leader: this model · 1749
- #1max of 1221813
Leader: this model · 1813
- #4max of 7566.6%
Leader: Sonnet 5.5 · 69.8%
- #1max of 1197.3%
Leader: this model · 97.3%
coding
FrontierCode 1.1 Extended
130 results- #1medium of 13065.3%
- #2high of 13065.2%
- #10max of 13063.6%
- #11xhigh of 13063.5%
- #27low of 13060.3%
Leader: this model · 65.3%
coding
FrontierCode 1.1 Main
130 results- #1medium of 13054.6%
- #2max of 13054.4%
- #3high of 13054.0%
- #10xhigh of 13051.4%
- #35low of 13047.3%
Leader: this model · 54.6%
- #2max of 2162.3%
Leader: GPT-6 Astra · 65.5%
- #4max of 4295.1%
Leader: Gemini 4 · 100.0%
coding
ProgramBench v1 Almost Resolved
62 results- #1max of 6265.0%
Leader: this model · 65.0%
coding
ProgramBench v1 Fully Resolved
62 results- #1max of 6218.5%
Leader: this model · 18.5%
coding
ProgramBench v1 Raw Pass Rate
62 results- #1max of 6287.0%
Leader: this model · 87.0%
- #1max of 22466.9%
- #2xhigh of 22465.0%
- #9high of 22460.4%
- #13medium of 22459.3%
- #20low of 22458.6%
Leader: this model · 66.9%
coding
Terminal-Bench 2.1 (Vals)
74 results- #1high of 7487.6%
Leader: this model · 87.6%
coding
Terminal-Bench Science-AA
45 results- #2xhigh of 4561.9%
- #3max of 4559.0%
- #7high of 4549.0%
- #9medium of 4543.3%
- #13low of 4524.3%
Leader: GPT-6 Astra · 63.3%
coding
Terminal-Bench v4-AA
214 results- #2max of 21459.6%
- #3xhigh of 21459.6%
- #8high of 21456.6%
- #13medium of 21452.5%
- #40low of 21431.3%
Leader: Sonnet 5.5 · 63.6%
coding
Vibe Code Bench 1-100
22 results- #1xhigh of 2230.4%
Leader: this model · 30.4%
coding
Vibe Code Bench v1.1
103 results- #5max of 10390.3%
Leader: Sonnet 5.5 · 92.4%
composite
AA Intelligence Index v4.3.2
607 results- #1max of 60757.6
- #3xhigh of 60756.0
- #4high of 60753.6
- #12medium of 60751.2
- #47low of 60742.3
Leader: this model · 57.6
- #2high of 2201507
Leader: Gemini 4 · 1525
composite
Arena Text — Multi-Turn
220 results- #10high of 2201498
Leader: Gemini 4 · 1553
- #3max of 4567.0%
Leader: Gemini 4 · 68.9%
composite
Vals RSI Index v1.1
24 results- #1max of 2437.3%
Leader: this model · 37.3%
composite
Vals RSI Index v1.1: Harness Engineering: Judge
24 results- #12max of 2414.5%
Leader: GPT-6 Sol · 22.8%
composite
Vals RSI Index v1.1: Post-training: Finance Agent
24 results- #1max of 2435.1%
Leader: this model · 35.1%
composite
Vals RSI Index v1.1: Pre-training: Compression
24 results- #1max of 2446.5%
Leader: this model · 46.5%
composite
Vals RSI Index v1.1: Pre-training: LM Training
24 results- #1max of 2453.0%
Leader: this model · 53.0%
cost
AA Intelligence Index Cost per Task
177 results- #95low of 177$0.55
- #131medium of 177$1.34
- #145high of 177$1.82
- #163xhigh of 177$3.46
- #175max of 177$5.98
Leader: GPT-6 Luna · $0.00
cost
Cost to Run AA Intelligence Index
177 results- #90low of 177$860
- #128medium of 177$1,627
- #140high of 177$2,172
- #160xhigh of 177$4,057
- #174max of 177$8,708
Leader: GPT-6 Luna · $10.63
- #3max of 722186
Leader: GPT-6 Astra Pro · 2278
creative
VoxelBench (text)
54 results- #2max of 542593
Leader: GPT-6 Astra · 2654
economics
Excel Modeling Benchmark
72 results- #2max of 7275.9%
Leader: Fable 5.1 · 76.7%
economics
Finance Agent v2
76 results- #9max of 7658.6%
Leader: Gemini 4 · 65.4%
- #10high of 20828.8%
- #18xhigh of 20826.6%
- #20max of 20826.2%
- #25medium of 20825.6%
- #26low of 20825.6%
Leader: GPT-6 Astra · 32.2%
- #1max of 2861866
- #3xhigh of 2861836
- #10high of 2861705
- #35medium of 2861584
- #108low of 2861234
Leader: this model · 1866
knowledge
AA-Omniscience Accuracy
523 results- #2max of 52366.2%
- #4xhigh of 52365.4%
- #7high of 52364.6%
- #8medium of 52364.5%
- #9low of 52363.5%
Leader: Fable 5.1 · 67.2%
knowledge
AA-Omniscience Non-hallucination
523 results- #146max of 52341.4%
- #184xhigh of 52334.3%
- #198high of 52332.4%
- #199low of 52332.4%
- #204medium of 52331.6%
Leader: MiniCPM5-1B (Non-reasoning) · 99.1%
knowledge
Harvey LAB-AA v1.1
25 results- #7max of 254.2%
Leader: Grok 4.7 · 9.4%
- #18max of 9449.8%
Leader: Opus 5 · 63.6%
- #1max of 9691.4%
Leader: this model · 91.4%
- #1max of 26487.7%
- #3xhigh of 26486.6%
- #7high of 26485.8%
- #8medium of 26485.7%
- #17low of 26484.7%
Leader: this model · 87.7%
knowledge
Public Benefits Bench v1.1
48 results- #3max of 4870.6%
Leader: Opus 5 · 76.9%
- #5max of 52084.7%
- #6xhigh of 52084.7%
- #8medium of 52084.3%
- #34high of 52082.7%
- #62low of 52080.7%
Leader: Kimi K3 · 88.7%
- #3max of 9366.7%
Leader: Sonnet 5.5 · 75.0%
- #2max of 72.9%
Leader: GPT-6 Astra · 2.9%
math
FrontierMath Tiers 1–3 (v2)
115 results- #3max of 11591.2%
Leader: GPT-6.1 Sol · 93.7%
- #2max of 50100.0%
Leader: Sonnet 5.5 · 100.0%
- #10high of 1161293
Leader: Fable 5 · 1308
- #1max of 8529.1%
Leader: this model · 29.1%
- #5high of 21598.5%
- #16max of 21597.5%
- #17xhigh of 21597.5%
- #18medium of 21597.5%
- #77low of 21588.5%
Leader: Fable 5 · 98.5%
- #4high of 21693.3%
- #6xhigh of 21692.5%
- #9max of 21691.7%
- #23medium of 21687.5%
- #52low of 21670.1%
Leader: GPT-6 Astra · 95.0%
reasoning
Blueprint-Bench 2
28 results- #2max of 2851.2%
Leader: Gemini 4 · 54.4%
- #2max of 52631.7%
- #3xhigh of 52631.7%
- #11high of 52630.9%
- #27medium of 52627.7%
- #71low of 52617.7%
Leader: GPT-5.6 Sol · 32.3%
reasoning
Furniture Assembly
31 results- #1max of 3183.3%
Leader: this model · 83.3%
reasoning
Humanity's Sixth Sense
25 results- #3xhigh of 2544.6%
Leader: GPT-6 Astra · 53.6%
reasoning
MysteryMechanism
25 results- #2max of 2549.5%
Leader: GPT-6 Astra · 53.2%
reasoning
Roboflow Visual Reasoning
102 results- #4high of 10285.9%
- #8low of 10283.0%
Leader: GPT-6 Astra · 91.2%
- #37max of 8045.8%
Leader: Opus 4.7 · 56.1%
reliability
Tool-call error rate
171 results- #45max of 1710.6%
Leader: GPT-5.4 Pro · 0.0%
- #85max of 18199.6%
Leader: GPT-5 Pro · 100.0%
- #7max of 1858.3%
Leader: Grok 4.7 · 68.3%
- #8max of 4374.6%
Leader: GPT-6 Sol · 78.0%
- #4max of 3233.6%
Leader: GPT-6 Astra · 56.9%
speed
Latency (Time to First Token)
355 results- #256low of 3558.7s
- #292medium of 35522s
- #312high of 35540s
- #342xhigh of 355128s
- #355max of 355704s
Leader: Gemini 2.5 Flash-Lite · 0.3s
speed
Output Speed (tokens/s)
355 results- #156max of 35597.0
- #193xhigh of 35583.2
- #203high of 35579.9
- #207medium of 35579.2
- #210low of 35578.2
Leader: Celeris-1 · 1461.1
- #145max of 1783.1s
Leader: Qwen 2.5 7B · 0.2s
- #53max of 17781.0
Leader: Llama 3.2 3b · 199.5