priorsBuilder

Meta · released 2026-09-02

Muse Spark 1.3

Measured on 74 benchmarks across 2 reasoning efforts.

General frontier

49.4

#15 of 32 models · best at max

Reasoning efforts in your mix

max49.4#52 of 128
xhigh46.8#60 of 128

Your mix’s benchmarks

5 of 8 measured

reasoning

Humanity's Last Exam

584 results
  1. #29max of 58448.7%
  2. #35xhigh of 58447.5%

Leader: Opus 5.5 · 61.4%

knowledge

AA-Omniscience

523 results
  1. #50max of 52325.0
  2. #53xhigh of 52323.1

Leader: Opus 5.5 · 46.4

math

FrontierMath Tier 4 (v2)

70 results
  1. #25max of 7046.3%
  2. #27xhigh of 7041.5%

Leader: GPT-6.1 Sol · 100.0%

reasoning

SimpleBench

97 results
  1. #8max of 9781.8%

Leader: Opus 5.5 · 88.4%

agentic

Terminal-Bench 4.0

35 results
  1. #32xhigh of 3514.6%

Leader: Opus 5.5 · 64.8%

All 74 benchmarks

show 69 more

agentic

AA-Briefcase v1.1

222 results
  1. #15max of 2221581
  2. #36xhigh of 2221486

Leader: Sonnet 5.5 · 1823

agentic

Agent Arena

49 results
  1. #15max of 494.0%

Leader: Opus 5.5 · 14.3%

agentic

AutomationBench (Zapier)

115 results
  1. #37max of 11520.7%

Leader: Gemini 4 · 51.3%

agentic

AutomationBench-AA

217 results
  1. #42max of 21757.9%
  2. #45xhigh of 21756.8%

Leader: Gemini 4 · 77.5%

agentic

CUA-bench

8 results
  1. #6max of 85.8%

Leader: GPT-6 Astra · 19.2%

agentic

Harvey's Legal Agent Benchmark

76 results
  1. #2max of 7623.8%
  2. #3xhigh of 7622.9%

Leader: Muse Spark 1.2 · 25.4%

agentic

ITBench-AA

45 results
  1. #29max of 4533.2%

Leader: GPT-5.6 Sol · 56.2%

agentic

Legal Research Bench

75 results
  1. #1max of 7555.3%
  2. #24xhigh of 7540.9%

Leader: this model · 55.3%

agentic

Tax Agent Bench

67 results
  1. #5max of 6742.4%
  2. #9xhigh of 6736.0%

Leader: Fable 5.1 · 49.2%

agentic

Terminal-Bench 4.0 (Vals)

45 results
  1. #17max of 4524.7%
  2. #33xhigh of 4510.6%

Leader: Opus 5.5 · 65.2%

agentic

Terminal-Bench Science (Vals)

39 results
  1. #13max of 3910.0%
  2. #23xhigh of 394.3%

Leader: GPT-6 Astra · 62.9%

agentic

WeirdML v3

17 results
  1. #14xhigh of 176.9%

Leader: GPT-6 Astra · 42.2%

agentic

τ³-Banking

206 results
  1. #3max of 20650.5%
  2. #9xhigh of 20647.2%

Leader: Qwen 3.8 Max · 51.3%

coding

Arena Image-to-WebDev

58 results
  1. #8max of 581642
  2. #15xhigh of 581598

Leader: Opus 5.5 · 1749

coding

Arena WebDev

122 results
  1. #12max of 1221657
  2. #19xhigh of 1221626

Leader: Opus 5.5 · 1813

coding

Code Migration

75 results
  1. #13max of 7547.4%
  2. #42xhigh of 7527.6%

Leader: Sonnet 5.5 · 69.8%

coding

FrontierSWE v2

21 results
  1. #13max of 2125.8%

Leader: GPT-6 Astra · 65.5%

coding

IOI

42 results
  1. #18max of 4256.6%
  2. #31xhigh of 4243.9%

Leader: Gemini 4 · 100.0%

coding

ProgramBench v1 Almost Resolved

62 results
  1. #13max of 6219.0%
  2. #24xhigh of 6213.0%

Leader: Opus 5.5 · 65.0%

coding

ProgramBench v1 Fully Resolved

62 results
  1. #8max of 622.5%
  2. #21xhigh of 620.5%

Leader: Opus 5.5 · 18.5%

coding

ProgramBench v1 Raw Pass Rate

62 results
  1. #18xhigh of 6271.7%
  2. #24max of 6268.7%

Leader: Opus 5.5 · 87.0%

coding

SciCode

224 results
  1. #11xhigh of 22459.7%
  2. #16max of 22458.8%

Leader: Opus 5.5 · 66.9%

coding

Terminal-Bench 2.1 (Vals)

74 results
  1. #12max of 7479.0%
  2. #24xhigh of 7472.3%

Leader: Opus 5.5 · 87.6%

coding

Terminal-Bench Science-AA

45 results
  1. #22max of 4511.0%

Leader: GPT-6 Astra · 63.3%

coding

Terminal-Bench v2.1-AA

238 results
  1. #25xhigh of 23885.4%
  2. #30max of 23884.3%

Leader: Fable 5.1 · 91.4%

coding

Terminal-Bench v4-AA

214 results
  1. #35max of 21433.3%
  2. #63xhigh of 21416.7%

Leader: Sonnet 5.5 · 63.6%

coding

Vibe Code Bench 1-100

22 results
  1. #6max of 2220.5%

Leader: Opus 5.5 · 30.4%

coding

Vibe Code Bench v1.1

103 results
  1. #12max of 10385.9%
  2. #17xhigh of 10382.9%

Leader: Sonnet 5.5 · 92.4%

composite

AA Intelligence Index v4.3.2

607 results
  1. #23max of 60748.1
  2. #34xhigh of 60745.1

Leader: Opus 5.5 · 57.6

composite

Arena Text

220 results
  1. #9max of 2201494

Leader: Gemini 4 · 1525

composite

Arena Text — Multi-Turn

220 results
  1. #23max of 2201487

Leader: Gemini 4 · 1553

composite

Vals Index

45 results
  1. #9max of 4558.2%
  2. #19xhigh of 4553.2%

Leader: Gemini 4 · 68.9%

composite

Vals RSI Index v1.1

24 results
  1. #17max of 2419.6%

Leader: Opus 5.5 · 37.3%

composite

Vals RSI Index v1.1: Harness Engineering: Judge

24 results
  1. #11max of 2414.5%

Leader: GPT-6 Sol · 22.8%

composite

Vals RSI Index v1.1: Post-training: Finance Agent

24 results
  1. #8max of 240.0%

Leader: Opus 5.5 · 35.1%

composite

Vals RSI Index v1.1: Pre-training: Compression

24 results
  1. #22max of 2423.1%

Leader: Opus 5.5 · 46.5%

composite

Vals RSI Index v1.1: Pre-training: LM Training

24 results
  1. #9max of 2440.9%

Leader: Opus 5.5 · 53.0%

cost

AA Intelligence Index Cost per Task

177 results
  1. #132xhigh of 177$1.37
  2. #140max of 177$1.60

Leader: GPT-6 Luna · $0.00

cost

Cost to Run AA Intelligence Index

177 results
  1. #130xhigh of 177$1,655
  2. #134max of 177$2,000

Leader: GPT-6 Luna · $10.63

creative

MineBench

72 results
  1. #20max of 721782

Leader: GPT-6 Astra Pro · 2278

creative

VoxelBench (text)

54 results
  1. #25max of 541680

Leader: GPT-6 Astra · 2654

economics

Excel Modeling Benchmark

72 results
  1. #15max of 7267.4%
  2. #29xhigh of 7262.7%

Leader: Fable 5.1 · 76.7%

economics

Finance Agent v2

76 results
  1. #4max of 7660.0%
  2. #6xhigh of 7658.9%

Leader: Gemini 4 · 65.4%

economics

GDP.pdf

208 results
  1. #19max of 20826.6%
  2. #33xhigh of 20824.2%

Leader: GPT-6 Astra · 32.2%

economics

GDPval-AA v2.1

286 results
  1. #13max of 2861681
  2. #20xhigh of 2861629

Leader: Opus 5.5 · 1866

knowledge

AA-Omniscience Accuracy

523 results
  1. #96max of 52343.6%
  2. #109xhigh of 52341.5%

Leader: Fable 5.1 · 67.2%

knowledge

AA-Omniscience Non-hallucination

523 results
  1. #47xhigh of 52368.5%
  2. #52max of 52367.1%

Leader: MiniCPM5-1B (Non-reasoning) · 99.1%

knowledge

GPQA Diamond

538 results
  1. #8xhigh of 53894.1%
  2. #14max of 53893.5%

Leader: GPT-6 Astra · 96.3%

knowledge

Harvey LAB-AA v1.1

25 results
  1. #2max of 258.9%

Leader: Grok 4.7 · 9.4%

knowledge

MMMU-Pro

264 results
  1. #34xhigh of 26482.0%

Leader: Opus 5.5 · 87.7%

long-context

AA-LCR v1.1

520 results
  1. #26max of 52083.0%
  2. #27xhigh of 52083.0%

Leader: Kimi K3 · 88.7%

long-context

Arena Document

43 results
  1. #15max of 431471

Leader: Opus 5 · 1516

long-context

MLCR-AA

93 results
  1. #15max of 9343.3%

Leader: Sonnet 5.5 · 75.0%

math

FrontierMath Tiers 1–3 (v2)

115 results
  1. #19xhigh of 11574.4%
  2. #20max of 11574.0%

Leader: GPT-6.1 Sol · 93.7%

math

ProofBench v1.1

50 results
  1. #24max of 5058.0%
  2. #27xhigh of 5055.0%

Leader: Sonnet 5.5 · 100.0%

other

Arena Vision

116 results
  1. #11max of 1161290
  2. #17xhigh of 1161286

Leader: Fable 5 · 1308

reasoning

CritPt

526 results
  1. #37xhigh of 52626.0%
  2. #42max of 52624.9%

Leader: GPT-5.6 Sol · 32.3%

reasoning

Humanity's Sixth Sense

25 results
  1. #7max of 2537.4%

Leader: GPT-6 Astra · 53.6%

reasoning

MysteryMechanism

25 results
  1. #9max of 2536.0%

Leader: GPT-6 Astra · 53.2%

reasoning

Roboflow Visual Reasoning

102 results
  1. #29low of 10273.3%
  2. #31high of 10273.1%

Leader: GPT-6 Astra · 91.2%

reliability

Tool-call error rate

171 results
  1. #170max of 17128.6%

Leader: GPT-5.4 Pro · 0.0%

reliability

Uptime (7d)

181 results
  1. #66max of 18199.8%

Leader: GPT-5 Pro · 100.0%

safety

CWE-Bench

18 results
  1. #6xhigh of 1860.8%

Leader: Grok 4.7 · 68.3%

safety

CyberBench v1.1

43 results
  1. #14max of 4372.7%
  2. #24xhigh of 4369.4%

Leader: GPT-6 Sol · 78.0%

safety

SRE Bench

32 results
  1. #19max of 323.8%

Leader: GPT-6 Astra · 56.9%

speed

Latency (Time to First Token)

355 results
  1. #284xhigh of 35519s
  2. #323max of 35557s

Leader: Gemini 2.5 Flash-Lite · 0.3s

speed

Output Speed (tokens/s)

355 results
  1. #17xhigh of 355258.9
  2. #113max of 355122.5

Leader: Celeris-1 · 1461.1

speed

Provider latency

178 results
  1. #164max of 1784.3s

Leader: Qwen 2.5 7B · 0.2s

speed

Provider throughput

177 results
  1. #81max of 17766.0

Leader: Llama 3.2 3b · 199.5