priors
01 Your calls02 Your ranking03 The evidence

Which is better?

Tap the better model, 6 quick calls at most. We find the benchmarks that agree with you. About a minute · no sign-in · 309 rankings so far

The closest callsCall 01 · up to 6

Your gut calls, backed by benchmarks.

About a minute. No sign-in. · Privacy

Better evidence, fewer benchmarks.

Your preferences guide the search. Quality helps break close calls.

Coverage counts how many models have measured results — a test most models took beats a niche one.

Recency rewards recent test releases. An unknown release date gets a neutral score. A fresh scrape is not a fresh benchmark.

Headroom measures saturation and score spread. Tests that separate models beat tests everyone passes.

Objectivity favors verifiable outcomes over preferences and subjective judging.

Quality combines 25% coverage, 25% recency, 30% headroom, and 20% objectivity — plus a bonus for preferred benchmarks, minus a penalty for superseded versions. Objectivity is an assigned judgment score. This heuristic does not measure contamination or scientific validity.

BenchmarkRecentHeadroomObjectiveQuality
✳ ProgramBench v1 Fully Resolved897798100
Released 2026-05-05 · Why this score

Released 2026-05-05. Headroom and score-spread quality 77/100 after collapsing effort/config variants by model. Objectivity 98/100 is an assigned judgment score based on grading method. Breadth 100/100 from 60 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher results updated 2026-09-30. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ FrontierMath Tier 4 (v2)95409694
Released 2026-06-12 · Why this score

Released 2026-06-12. Headroom and score-spread quality 40/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Breadth 100/100 from 64 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ AA-Omniscience63508590
Released 2025-11-16 · Why this score

Released 2025-11-16. Headroom and score-spread quality 50/100 after collapsing effort/config variants by model. Objectivity 85/100 is an assigned judgment score based on grading method. Breadth 100/100 from 358 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ Terminal-Bench 4.0100449289
Released 2026-08-01 · Why this score

Released 2026-08-01. Headroom and score-spread quality 44/100 after collapsing effort/config variants by model. Objectivity 92/100 is an assigned judgment score based on grading method. Breadth 57/100 from 23 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ Humanity's Last Exam16618884
Released 2025-01-23 · Why this score

Released 2025-01-23. Headroom and score-spread quality 61/100 after collapsing effort/config variants by model. Objectivity 88/100 is an assigned judgment score based on grading method. Breadth 100/100 from 413 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ BullshitBench v279426583
Released 2026-03-02 · Why this score

Released 2026-03-02. Headroom and score-spread quality 42/100 after collapsing effort/config variants by model. Objectivity 65/100 is an assigned judgment score based on grading method. Breadth 100/100 from 105 measured models; frontier coverage 80/100 from 8/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ ARC-AGI-383579080
Released 2026-03-25 · Why this score

Released 2026-03-25. Headroom and score-spread quality 57/100 after collapsing effort/config variants by model. Objectivity 90/100 is an assigned judgment score based on grading method. Breadth 48/100 from 19 measured models; frontier coverage 40/100 from 4/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ DeepSWE v1.196379680
Released 2026-06-14 · Why this score

Released 2026-06-14. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Breadth 70/100 from 28 measured models; frontier coverage 30/100 from 3/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗

✳ SimpleBench7378673
Released 2024-11-25 · Why this score

Released 2024-11-25. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 86/100 is an assigned judgment score based on grading method. Breadth 100/100 from 96 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗