✳ ProgramBench v1 Fully Resolved897798100
Released 2026-05-05 · Why this score
Released 2026-05-05. Headroom and score-spread quality 77/100 after collapsing effort/config variants by model. Objectivity 98/100 is an assigned judgment score based on grading method. Breadth 100/100 from 60 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher results updated 2026-09-30. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ FrontierMath Tier 4 (v2)95409694
Released 2026-06-12 · Why this score
Released 2026-06-12. Headroom and score-spread quality 40/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Breadth 100/100 from 64 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ AA-Omniscience63508590
Released 2025-11-16 · Why this score
Released 2025-11-16. Headroom and score-spread quality 50/100 after collapsing effort/config variants by model. Objectivity 85/100 is an assigned judgment score based on grading method. Breadth 100/100 from 358 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ Terminal-Bench 4.0100449289
Released 2026-08-01 · Why this score
Released 2026-08-01. Headroom and score-spread quality 44/100 after collapsing effort/config variants by model. Objectivity 92/100 is an assigned judgment score based on grading method. Breadth 57/100 from 23 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ Humanity's Last Exam16618884
Released 2025-01-23 · Why this score
Released 2025-01-23. Headroom and score-spread quality 61/100 after collapsing effort/config variants by model. Objectivity 88/100 is an assigned judgment score based on grading method. Breadth 100/100 from 413 measured models; frontier coverage 100/100 from 10/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ BullshitBench v279426583
Released 2026-03-02 · Why this score
Released 2026-03-02. Headroom and score-spread quality 42/100 after collapsing effort/config variants by model. Objectivity 65/100 is an assigned judgment score based on grading method. Breadth 100/100 from 105 measured models; frontier coverage 80/100 from 8/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ ARC-AGI-383579080
Released 2026-03-25 · Why this score
Released 2026-03-25. Headroom and score-spread quality 57/100 after collapsing effort/config variants by model. Objectivity 90/100 is an assigned judgment score based on grading method. Breadth 48/100 from 19 measured models; frontier coverage 40/100 from 4/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ DeepSWE v1.196379680
Released 2026-06-14 · Why this score
Released 2026-06-14. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Breadth 70/100 from 28 measured models; frontier coverage 30/100 from 3/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗
✳ SimpleBench7378673
Released 2024-11-25 · Why this score
Released 2024-11-25. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 86/100 is an assigned judgment score based on grading method. Breadth 100/100 from 96 measured models; frontier coverage 90/100 from 9/10 current top-lineup models. Publisher update date unknown. Quality combines 15% breadth, 15% frontier coverage, 20% benchmark recency, 30% headroom and spread, and 20% objectivity, plus a preferred-benchmark bonus. It does not prove validity or freedom from training contamination. Source ↗