AnalysisAI ModelsSeptember 12, 2026

Dan Luu critiques Senior SWE-Bench, napkin math, and tire benchmarks

Dan Luu's post examines three benchmarks: the napkin-math README (5.4k stars) listing random memory R/W at 20ns, DeepSWE and Senior SWE-Bench model evals, and winter-tire claims. He argues each is flawed as a benchmark.

1 source

More stories today

Open the live feed