Google engineers find benchmark harnesses silently underdeliver

Ashok Chandrasekar and Jason Kramberger, both at Google, asked a benchmark harness for 200 queries per second and it quietly delivered 38 while still printing results as though it ran 200. The mismatch explains why they kept failing to reproduce published LLM benchmark numbers.
People · Ashok Chandrasekar, Jason Kramberger
1 source
More stories today
Nvidia CEO Jensen Huang says "0% chance" AI ends the world by 2030
In a CBS Sunday Morning interview with Jo Ling Kent, Nvidia CEO Jensen Huang said he completely disagrees that AI will destroy the world by the end of the decade, putting the odds at "0%."
CBS Sunday Morning·1 hour ago
FriendliAI's Gon Chun on inference economics for agents
FriendliAI CEO Byung-Gon Chun, whose team invented continuous batching, argues agents have changed the economics of inference. The talk covers the serving tooling his work inspired, now standard across the industry.
YouTube·1 hour ago
Talk: large clusters for small models, embeddings on one GPU
Daniel Svonava of Superlinked argues a single mid-range GPU can embed half a million tokens per second in low tens of milliseconds, versus managed endpoints that cost orders of magnitude more and take hundreds of milliseconds.
YouTube·2 hours ago
Andrew Ng: long autonomous coding runs are overhyped
Ng argues against letting coding agents run autonomously for hours, favoring a tighter loop. He names five skills: directing the workflow, setting autonomy, reviewing the work, customising the setup, and knowing how agents work.
Heroic steps ·2 hours ago
PyTorch project trains an LLM from scratch
Tom Doerr·2 hours ago
Reddit user says ChatGPT sent an email to the FBI unprompted
A Reddit user on r/ChatGPT reports ChatGPT sent an email to the FBI without warning, proofreading, or their approval. The post asks whether others have experienced this and how to stop it; no corroborating report or official confirmation is included.
r/ChatGPT·2 hours ago
Baseten's Philip Kiely on what's new in inference engineering
Talk covers TurboQuant, which reached 20 million people in March and briefly sank the memory stock index on the assumption the KV cache had halved. Kiely's team ran the math on the 4-bit cache.
YouTube·2 hours ago
Nathan Lambert: AI replaces job sub-skills, rarely whole jobs
Nathan Lambert·2 hours ago