Reddit users question METR's benchmark inactivity
A Reddit post criticizes METR for abandoning its benchmark after Opus 4.6, despite receiving millions in investments. The user argues METR could still benchmark current models for 80/90/99% accuracy with its existing task set.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Abliterlitics compares 12 abliterated Gemma 4 12B variants
- ElevenLabs launches Composer section-by-section song editor in ElevenMusic
- Perplexity CEO thanks Nvidia for DGX Spark research support
- LangChain cookbook builds due-diligence agent with Deep Agents and Parallel
- LangSmith and LangChain OSS help meet EU AI Act requirements