Prime Intellect benchmarks 18 frontier models on nanoGPT optimizer

The study evaluated 153 autonomous runs across 18 models, with Fable achieving the top validated result of 2,726. The benchmark measured performance on the nanoGPT optimizer speedrun, with Claude-code, Codex, and Kimi-code agents used for execution.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Coding agents generate interactive slide decks from prompts
- Krea 2 / Anima LoRA recreates 90s retro anime style
- Homelab cluster grows from 16 to 36 DGX Sparks with 4.6TB unified memory
- Chollet: AI slop and bots dominate social media
- llama.cpp fork optimizes AMD GFX906 GPUs, doubling prompt speeds