Models Are Getting Dumber on Purpose
GLM-5.2 scores 99.2% on AIME 2026 with ~40B active parameters per token, while SimpleQA leader Gemini 2.5 Pro recalls only 53% of facts. Essay argues labs deliberately trade world knowledge for reasoning: Qwen3.5 4B/9B hallucinate at 80-82% on Artificial Analysis's knowledge benchmark.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills