Models trade world knowledge for reasoning, study finds
GLM-5.2 scores 99.2% on AIME 2026 with ~40B active parameters, while Qwen3.5 9B fits in 6GB VRAM and doubles the next best sub-10B model. Yet SimpleQA leader Gemini 2.5 Pro still misses half of factual questions, and small models hallucinate 80-82% of the time.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Redditor tests AI agents with $1 online task
- AI agents need their own identity before a gateway
- Claude Max users find default $200K spend limit
- TTFT-First Benchmark Ranks Lowest-Latency Voice and Realtime Agent APIs
- AI training demand causes Mac Mini shortages