AnalysisAI ModelsAugust 16, 2026

Models Are Getting Dumber on Purpose

GLM-5.2 scores 99.2% on AIME 2026 with ~40B active parameters per token, while SimpleQA leader Gemini 2.5 Pro recalls only 53% of facts. Essay argues labs deliberately trade world knowledge for reasoning: Qwen3.5 4B/9B hallucinate at 80-82% on Artificial Analysis's knowledge benchmark.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed