Qwen 3.8-Max and Claude Opus 5: benchmark scores don't predict cost

Alibaba's Qwen 3.8-Max, marketed as second only to Claude Fable 5, led on just one of 12 coding-agent rows in an independent harness, showing raw benchmark scores don't predict real-world cost.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Agentic Cloud concept introduced
- OpenAI's Codex client reveals GenUI interface platform
- Weaviate 1.39 adds Boost API, MMR, 4-bit quantization
- JPMorgan leads $5B debt package for Volta AI data centers
- Vercel Chat SDK adds Claude Managed Agents and Notion adapter