AnalysisAI ModelsAugust 6, 2026

Qwen 3.8-Max and Claude Opus 5 show raw benchmark scores don't predict cost

Alibaba's Qwen 3.8-Max preview was marketed as second only to Claude Fable 5, but an independent harness run suggested the opposite on coding-agent tasks. The analysis argues raw benchmark scores don't predict real-world cost or performance.

1 source

More stories today

Open the live feed