AnalysisAI ModelsAugust 6, 2026

Qwen 3.8-Max and Claude Opus 5 show raw benchmark scores don't predict cost

Alibaba released Qwen 3.8-Max, marketed as second only to Claude Fable 5, but an independent harness suggested the opposite on coding-agent tasks. The analysis argues raw benchmark scores don't predict real-world cost.

1 source

More stories today

Open the live feed