AnalysisAI ModelsAugust 6, 2026

DeepSeek-V4 Flash 0731 outperforms GPT-5.6 Luna in cost-efficiency

DeepSeek-V4 Flash 0731 costs $0.10 per task on DeepSWE, delivering 532 solves per $100 compared to 110 for GPT-5.6 Luna. While Luna leads in pass@1 accuracy at 67.2% versus 53.3%, a cascaded approach using both models achieves 78.9% accuracy at 37% lower cost than Luna alone.

3 sources

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed