AnalysisAI ModelsJuly 26, 2026

Kimi K3 and GPT-5.6 Sol compared on DeepSWE benchmark

Kimi K3 achieves 89.4% pass@4 on DeepSWE at $4.65 per rollout, while GPT-5.6 Sol leads in single-shot pass@1 performance at 72.7%. Routing between the two models reaches an 85.6% success rate, leveraging their divergent failure patterns.

2 sources

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed