AnalysisAI ModelsJuly 26, 2026
DeepSWE: Kimi K3 tops GPT-5.6 Sol on pass@4 and cost

GPT-5.6 Sol leads pass@1 72.7% to 68.5%, but Kimi K3 wins pass@4 89.4% vs 85.8% while costing $4.65 per rollout to Sol's $8.37 — 2.8x more solved tasks per dollar. Routing between the two models reaches ~85.6% across 113 DeepSWE tasks.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Daniel Kokotajlo and AI Futures Project release AI 2040: Plan A
- LearnHouse open-source LMS features AI integration and whiteboards
- Practical guide to Claude Code custom slash commands
- Reddit debate over xAI's Grok 4.5 vs Meta's AI spending
- Aravind Srinivas: two orders of magnitude improvements 'a big deal'