Kimi K3 vs GPT-5.6 Sol on DeepSWE: Kimi wins pass@4, Sol takes pass@1

Across 904 rollouts on DeepSWE's 113 tasks, GPT-5.6 Sol leads pass@1 at 72.7% vs 68.5%, but Kimi K3 wins pass@4 (89.4% vs 85.8%) at $4.65 per rollout to Sol's $8.37, delivering 2.8x solves per dollar. A Kimi-first cascade with test-suite verification reaches ~85.6%, covering 108 of 113 tasks.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude detects failing drive, saves user's data
- AI and satellite guidance could bring robot mowers to half of US lawns
- OpenAI acquires Instant backend team
- China unveils AI-powered flying lifebuoy for water rescues
- Apollo's Slok: AI weighs on pay without cutting jobs