GLM-5.3 vs GPT-5.6 Sol on DeepSWE: cost, coding, routing

Together AI ran 904 DeepSWE rollouts: GPT-5.6 Sol leads pass@1 by 3.7 points (72.7% vs 69.0%), but GLM-5.3 wins pass@4 (87.6% vs 85.8%) at half the cost ($3.99 vs $8.37 per rollout). A GLM-first cascade solves 85.9% of tasks at $6.61 each.
How this story unfolded
2 days · 1 report · 4 community posts · from Aug 22
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Apple's Luce generates relightable 3D assets from single images
- Claude gets its own browser in Cowork
- Steve Case discusses AI buildout and Nvidia's role
- DHH discusses AI agents, vibe coding, and the future of programming on Lex Fridman Podcast
- Dwarkesh Patel interviews Ryan Greenblatt on AI and politics