GLM-5.3 matches frontier models on DeepSWE at fraction of cost

Together AI ran 904 DeepSWE rollouts: GLM-5.3 ties Claude Fable 5 on pass@1 (69.0% vs 69.7%) and leads pass@4 (87.6% vs 84.1%), at $3.99 per rollout vs $21.63. Against GPT-5.6 Sol, GLM-5.3 trails pass@1 by 3.7 points but wins pass@4 at half the cost.
How this story unfolded
2 days · 2 reports · 4 community posts · from Aug 22
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- ChatGPT users complain it forces local language despite English setting
- New methods steer recurrent reasoners at inference time
- Tencent releases WeMM-Embedding multimodal embedding models
- FireRedTeam releases FireRedAudio and FireRedTTS3
- Mindrank AI CEO discusses AI drug discovery strategy