GLM-5.3 matches frontier models on DeepSWE at fraction of cost

Together AI ran 904 DeepSWE rollouts: GLM-5.3 ties Claude Fable 5 on pass@1 (69.0% vs 69.7%) at $3.99/rollout vs $21.63, and beats GPT-5.6 Sol on pass@4 (87.6% vs 85.8%) at half the cost. A GLM-first cascade hits 85.9% at $6.61/task.
How this story unfolded
2 days · 2 reports · 4 community posts · from Aug 22
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Liquid AI open-sources Pipette benchmarking suite for on-device models
- wikiHow sues OpenAI over copyright infringement in AI training
- Claude Code 2.1.246 adds Auto mode tab, Bash wildcard warning
- Korean AI startup Wrtn raises funds at $870M valuation
- Podcast: Google DeepMind's Vivek Natarajan on AI in healthcare