GLM-5.3 matches GPT-5.6 Sol on DeepSWE at half the cost

Together AI ran 904 DeepSWE rollouts: GLM-5.3 trails GPT-5.6 Sol by 3.7 points on pass@1 (69.0% vs 72.7%) but leads pass@4 (87.6% vs 85.8%) at $3.99 per rollout vs $8.37. A GLM-first cascade solves 85.9% of tasks at $6.61 each.
How this story unfolded
2 days · 1 report · 5 community posts · from Aug 21
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- ChatGPT user reports 5-hour limit consumed by single prompt
- Google's Gemini 3.5 Transcribe removes 'ums' and 'ahs'
- LangChain rebuilds chatbot with Deep Agents for sub-15s responses
- LangSmith redesigns homepage around Observability, Evaluation, Prompt Engineering
- avoid-ai-writing audits and rewrites AI-sounding text