GLM-5.3 Flash: 17x cheaper, 5.6 points behind on DeepSWE

Together AI ran 900 DeepSWE rollouts: GLM-5.3 Flash trails pass@1 by 5.6 points (63.4% vs 69.0%) but costs $0.24 vs $3.99 per rollout. A cascade starting with Flash solves 80.9% of tasks at $1.70 each.
How this story unfolded
3 days · 3 reports · 4 community posts · from Aug 29
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- llama.cpp adds fixes for Qwen Flash Next
- OpenAI's Cursor Ban Is About Astra
- Podcast explores AI's progress in mathematical intuition
- Matt Wolfe builds SaaS dashboard with one ChatGPT prompt
- HuggingFace attack postmortem: OpenAI agents hacked platform