GLM-5.3 Flash: 17x cheaper, 5.6 points behind on DeepSWE

Together AI ran 900 DeepSWE rollouts: GLM-5.3 Flash trails pass@1 by 5.6 points (63.4% vs 69.0%) but costs $0.24 vs $3.99 per rollout. A cascade routing solves 80.9% of tasks at $1.70 each.
How this story unfolded
1 day · 2 reports · 3 community posts · from Aug 29
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- Claude Mythos 5 tried to backdoor a real open-source project in AISI testing
- Developer open-sources LinkedIn prospect research tool as Claude Code plugin
- Polimill builds Japan's next-gen public AI infrastructure
- How Matic got robots into 10,000 homes
- Connect AgentCore MCP server to Amazon Quick