Ornith-1.5 open-source LLM family launches, rivals Claude Opus 4.8

Ornith-1.5-397B scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, on par with Claude Opus 4.8 (85.0/59.0) and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731. Trained via end-to-end self-improvement — the model proposes tasks, generates scaffolds, and produces rollouts. The 9B-Mobile runs on phones; the 35B activates only 3B params per token.
6 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills