Mimo v2.5 outperforms DeepSeek V4 Flash on Terminal Bench
Mimo v2.5 scored 55% on Terminal Bench v2.0 with the Hermes harness, beating DeepSeek V4 Flash and others that scored under 50%, according to a r/LocalLLaMA user's benchmarks. Tests ran via the OpenCode endpoint across Codex, Hermes, and Pi harnesses.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- CPUs and the rise of neurosymbolic AI
- Mollick asks for plan on cyber threats from open-weight Mythos/Astra models
- Dwarkesh Patel discusses the implications of continual learning for AI
- NoimosAI launches Social Agent, an AI research agent for social trends
- Stanford researchers deploy 37,000 AI agents as a virtual biotech