AnalysisAI ModelsJuly 29, 2026
Claude Opus 5, Kimi K3, Grok 4.5, Gemini 3.6 Flash benchmarked on Baba Is You

An open-source benchmark evaluates four recent large language models on the puzzle game Baba Is You, testing their reasoning capabilities. The benchmark, baba-is-harbor, compares Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- PortSwigger explains safety design for agentic pentesting
- Kimi K3 distillation into Laguna 2.1 requested
- Cursor and Anthropic launch localized India pricing plans
- SpaceXAI releases Grok Voice Think Fast 2.0
- Sam Altman to brief White House on OpenAI's next AI model