AnalysisAI ModelsJuly 29, 2026

Claude Opus 5, Kimi K3, Grok 4.5, Gemini 3.6 Flash benchmarked on Baba Is You

An open-source benchmark evaluates four recent large language models on the puzzle game Baba Is You, testing their reasoning capabilities. The benchmark, baba-is-harbor, compares Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed