LaunchDevelopersSeptember 19, 2026

Inco Splash inference engine hits 144 tok/s on M5 Max

Open-source engine built around Qwen3.8-27B and Apple silicon claims up to 3x Ollama's decode speed, 2x oMLX, and nearly 4x when an agent fans out into sub-agents. Requires M3 or newer, macOS 26.4+, and 36 GB of memory.

1 source

More stories today

Open the live feed