LaunchDevelopersSeptember 18, 2026

Flyweight ships C++/CUDA engine to run MoE models beyond VRAM

Flyweight is a native GGUF inference runtime with OpenAI/Anthropic-compatible APIs and a chat UI, built to run MoE models that don't fit in VRAM using one consumer NVIDIA card plus system RAM. First PyPI release; author is seeking contributors.

1 source

More stories today

Open the live feed