French startup Kog promises 30x faster LLM inference on standard GPUs

Kog's demo hit 3,000 tokens/sec single-request decoding on AMD MI300X and Nvidia H200 GPUs, toward its promised 30x faster inference. CEO Gaël Delalleau says 200 business leads came in, with software engineering the likely first use case. Since customers won't fine-tune small models, the startup now targets larger ones.
Featured · Gaël Delalleau
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs