Moonshot AI releases Kimi K3, a 2.8T parameter sparse MoE model
The 2.8T parameter model features native vision support, a 1M token context window, and 2.5x better scaling efficiency than Kimi K2. It introduces Kimi Delta Attention and Attention Residuals architectures, and is now available on Together AI, AWS, and Nebius.
15 sources
Deploying Kimi K3 on AWSaws.amazon.com
audnai/penclaw-Kimi-K3.0-abliterated-GGUFhuggingface.co
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi modelstogether.ai
We've added Kimi K3 to Perplexity and Perplexity Computer for Pro and Max subscribers. Kimi K3 in...x.com
Kimi K3 Architecture Notessebastianraschka.com
moonshotai/Kimi-K3simonwillison.net
What Is Moonshot AI’s Kimi K3 Model and Why Is It Making Waves?bloomberg.com
On Kimi K3: Its Capabilities And Related Discontentsthezvi.substack.com
Moonshot AI by email
Get an email when Moonshot AI has news
No news that day, no email.
More stories today
- Cogent AI releases VR-1 cyber reasoning model
- Orchestrator tool integrates 12 AI coding agents in Visual Studio Code
- Taste Skill rules file for AI coding agents crosses 70,000 GitHub stars
- Hugging Face Diffusers flaws allow arbitrary code execution
- 139 Agent Skills bring legal workflows to Claude, Codex and Gemini CLI