Moonshot AI releases Kimi K3, a 2.8T parameter sparse MoE model
The 2.8T parameter model features a 1M token context window and native vision support, delivering 2.5x better scaling efficiency than Kimi K2. It introduces Kimi Delta Attention and Attention Residuals architectures, achieving top rankings on Agent Arena and full-stack coding benchmarks.
15 sources
Deploying Kimi K3 on AWSaws.amazon.com
audnai/penclaw-Kimi-K3.0-abliterated-GGUFhuggingface.co
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi modelstogether.ai
We've added Kimi K3 to Perplexity and Perplexity Computer for Pro and Max subscribers. Kimi K3 in...x.com
Kimi K3 Architecture Notessebastianraschka.com
moonshotai/Kimi-K3simonwillison.net
What Is Moonshot AI’s Kimi K3 Model and Why Is It Making Waves?bloomberg.com
On Kimi K3: Its Capabilities And Related Discontentsthezvi.substack.com
Moonshot AI by email
Get an email when Moonshot AI has news
No news that day, no email.
More stories today
- GLM-5.3 launch appears imminent, AI commentator predicts
- Recursive AI agents explore questions, synthesize comprehensive answers
- MiniMax-generated The Office scenes draw praise
- OpenCode workflow uses parallel agents for code review and security audits
- 80 skills clone founder, philosopher, scientist minds in coding agents