AnalysisAI ModelsJuly 28, 2026
Kimi Linear: hybrid-linear attention with 256K context

The architecture uses a 5:1 stack of KDA+MLA layers with 1/64 experts per token, native 256K context scalable to 1M. Paper available on arXiv.
2 sources
More stories today
- Hybe liquidates AI startup Supertone after $32M acquisition
- Perplexity adds Kimi K3 for Pro and Max subscribers
- Lecture 10 on RL regularization: KL penalty role
- Cohere talk explores the 'death of NLP' in the LLM era
- Google's higher AI capex estimate spooks Wall Street