AnalysisAI ModelsJuly 28, 2026

Kimi Linear: hybrid-linear attention with 256K context

The architecture uses a 5:1 stack of KDA+MLA layers with 1/64 experts per token, native 256K context scalable to 1M. Paper available on arXiv.

2 sources

More stories today

Open the live feed