AnalysisAI ModelsOctober 6, 2026

Six papers probe looped transformers and looped MoE scaling

Read original source →arxiv.org

A wave of arXiv preprints examines recurrence as a scaling axis: looping shared blocks raises effective depth at fixed parameter count, and looped MoE combines recurrence with sparse experts. Papers cover truncated backprop and terminal KV sharing at fixed points, decoding earlier loop states, and FLOPs-matched scaling laws for looped MoE.

How this story unfolded

5 days · 6 reports · 2 community posts · from Oct 1

  1. Oct 1
  2. Oct 2
  3. Oct 5
  4. Oct 6

More stories today

Open the live feed