AnalysisAI ModelsSeptember 12, 2026

Anthropic's 2021 transformer circuits paper resurfaces

The paper studies transformers with two layers or fewer and only attention blocks, versus GPT-3's 96 layers alternating attention with MLP blocks. It identifies "induction heads" that explain in-context learning and only develop in models with at least two attention layers.

1 source

More stories today

Open the live feed