AnalysisAI ModelsOctober 5, 2026

Dust pretrains transformers without backpropagation

Read original source →qlabs.sh

Dust perturbs activations per token, making each token a virtual population member so one forward pass evaluates all in parallel. It is 10^3–10^4x more efficient than EGGROLL, a state-of-the-art ES method, from 1M tokens up, and a 243M-parameter model outperforms a 120x smaller one at most population sizes.

1 source

More stories today

Open the live feed