Dust: Pretraining Transformers Without Backpropagation
2 hours ago
- Dust is a zeroth-order optimization algorithm that perturbs activations (node perturbation) independently at each token, treating each token as a virtual population member for parallel evaluation.
- Dust is competitive with backpropagation for pretraining transformer language models; at large populations it can even exceed backprop, suggesting a compute-rich regime may surpass backprop.
- Dust is orders of magnitude more efficient than weight-space evolution strategies (ES), e.g., 10^3 to 10^4 times more efficient than EGGROLL for transformer pretraining.
- Contrary to conventional wisdom, larger models are more population-efficient with Dust; a 243M-parameter model outperforms a 120× smaller model at most population sizes.
- Dust's gradient estimates align more closely with backprop's as population grows, and this alignment persists up to 1 billion tokens, encouraging scaling.
- The method reduces interference by jittering different layer types in separate forward passes, using credit decay, and evaluating the language modeling head separately.
- Dust's gradients approach but do not match backprop exactly, leading to a different optimization trajectory that can yield better results in some settings.