Hasty Briefsbeta

Bilingual

Dust: Pretraining Transformers Without Backpropagation

2 hours ago
  • Dust is a zeroth-order optimization algorithm that perturbs activations (node perturbation) independently at each token, treating each token as a virtual population member for parallel evaluation.
  • Dust is competitive with backpropagation for pretraining transformer language models; at large populations it can even exceed backprop, suggesting a compute-rich regime may surpass backprop.
  • Dust is orders of magnitude more efficient than weight-space evolution strategies (ES), e.g., 10^3 to 10^4 times more efficient than EGGROLL for transformer pretraining.
  • Contrary to conventional wisdom, larger models are more population-efficient with Dust; a 243M-parameter model outperforms a 120× smaller model at most population sizes.
  • Dust's gradient estimates align more closely with backprop's as population grows, and this alignment persists up to 1 billion tokens, encouraging scaling.
  • The method reduces interference by jittering different layer types in separate forward passes, using credit decay, and evaluating the language modeling head separately.
  • Dust's gradients approach but do not match backprop exactly, leading to a different optimization trajectory that can yield better results in some settings.