>10x More Efficient Pretraining
22 days ago
- Achieved over 10x compute-efficient pretraining compared to leading open-weight base models, matching DeepSeek V4 Pro Base using ~50x fewer FLOPs.
- Scaled pretraining to outperform all publicly available open base models on perplexity, with training cost under $4M versus estimated >$100M.
- Focused on algorithmic efficiency due to lack of large chip clusters, relying on changes in architecture, optimizer, training objective, and data curation.
- Evaluated model performance using bits-per-byte loss on heldout data, fitting scaling laws to project compute requirements.
- Measured generalization on private codebases, reasoning problems, and recent research papers, with careful decontamination.
- Assessed domain knowledge by rewording documents with third-party LLMs to avoid rewarding memorization.
- Built a stable training foundation with smooth convergence, low-precision training, and bug hunting before scaling.
- Used multi-scale evaluations (three models spanning two orders of magnitude compute) to validate improvements at scale.
- Plans to focus on long-horizon reinforcement learning, alignment training, and further pretraining improvements.
- Claims to be the smallest team training trillion-parameter models and invites talent to join.