Hasty Briefsbeta

Bilingual

>10x More Efficient Pretraining

22 days ago
  • Achieved over 10x compute-efficient pretraining compared to leading open-weight base models, matching DeepSeek V4 Pro Base using ~50x fewer FLOPs.
  • Scaled pretraining to outperform all publicly available open base models on perplexity, with training cost under $4M versus estimated >$100M.
  • Focused on algorithmic efficiency due to lack of large chip clusters, relying on changes in architecture, optimizer, training objective, and data curation.
  • Evaluated model performance using bits-per-byte loss on heldout data, fitting scaling laws to project compute requirements.
  • Measured generalization on private codebases, reasoning problems, and recent research papers, with careful decontamination.
  • Assessed domain knowledge by rewording documents with third-party LLMs to avoid rewarding memorization.
  • Built a stable training foundation with smooth convergence, low-precision training, and bug hunting before scaling.
  • Used multi-scale evaluations (three models spanning two orders of magnitude compute) to validate improvements at scale.
  • Plans to focus on long-horizon reinforcement learning, alignment training, and further pretraining improvements.
  • Claims to be the smallest team training trillion-parameter models and invites talent to join.