Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes
20 days ago
- A single developer trained a 3.8B-parameter language model to a CORE score of 0.384 in 43 hours for $998, using rented B200 GPUs.
- The model architecture is Llama-style with key additions: ResFormer value embeddings, QK-norm, logit softcap, and relu² non-gated MLPs.
- Key optimizations included FP8 training, trapezoidal learning rate schedule, Muon optimizer for matrices, and ClimbMix dataset, which dramatically improved convergence.
- Training at 2048-token context significantly boosted CORE scores (0.384 vs 0.338 at 1024), mainly by enabling proper evaluation on long-context tasks like SQuAD and BoolQ.
- Value embeddings added 19% more parameters with negligible throughput cost, yielding ~5% effective training gain without requiring additional compute.
- The author emphasizes the value of clean infrastructure, local data sharding, and avoiding misleading micro-benchmarks for efficient training on a limited budget.