A study of sequence weighting at scale
a day ago
- Study scaling laws of data weighting across model scales, finding non-monotonic behavior.
- Small models learn general patterns independent of data weight; medium models learn patterns proportional to weight; large models learn all patterns independent of weight.
- Effective sequence weight exponent rises then falls with model scale, with epoching shifting the peak towards smaller models.
- Data mixing exhibits aberrant scaling, complicating extrapolation from small to large models.
- Experiments across three model families (in-house dense, in-house MoE, Qwen 2.5) confirm these trends.