Hasty Briefsbeta

Bilingual

A study of sequence weighting at scale

a day ago
  • Study scaling laws of data weighting across model scales, finding non-monotonic behavior.
  • Small models learn general patterns independent of data weight; medium models learn patterns proportional to weight; large models learn all patterns independent of weight.
  • Effective sequence weight exponent rises then falls with model scale, with epoching shifting the peak towards smaller models.
  • Data mixing exhibits aberrant scaling, complicating extrapolation from small to large models.
  • Experiments across three model families (in-house dense, in-house MoE, Qwen 2.5) confirm these trends.