Hasty Briefsbeta

Bilingual

RL economics, morally charged terms, and "distillation"

a day ago
  • Advancements in LLM capabilities, particularly in coding and mathematics, are driven by reinforcement learning due to the exhaustion of human-curated data.
  • Reinforcement learning involves generating multiple solutions to problems, using successful ones as training data, which is computationally expensive for pioneers but cheaper for followers.
  • Model weights are not copyrightable, but prompts and outputs may be copyrightable to users, not model providers; terms of service govern usage, not copyright law.
  • Model providers cannot legally prevent third parties from using published model outputs as training data, even if it benefits competitors.
  • The term 'distillation' is misused to morally frame training on model outputs as 'attacks,' but it should be called 'training on model output' to avoid moral bias.
  • The current legal system allows training on publicly available model outputs, and model providers' business concerns do not justify moral claims against such practices.