Hasty Briefsbeta

Bilingual

Revealed: The Authors Whose Pirated Books Are Powering Generative AI (2023)

20 hours ago
  • Generative AI systems like ChatGPT are trained on vast amounts of text, including pirated books, without public transparency.
  • The Books3 dataset, containing over 170,000 copyrighted books (e.g., by Stephen King, Zadie Smith), was used to train Meta's LLaMA, BloombergGPT, and EleutherAI's GPT-J.
  • Meta and other companies face lawsuits for copyright infringement, claiming fair use, but the unauthorized source of the books may weaken that defense.
  • The dataset was created by developer Shawn Presser to democratize AI training, but it relies on pirated content from Bibliotik and other sources.
  • The article highlights a clash between open-source culture, which prioritizes free access, and the need for copyright protection to support authors' livelihoods.