Revealed: The Authors Whose Pirated Books Are Powering Generative AI (2023)
20 hours ago
- Generative AI systems like ChatGPT are trained on vast amounts of text, including pirated books, without public transparency.
- The Books3 dataset, containing over 170,000 copyrighted books (e.g., by Stephen King, Zadie Smith), was used to train Meta's LLaMA, BloombergGPT, and EleutherAI's GPT-J.
- Meta and other companies face lawsuits for copyright infringement, claiming fair use, but the unauthorized source of the books may weaken that defense.
- The dataset was created by developer Shawn Presser to democratize AI training, but it relies on pirated content from Bibliotik and other sources.
- The article highlights a clash between open-source culture, which prioritizes free access, and the need for copyright protection to support authors' livelihoods.