Microsoft exec called AI scraping 'the largest theft of labor in human
9 hours ago
- Microsoft executive privately described AI training practices as 'theft' and OpenAI leadership said models pose an 'existential threat' to publishers.
- Companies allegedly bypassed paywalls, mass-scraped content, and stripped copyright notices from training data.
- New York Times lawsuit claims copyright violation; fair use defense is challenged by admissions of market harm.
- Microsoft's Copilot caused up to 93% drop in click-through rates for NYT, described as a 'doom loop'.
- Microsoft CEO Nadella stated paywalled content must be licensed and would have required retraining if aware of scraping.
- OpenAI's head of ChatGPT noted chatbots are 'largely substitutive' and pose existential threat to publishers.
- Documents reveal massive scale: over 91,692 copies of plaintiff works in training datasets and 2 million documents from nytimes.com.
- Internal memo called it 'the largest theft of labor in human history'.
- OpenAI and Microsoft exchanged training data via projects like Taxi and Mango, including 160,903 works.
- Employees planned to circumvent paywalls without detection; intentionally removed copyright notices from training data.