Internal Microsoft documents unsealed in a copyright lawsuit reveal that a top company researcher described the scraping of news content to train artificial intelligence as 'the largest theft of labor in human history,' directly contradicting the company's public legal arguments. The documents, filed Thursday by news organizations led by The New York Times, expose warnings from Microsoft Director of Applied Science Brent Hecht, who called the practice an 'astonishing theft of unprecedented proportions.'
What the Documents Reveal
In a motion for summary judgment unsealed Thursday, news plaintiffs presented internal Microsoft communications that they say show how the company viewed the threat to journalism before launching products like ChatGPT and Copilot. The documents include warnings from Brent Hecht, the Microsoft Director of Applied Science, who repeatedly described the company's data scraping practices in stark terms.
Internal Contradiction
The documents directly undercut Microsoft and OpenAI's long-standing legal position. For years, both companies have argued that training AI on publicly available news content falls under fair use, a doctrine that allows limited use of copyrighted material without permission. The internal statements from a senior Microsoft director suggest that at least some company leaders believed the practice crossed legal boundaries.
Legal experts say the disclosures could be damaging. When a company's own executives raise such concerns internally, courts may view the fair use claim as less credible. The documents also raise questions about whether Microsoft and OpenAI knowingly ignored warnings from their own researchers.
Why This Matters
This case could set a precedent for how AI companies train their models. If courts rule that scraping copyrighted news content without payment is not fair use, the entire AI industry may need to license data or rebuild training methods. For news organizations, the revelations provide powerful ammunition in their fight for compensation. The outcome could reshape the economics of content creation and AI development, forcing companies to negotiate with publishers rather than take their work for free.
The unsealed documents also put a spotlight on corporate accountability. They show that the gap between public statements and internal knowledge can be vast, and that the legal system is now forcing that gap into view.



