Policy

Unsealed filings show Microsoft scientist called AI web scraping a theft of labor

Court documents unsealed in the New York Times copyright case against OpenAI and Microsoft include a January 2023 internal memo in which a Microsoft research director described scraping the open web for training data as theft on an unprecedented scale, TechCrunch reported.

TechCrunch reported that the memo was written by Brent Hecht, a director of applied science at Microsoft, who warned that the practice could set off a cycle damaging to news organizations. The filings belong to a suit brought in 2023 by the Times, in litigation that also involves the Daily News and the Center for Investigative Reporting.

According to the same report, the unsealed material describes training sets holding more than 91,000 copies of works by the named plaintiffs, over two million documents drawn from the newspaper's website in one public crawl, and roughly 160,000 unique news works in a separate internal dataset. It also cites remarks attributed to OpenAI executives about the commercial threat chatbots pose to publishers. TechCrunch said neither Microsoft nor OpenAI answered requests for comment. The material consists of filings made by one side in active litigation, and the companies have not conceded the underlying claims.

Source details

Source reporting

Read the original reporting and research behind this briefing.