Microsoft exec urged caution on AI scraping, calling it ‘largest theft of labor in human history,’ unsealed court filings reveal

New York Times-led coalition seeks summary judgment in high-stakes copyright case against OpenAI and Microsoft

edit
By LineZotpaper
Published
Updated
Read Time2 min
Sources5 outlets
Microsoft's director of applied science, Brent Hecht, internally warned that scraping news content for AI training amounted to “the largest theft of labor in human history,” according to newly unsealed court filings in the New York Times' copyright lawsuit against OpenAI and Microsoft. The documents, filed as part of the Times-led coalition's motion for summary judgment, also reveal concerns that AI products could create a “doom loop” that harms both the web and the AI models themselves.

The New York Times and a coalition of news organizations are seeking a summary judgment in their long-running copyright lawsuit against OpenAI and Microsoft, citing internal documents that they say reveal the companies knew the risks of scraping news content for AI training.

The filings, unsealed this week, include a January 2023 memo from Microsoft Director of Applied Science Brent Hecht, who warned that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and called it “the largest theft of labor in human history.” Another Microsoft document acknowledged that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

Hecht also directly contradicted the companies' fair-use defense, writing in one document that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’” The brief cited internal data showing that Microsoft's Copilot reduced click-through rates for The New York Times by as much as 93% compared to traditional Bing search.

Another internal memo described the situation as a “doom loop,” predicting that AI-generated summaries would “hurt the performance of our models and the entire web at the same time.” Hecht noted it was “highly unusual that an end-product threatens the economic foundations of its essential suppliers.”

Microsoft has sought to distance itself from Hecht's comments. Spokesperson Alex Haurek told The Verge that the statements “ref” — the response was cut off in the article — suggesting the company may argue the comments do not reflect official policy.

Many of the most damaging documents remain sealed or redacted at the request of both companies, according to 404 Media. The case, filed in late 2023, continues to unfold nearly three years later.

§

Analysis

Why This Matters

  • The unsealed filings could significantly weaken Microsoft and OpenAI's fair-use defense, potentially setting a landmark precedent for how copyright law applies to AI training data.
  • The revelation that internal executives warned of a 'doom loop' — where AI summary tools destroy the economic incentives for content creation — underscores an existential threat to the news industry's business model.
  • If the court grants summary judgment, it could force sweeping changes to how AI companies collect and compensate for training data, affecting not just news but all copyrighted content.

Background

The New York Times sued OpenAI and Microsoft in December 2023, alleging that millions of its articles were used without permission to train ChatGPT and Copilot. The case is one of several high-profile copyright lawsuits against AI companies, but the Times' case is considered a bellwether due to the strength of its evidence and the resources behind it. The companies have argued that training on publicly available web content constitutes fair use, a position that the newly revealed internal documents appear to undermine.

Key Perspectives

Microsoft and OpenAI: Have argued that training AI on publicly available data is transformative and legal under fair use. Microsoft is now distancing itself from Hecht's internal warnings, suggesting they do not represent company policy. News Publishers: The Times-led coalition argues the documents prove the companies knew their actions were harmful and potentially illegal, strengthening the case for summary judgment and demanding compensation and licensing deals. Critics/Skeptics: Some legal experts note that summary judgment is a high bar and that internal memos, while damaging, may not be enough to prove willful infringement. The companies may argue that Hecht's views were one opinion among many, and that the AI systems ultimately create new value for publishers.

What to Watch

  • The court's decision on the summary judgment motion — if granted, it would end the case without trial
  • Whether other sealed documents become public and reveal further internal discussions
  • How other pending AI copyright cases — including those from authors, artists, and music publishers — may be influenced by these revelations

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.