In the ongoing legal battle over the use of copyrighted material to train and operate AI systems, Microsoft has submitted evidence it says shows its Copilot tool does not routinely reproduce protected content. The company provided 8.2 million chat logs to an expert hired by the news publishers, selected specifically because they involved keywords related to the plaintiffs' websites. Microsoft says the resulting analysis found that only 59,545 of those logs contained any text from the plaintiffs' works, and that even those instances rarely involved full sentences or substantial excerpts that could act as a substitute for the original. The lawsuit, filed by The New York Times and later joined by a group of authors, alleges that Microsoft and OpenAI used copyrighted material without permission in developing their AI products. Microsoft contends the evidence undermines claims of widespread infringement.
Microsoft Says Copilot Rarely Reproduces Copyrighted Content in Legal Filing
Company provides 8.2 million chat logs to publishers' expert, pointing to low rate of matching
Analysis
Why This Matters
- The outcome could set a precedent for how AI companies use copyrighted material for training and output generation, affecting publishers, authors, and tech firms alike.
- Microsoft's data-driven argument challenges the plaintiffs' claims of systematic infringement, potentially shifting the burden of proof.
- The case is one of several high-profile copyright lawsuits against AI developers, and a ruling could influence future litigation and licensing practices.
Background
Microsoft and OpenAI face multiple copyright lawsuits from news publishers and authors who allege that their AI models were trained on copyrighted works without permission and that they generate outputs that reproduce protected content. The New York Times sued in December 2023, and a group of authors added Microsoft to their existing suit against OpenAI. Microsoft has consistently denied wrongdoing, arguing that its AI tools operate within fair use boundaries and that they do not meaningfully reproduce third-party content. The recent filing is part of the discovery phase, where both sides submit evidence.
Key Perspectives
Microsoft: The company argues that the chat log data shows the incidence of reproduced text is minuscule (0.7% of the targeted logs) and that the outputs are not substitutes for the original works. It frames the suit as an overreach that could stifle innovation. News Publishers and Authors (Plaintiffs): They maintain that even occasional reproduction of their content without permission constitutes infringement, and that the training process itself violated copyright. They are likely to scrutinize Microsoft's methodology for selecting logs and challenge the low figure. Critics/Skeptics: Some legal experts question whether the chat log sample is representative, noting that it was chosen by Microsoft based on keywords, which could bias the results. Others argue that the case may hinge on the broader question of fair use rather than the volume of matches.
What to Watch
- Further expert analysis of the chat logs by the plaintiffs, which could yield different conclusions.
- The court's ruling on motions for summary judgment, which could narrow or expand the scope of the case.
- Any settlement discussions, as both sides may seek to avoid a lengthy trial with uncertain outcomes.