Microsoft is arguing in new court filings that its Copilot chatbot rarely reproduces copyrighted material in a way that could substitute for the original work. The company says even full sentences from news articles and books appear infrequently, and that larger, more substantive excerpts are rarer still.
As part of discovery in lawsuits involving publishers including The New York Times and book authors, Microsoft provided 8.2 million Copilot chat logs to an expert hired by the news plaintiffs. Notably, the company says those logs were chosen because they were especially likely to involve publisher-related keywords, making them a high-risk sample for finding potential reproductions.
Why this matters
The positive signal is practical: real-world usage data may help courts, publishers, authors, and AI companies better understand how often chatbots actually output protected material verbatim. That kind of evidence can move the debate beyond hypotheticals and toward clearer rules for safe, useful AI deployment.
- Large-scale log analysis can inform copyright policy with concrete data.
- Low rates of verbatim reproduction would support AI tools that assist users without replacing original reporting or books.
- The outcome may encourage stronger safeguards, better licensing discussions, and more transparent standards across the industry.
The legal fight is far from over, and Microsoft’s claims will be tested in court. Still, the filing highlights an important step for AI accountability: measuring actual system behavior at scale and using that evidence to improve how AI products coexist with creators and publishers.