Microsoft Says Copilot Almost Never Copies News Articles Word for Word. Publishers Aren't Convinced.

New court filings show Microsoft analysed 8.2 million chat logs to prove its AI assistant rarely reproduces copyrighted text. The legal battle with The New York Times and book authors is far from over.

AI2Day Newsdesk3 min read
A large server room bathed in cold blue emergency lighting, rows of inactive server racks with dark indicator panels, a single red warning light reflected acros
Share

Key points

  • Microsoft handed 8.2 million Copilot chat logs to publishers' legal experts as part of court-ordered discovery.
  • Of those logs, 59,545 contained at least 16 words matching news content; only 24 contained 30 or more matching words from books.
  • Microsoft argues these figures support a "fair use" defence, meaning AI training on copyrighted material should be treated like research or commentary, not theft.
  • The case involves claims by The New York Times, the Center for Investigative Reporting, and book authors that Microsoft and OpenAI built profitable products on their work.
  • Microsoft filed for summary judgement on Friday, asking the judge to end the case early before it goes to trial.

Microsoft wants a federal judge to close a major copyright lawsuit before it ever reaches a jury. To do that, the company is leaning hard on its own chat records.

As part of the ongoing legal fight with news publishers and book authors, Microsoft turned over 8.2 million conversation logs from Copilot, its AI assistant, to an expert hired by the plaintiffs. The company says those logs were deliberately selected because they were the ones most likely to contain the publishers' material, picked by searching for keywords tied to the news outlets involved.

What did the chat logs actually show?

Far less copying than the publishers allege, according to Microsoft. Of the 8.2 million conversations, 59,545 included at least 16 words in a row that matched news content used to train the model. For books, the number was smaller still: only 24 responses across all 8.2 million chats contained 30 or more matching words, and only 10 of the 212 books examined had any matches at all.

An expert working for the Center for Investigative Reporting, one of the news outlets suing, found 51 cases of "substantial overlap" with its published work inside that same dataset.

Microsoft's lawyers call those figures modest. The fact that Copilot occasionally mirrors a phrase or sentence, they argue, "hardly undermines the transformative purpose of LLM training", meaning that building a large language model, the technology behind AI chatbots like Copilot and ChatGPT, is different enough from simply reprinting an article that copyright law should not treat it the same way.

What is the fair use argument, and why does it matter?

Fair use is a legal defence that allows limited use of copyrighted material without permission, provided the use is sufficiently different in purpose from the original. A critic quoting a paragraph to review a book is a classic example. Microsoft is arguing that training an AI on news articles falls into the same category: the model learns patterns from the text rather than republishing it.

The publishers and authors see it differently. The New York Times, the Authors Guild, and others claim that Microsoft and OpenAI built commercially valuable products on their work and now compete directly with them, in part by spitting out content that substitutes for the original.

The cases were consolidated under one judge to speed things up, over the objections of the plaintiffs. As first reported by The Verge, the Trump administration also filed a statement of interest this week, backing OpenAI's position in the Times case.

What happens next?

If the judge grants summary judgement, the case ends now and Microsoft wins. If not, the full legal battle continues. The New York Times, the Center for Investigative Reporting, and the Authors Guild did not respond to requests for comment before publication.

For readers who subscribe to newspapers or buy books, the outcome could shape what AI companies are legally allowed to do with published work going forward, and whether creators see any compensation for it.

© 2026 AI2Day