Eigen RadarAI
Analysis

Microsoft tells court its Copilot rarely echoes news text or book passages

In a 4 September summary judgment filing before a Manhattan federal court, Microsoft said less than 1 percent of the 8.2 million Copilot conversations it gave plaintiffs' experts reproduced even 16 words of news content, and a separate expert found just 24 matching responses on 212 books, with no match at all for 202 of them.

Artificial Intelligence··Night
Three lawyers push two steel carts stacked with white file boxes toward a columned federal courthouse on a sunny Manhattan afternoon.

Microsoft counts a thin overlap with news content

In new filings in the consolidated copyright case brought by The New York Times, other news publishers and book authors, Microsoft said less than 1 per cent of the 8.2 million Copilot conversations it gave the plaintiffs' expert reproduced even 16 words of the news content used to ground the model, The Verge reports. The company's own count, cited by The Verge, found 59,545 of those 8.2 million logs sharing 16 words with news text, 51 passages substantially overlapping with Center for Investigative Reporting work, and 24 responses matching at least 30 words. The Times' lead counsel, Ian Crosby, rejected that framing and said discovery shows Microsoft and OpenAI stole from the paper to build commercial products that substitute for its journalism.[1]

The same filing reports an even thinner overlap with books

TNW reports that the same 4 September summary judgment memorandum told the Manhattan court that training large language models on books is fair use, citing an expert who found just 24 matching responses across the 8.2 million Copilot conversations, with no match at all for 202 of the 212 books at issue. TNW puts that 24-match figure at roughly 0.00029 per cent of the reviewed conversations. The case sits in the multidistrict litigation before Judge Sidney Stein, which gathers claims from authors and publishers including 400 local newspapers, and Microsoft is asking the court to decide the whole question as a matter of law.[2]

Microsoft is pushing for an early exit from the case

Both filings serve the same goal: The Verge reports Microsoft filed as it argues for a summary judgment that would end the case early, while TNW notes Microsoft wants the whole fair-use question decided as a matter of law rather than going to trial. The two figures Microsoft submitted, 59,545 matches out of 8.2 million conversations with news text and 24 matches out of 212 books, both serve that same argument that reproduction is rare enough to qualify as fair use, even though Ian Crosby's response, carried by The Verge, disputes that framing for the news-content side of the case.[1], [2]

References

  1. News sourceThe VergeMicrosoft's court filing counts 59,545 of 8.2 million Copilot chat logs sharing 16 words with news text↩1↩2
  2. News sourceTNWMicrosoft tells the same court its Copilot barely echoes books either↩1↩2