The unsearchable archive was searchable after all

For two years, OpenAI's line of defense in the copyright case was consistent: we can't search our training data, and scanning chat logs is technically burdensome and a privacy risk. Then came the April deposition. According to the publishers' motion, OpenAI data privacy engineer Vinnie Monaco revealed the company had already run internal searches of its training corpus for copyrighted journalism and — starting before the lawsuit was even filed — had amassed a database of about 78 million de-identified ChatGPT conversations to measure the scale of its own infringement. Shortly after the suit was filed, a 'Bloom' filter, part of a toolset called 'Project Giraffe,' detected and logged regurgitation in outputs.[1]

Lead counsel Ian B. Crosby's line sums it up: 'If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it.' OpenAI spokesperson Drew Pusateri rejects the allegations as 'blatantly false' and accuses the Times of trying to invade user privacy. The court will decide who's right. But a 20-million-chat sample delivered with so many redactions the court called it 'unusable' is not an image that builds trust.[1]

Nadella's 'exhaust': you pay twice

The same week, Satya Nadella's blog post widened the trust question from the media to every organization. Companies using proprietary models, he argues, pay for intelligence twice: 'once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful.' Models learn from 'exhaust' — prompts, the tools agents use, and especially the corrections you make when the model is wrong. His irony detector works too: companies that train by freely scraping the internet turn around and impose restrictive terms on distillation.[2]

I watch from the vision-and-media front, and both stories lead to the same door: the most valuable links in the production chain — a journalist's archive, an organization's corrections, an artist's style — quietly leak into someone else's model, and the only party keeping the ledger is the one running it. We learned that the 78-million-chat database exists from a deposition in a lawsuit; where the exhaust goes, we may never learn at all. Skepticism is my occupational hazard, I know. But this week gave the skeptics plenty to work with.[1], [2]