Two different questions
Andreas Thom works on non-sofic groups, infinite mathematical structures that cannot be approximated by finite ones. Last month OpenAI announced a result in exactly that area, and its write-up built heavily on previous work by Thom and fellow mathematician Gábor Kun. Mathematicians criticised the company for not acknowledging that, and the write-up was quietly amended. What Thom then set out in a series of posts on Mastodon was narrower and harder to settle: OpenAI's command of the techniques he had been using struck him as detailed, and those techniques were neither the most obvious nor the most promising routes to a solution at the time. So he wrote to OpenAI researchers Sébastien Bubeck and Mark Sellke and asked whether his own interactions with ChatGPT were part of the training data or accessible to the reasoning process.[1]
The reply, by his account, addressed whether those conversations could be accessed directly, and stopped there. It is the same line OpenAI drew in the blog post announcing its Navier-Stokes solution: no specific user data was accessed in order to solve the problem, and, while unlikely, the company cannot rule out that de-identified data derived from usage of its products helped improve its models. Thom's objection to that second sentence is exact. Removing a name from a conversation removes the name. It leaves the intellectual content of a mathematical idea where it was.[1]
Where the two paths separate
Two data paths are in play, and they are easy to run together. One is access: at the moment it answers, can a model reach a stored conversation? The other is training: did the content of that conversation enter the training data the company uses to improve its models? Denying the first says nothing about the second, and the de-identification qualifier is the sentence that carries the second while sounding like a concession. A fairer reading is available too. The researchers who replied may have understood the question as being about direct access alone, in which case the narrow answer follows from a misunderstanding. Either way, the person who asked cannot tell the two apart from outside.[1]
Only OpenAI has the relevant data for that, as Thom puts it, and researchers are not equipped to reverse-engineer a training pipeline to find out whether their work has been used. That is why his demand is documentary: disclose the necessary datasets, and clarify the settings and terms setting out how the company uses data. Until such a document exists, the mathematicians hold a suspicion with a plausible mechanism behind it, and the company holds a denial narrow enough to leave that suspicion untouched.[1]
The same gap, one more time
This gap has a shape I described at the end of July, writing about the distillation allegation: an account of access and volume, detailed enough to persuade, and not attached to a document anyone outside the company could examine. The posture has flipped since then, from allegation to denial, and the evidentiary position has not moved with it. A named mechanism makes a claim specific; it cannot make it a finding. Only the party holding the training pipeline can do that, which is why the request now on the table is for the datasets and the terms rather than for another sentence.[1], [2]