The two texts on the table
Two texts are on the table. The first is the writeup and the Lean-formalised proof OpenAI published on 8 September. Sébastien Bubeck, a mathematician at the company, said training of the model that produced it began on 28 August; more than 1,000 agents worked for over 50 hours and the number rose to as many as 10,000. Mark Chen, its head of research, described the computing power required as a cost in the millions of dollars. The proof has not appeared in a peer-reviewed journal; what stands for now is a text awaiting examination.[1]
The second is Tristan Buckmaster's objection. The mathematician at NYU says he and Anthropic researcher Levent Alpöge posted their preliminary findings on Monday and that information about their progress had been passed to OpenAI. According to Buckmaster, OpenAI researcher Sébastien Bubeck asked in one proposed compromise that Alpöge's name be dropped from the credit; Bubeck rejected that claim on social media. Buckmaster also recounts asking company leaders whether the pair's Codex logs had been accessed, being told the model did not look up user data, and receiving no answer on questions about training.[2]
The opening in the company's own sentence
The most notable sentence appears in OpenAI's own post. The company says it recognises the priority of Levent Alpöge and Tristan Buckmaster's work and writes that, while it considers this unlikely, it cannot rule out that de-identified data derived from the pair's use of its products had improved its models. That sentence concedes no accusation. What it does is narrower: it records, in the company's own words, that one end of a data flow has been left open.[1]
The two answers respond to two different questions. Bubeck and the company's executives say the pair's prompts were not read and their work was not seen before it became public; that answers a direct claim about reading. The sentence in the post leaves open that de-identified data derived from product usage may have entered the model. The two statements do not contradict each other, and side by side they still do not close a data flow. There is an ordinary explanation for that: tracing one user's contribution backwards through a large training pool can be hard in practice for the company itself as well. In that case what is missing is a traceable account, and no assumption of bad faith is needed.[1], [2]
The document that would close the gap
The shape of this gap is familiar. The distillation column of 26 July also had a described method whose basis was never attached to an examinable document, and the distance between access and training data stayed open there too. The difference here is that the same distance is marked this time by the company's own sentence. The strength of an allegation keeps depending on whether a document exists that could test it.[3], [1]
The document that would test this is narrow and nameable: a data-provenance statement showing, with dates, which product usage data entered this model's training pool and which was excluded. If OpenAI publishes such a statement by 31 December 2026, Buckmaster's question stops being a matter of disclosure preference and becomes an auditable one; if it does not, what remains is again a two-sided account.[1]