OpenAI leaves open a training-data possibility behind its Navier–Stokes model
OpenAI published a Lean-formalised proof for the Navier–Stokes problem and recognised the priority of Levent Alpöge and Tristan Buckmaster. The company said it considered it unlikely, but could not rule out, that de-identified data derived from the pair's product usage had improved its models. Buckmaster says information about their progress reached OpenAI, bringing the distinction between direct prompt access and training-data provenance to the centre of the dispute.
Artificial Intelligence··Midday
The proof is published and priority recognised
OpenAI published a writeup and a Lean-formalised proof for the Navier–Stokes Millennium Prize Problem. The company says training of the model that produced it began on 28 August and that more than 1,000 agents worked for over 50 hours. The same text recognises the priority of work by Levent Alpöge and Tristan Buckmaster. Buckmaster's objection, reported by TechCrunch, rests on his claim that information about the pair's progress reached OpenAI after their preliminary findings.[1], [2]
The training-data possibility left open
The new distinction in OpenAI's post separates direct access to prompts from training data derived from product usage. The company writes that, although it considers the possibility unlikely, it cannot rule out that de-identified data derived from Alpöge and Buckmaster's use of its products improved its models. Buckmaster says he asked company leaders whether the pair's Codex logs had been accessed, was told user data had not been viewed, and did not receive an answer to questions about training. The two accounts therefore answer different questions.[1], [2]
The dispute shifts to the data flow
Buckmaster also says OpenAI researcher Sébastien Bubeck asked in one proposed compromise for Alpöge's name to be removed from the credit; Bubeck denies that allegation. These competing accounts make the proof's mathematical validity and the provenance of training data separate issues. OpenAI's own sentence does not admit wrongdoing; it only stops short of definitively excluding an effect from data derived from product usage. For readers, the new development is that the traceability of the training pool has become an explicit question alongside the priority dispute.[1], [2]