AI & ML
OpenAI can't rule out that it stole its most recent breakthrough
Jonathan Murray DEV Community
2 views
A mathematician spent a year on one of the hardest open problems in math. He asked OpenAI a simple question. Did you train on my sessions? Today they answered. Sort of.
Here is what happened.
The setup
Tristan Buckmaster is a math professor at NYU. He and Levent Alpöge spent most of the past year attacking finite-time blowup for fluid equations, the family of problems that includes the Navier-Stokes Millennium Prize. They worked with LLMs the whole way. Claude, Codex, GPT-5.6 Sol, Astra. By August 15 they had blowup results for Boussinesq and 3D Euler. By August 22 they had a proof verified in Lean.
Every draft of the project went through Codex. His words, from his statement: "our sessions in Codex, into which we had been putting all our drafts for the whole of this project."
The call
September 3. Rumors start moving. Buckmaster emails his contact at OpenAI to ask what's going on. Within days he's on a call with Sébastien Bubeck. He's told an internal OpenAI model has produced a proof of finite-time blowup for forced Navier-Stokes. About 100 pages. Same smooth-forcing setup he and Alpöge had quietly chosen.
He asks when the first prompt was sent. It takes a while to get an answer. Eventually it's agreed: in the past few days, after information about their work had reached OpenAI.
Then he asks the real question. Was the model trained on, or did it have access to, their Codex sessions?
He's told the model did not look up user data.
He asks again. About training specifically.
No answer.
That was the state of things when he went public last night. Alongside it: two proposals to coordinate release, a request to drop Alpöge from authorship because Alpöge works at Anthropic, and a line he quotes directly: "If you don't want me to be nice, then I don't have to be nice." OpenAI's first response called the allegations "false and inflammatory." Bubeck has since called the career remark "ill-chosen" and retracted it.
The answer
Today OpenAI posted a statement. Read this part slowly.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
That is the training question. Answered.
Not "we did not train on it." Not "their sessions were excluded." They cannot rule it out. A year of unpublished work on a Millennium Prize problem, and the company whose model produced the result cannot rule out that the work was in the training set.
Notice how carefully the statement is built. It says no specific user data was "accessed." That's a retrieval claim. It says the researchers and agents did not "see" the work. That's a visibility claim. Training is a third thing, and on training the answer is a shrug.
They also say the proofs differ significantly and the Euler results are different, forced versus unforced. Maybe so. Nobody outside OpenAI has seen their proof yet, including Buckmaster. That part will get sorted out by mathematicians.
Why this matters to you
You don't have a Millennium Prize problem in your Codex history. You have your codebase. Your architecture decisions. The thing you've been building for eight months that isn't public yet.
Consumer Codex sessions are training data by default. That is not a leak. That is the product. Buckmaster's drafts were "de-identified data derived from usage." So are yours.
The lesson isn't that OpenAI did something exotic here. It's that the default did exactly what the default does, and for once it happened to someone whose work was important enough that the question got asked out loud, and answered in writing.
"We cannot rule it out" is the honest answer. It's also the only answer they can give. Think about what that means for everything you've ever pasted in.
Sources: Buckmaster's statement, the OpenAI statement, TechCrunch, Fortune, The Decoder, OpenAI data policy.
Read original: https://dev.to/jon_at_backboardio/openai-cant-rule-out-that-it-stole-its-most-recent-breakthrough-12d2
Related
Does an LSP help a coding agent?
AI & ML
2
DEV Community
My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.
AI & ML
2
Dev.to (EN Zone)
How we made 2,000 customer conversations queryable in a few hours
AI & ML
2
Dev.to (EN Zone)
Robotics Concepts for Beginners
AI & ML
2
Dev.to (EN Zone)
Comments0
No comments yet — be the first