Can anyone explain to me why Tristan simply didn't go to settings page and turn off the thing? Especially when Levent was collaborating with him and Levent is completely aware of how the training data is used and the implications thereby?
If it is oversight, then that's ok and OpenAI can volunteer to make him the lead author which they did. But he's pissed that OpenAI is not letting Levent as well, who had access to internal Anthropic models. So this guy thinks
1. oh my bad i forgot to turn off the consent thing in settings page
2. also i'll collaborate with a literal Anthropic employee who has access to their internal models
3. i'll also reject OpenAI's deal to be the lead author because i want an employee of the competitor to be a part of it
The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.
This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.
oAI has made clear they did not specifically pull in any user data to context for this run.
Tristan Buckmaster’s post cited extensive use of LLMs in the process of his collaboration with Levent:
“We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.
The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).
This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”
The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.
“We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”
For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.
That OpenAI setting helps, but there is a better way to do it. To completely opt-out of training, submit a request via the OpenAI privacy portal.
Visit this website https://privacy.openai.com/policies/en/ , click "Make a Privacy Request", choose "Do not train on my content", and complete the form. That submits a formal objection to training on your data, as required by GDPR/your local legislation.
Haha. I am 100% these companies will ignore this if they choose to. Just as they played fast and loose with copyright rules.
They would do it, the say “ah sorry chaps, impossible to extract it from the dataset by now, anyway we anonymized it so can’t tell what’s what, and we can’t risk losing to China. Oh look - did you see Superman fly outside?”.
This form is legally binding and has more legal weight than just clicking a toggle. If they still train on my data, they can get sued, and I'll get a payout.
This looks incorrect. OpenAI have flat out said that there's no need of doing this and both ways are equivalent
> We respect our users' choice whether to use their data to “improve our models for everyone” regardless of where they express that choice. Users can opt out in the in-app settings or indeed also in our privacy portal. They do not need to opt out in both places, and we will make this clearer in our Help Center.
Can anyone explain to me why Tristan simply didn't go to settings page and turn off the thing? Especially when Levent was collaborating with him and Levent is completely aware of how the training data is used and the implications thereby?
If it is oversight, then that's ok and OpenAI can volunteer to make him the lead author which they did. But he's pissed that OpenAI is not letting Levent as well, who had access to internal Anthropic models. So this guy thinks
1. oh my bad i forgot to turn off the consent thing in settings page
2. also i'll collaborate with a literal Anthropic employee who has access to their internal models
3. i'll also reject OpenAI's deal to be the lead author because i want an employee of the competitor to be a part of it
I don't get the mindset.
This post is out of date. OpenAI quietly updated the references on their paper earlier today and added several authors.
The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.
> our hypodissipative result (which is not public), but as I understand it part of their training data
How did it become part of their training data if it wasn't public? /confused
This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.
oAI has made clear they did not specifically pull in any user data to context for this run.
Tristan Buckmaster’s post cited extensive use of LLMs in the process of his collaboration with Levent:
“We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.
The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).
This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”
The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.
“We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”
For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.
Tristan Buckmaster, the mathematician at the center of it (https://cims.nyu.edu/~tristanb/) put out a public statement that everybody should read (pdf) - https://cims.nyu.edu/~tristanb/statement.pdf
So what might have been the incentive for OpenAI to do all this shenanigans? It might have to do with getting its models certified for AGI and getting out of lockin with Microsoft - https://deadneurons.substack.com/p/the-quiet-unwinding-of-mi...
One of the researchers was using the product. They train on your private interactions unless you explicitly opt out in the settings.
Loosely related, you may want to check your own privacy settings at
https://chatgpt.com/codex/cloud/settings/data#settings/DataC...
https://claude.ai/new#settings/data-privacy-controls
I just realized I've been happily "improving the model for everyone"...
That OpenAI setting helps, but there is a better way to do it. To completely opt-out of training, submit a request via the OpenAI privacy portal.
Visit this website https://privacy.openai.com/policies/en/ , click "Make a Privacy Request", choose "Do not train on my content", and complete the form. That submits a formal objection to training on your data, as required by GDPR/your local legislation.
Haha. I am 100% these companies will ignore this if they choose to. Just as they played fast and loose with copyright rules.
They would do it, the say “ah sorry chaps, impossible to extract it from the dataset by now, anyway we anonymized it so can’t tell what’s what, and we can’t risk losing to China. Oh look - did you see Superman fly outside?”.
European users have right to be forgotten. Waiting for the court order to delete all models.
This form is legally binding and has more legal weight than just clicking a toggle. If they still train on my data, they can get sued, and I'll get a payout.
How would you prove they trained on your data specifically?
> but there is a better way to do it.
This looks incorrect. OpenAI have flat out said that there's no need of doing this and both ways are equivalent
> We respect our users' choice whether to use their data to “improve our models for everyone” regardless of where they express that choice. Users can opt out in the in-app settings or indeed also in our privacy portal. They do not need to opt out in both places, and we will make this clearer in our Help Center.
https://x.com/thsottiaux/status/2097746417012166816