If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.
What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!
It might be a lost cause regardless, because this is under the assumption that we can trust OpenAI to be honest about their own investigation, which is unlikely.
Only under gross negligence would it be unprovable: Did you use a model whose training set included user data? Did the transcripts of any of the agents include a tool call whose result including user data?
>Did you use a model whose training set included user data?
OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out.
>Did the transcripts of any of the agents include a tool call whose result including user data?
>>Did the transcripts of any of the agents include a tool call whose result including user data?
> They have already explicitly denied this.
AFAICT they only denied accessing data through a request targeting a user, not that they accessed (users) data through targeting (an extremely niche) topic, which is the relevant part here.
So, unless I'm mistaken, a "result including user data" is very much still in the air.
OpenAI is the possessive actor here. They were supposed to develop the AI for the benefit of humanity. Now they do it closed and do it only for their private profit.
Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.
Yes you're missing several obvious things. Even saving the last changed date (nevermind every change date or what the change was) for every setting for every user would be earth crushingly wasteful. The value by itself isn't even worth including in backups.
Storing the last changed date for every single person on earth (even though not every person is an OpenAI customer) is something you could easily do on a laptop. It would be a rounding error for OpenAI.
I don't know what format they use for storage, but Iceberg would be a reasonable choice. A date in iceberg format is 4 bytes[1]. I checked postgres as well as a reference point. It also uses 4 bytes for a date, so whatever they use it's going to be about that.
Current world population is just shy of 8.3 Billion people [2].
4 bytes times 8.3 billion people gives 30.92 GiB. [3] OpenAI's training data will be in the petabyte range at least.
This is actually bigger than that. This is an "AI is eating the world" situation.
The math community was one of the first to be affected because RLVR makes Math an easier target for AI. We witness AI eating the math community.
Software community was also affected due to similar and other reasons. All other professional communities will face a similar challenge in the near future.
Do I understand it right that people now claim an ownership of the actual mathematical methods? What next, people patenting the letters and the numbers? And then the sounds? This starts going ridiculous.
If OpenAI remained a full nonprofit looking to build an "OPEN" AI for the benefit of all humanity (not only the US or a few shareholders), I would have been happy to share my code, my work, and even label their data... This said, I don't blame them. It's a difficult mission to remain a nonprofit and, at the same time, have the required capital investment to build AGI.
I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.
Imagine what a genuinely openness-focused organization of this sort could be. Even if we imagined a commercial half, we could imagine a foundation with mass-membership, perhaps with a membership fee equal to 1/2 the typical personal subscription and functioning to set the direction, elect the board, etc., and then a commercial half which might be rough, tricky, deceptive, making deals with anybody.
I think I'd have been fine with the commercial half being a bit of a monster, as long as I'm part of the members and we decide what sort of board it gets and there's a clear "this is basically controlled by the public" and if I were part of a club of this sort, I would, like you absolutely fill up a directory with texts and computer programs and careful annotations to aid training.
and they could have had it. It could have been easy to make an organization like this. I think you still can. An international AI club, the members vote on what sort of training material may be supplied and for what intents, create some committees to review quality, and then everyone starts making their little games and RL environments and annotated stories and programs that ordinary LLMs misunderstand, and then they get together and fine-tune something, and if that works well they then get some staff and better training infrastructure and end up with a commercial half.
Both you and the parent comment give Sam Altman a lot of undue credit. He was given funding specifically because his sociopathic tendencies were considered useful for the job. It was only after he had connections through YC that he started fearmongering around AI. The "AGI" concept as he describes it is part and parcel with this deceptive marketing that he does.
How can any of us act surprised, now that OpenAI has gone mask-off? Sam made those Worldcoin orbs. He's Paul Graham's cannibal king. You have been warned at every stop along this road, and yet we still take him seriously when he says empty platitudes to placate investors. AGI is like "China's Final Warning", a plain lie that only teases a potential innovation but never precedes it.
Ah, I didn't intend to imply anything of that sort.
It's just that I, having never seen what that would actually be like, imagine that I'd be fine with a genuinely democratic mass-organization for LLM development which has a cutthroat, even OpenAI-level cutthroat commercial arm.
If you run a model locally to work on your area of expertise in mathematics, and you make a discovery, haven't you stood on the shoulders of others that created that model you are using? Playing devil's advocate here, are we only in for human collaboration, but not machine contributions in this case? Scholars get cited, but do they normally get paid for others citing them and using their work to advance their own? I get the whole sneakiness about how these companies like OpenAI built their models upon IP and data without a clear trail of attribution or compensation where it would normally be present. I am not a proponent for either side at the moment. I am trying to grapple with this whole new world of shared knowledge and how it is produced and shared and profited from especially when there is a $1m prize to distribute!
I was trying to patent our maintenance tracking algorithm, which produces guaranteed weight loss or gain within 2–4 weeks by producing accurate calorie and macro targets for people to follow; in our test, it beats GLP-1s like Ozempic, Tirzepatide, and Retatrutide in results.
But later we found that algorithm and math cannot be patented.
We shouldn't forget that the only way this unpublished research is supposedly getting into the training data in the first place is because the researchers making the accusations were conversing with ChatGPT in the process of doing their own research on these problems.
But if they are saying this contaminated the model's training data with knowledge of their ideas, who's to say the model they were developing their research ideas with wasn't already contaminated through prior discussions with other researchers about the same topics?
So if contamination is proved, or cannot be disproved, and if these researchers want OpenAI to relinquish its claim to have solved these problems independently, then it would seem they also have to give up their claim to have solved them independently?
1. Tristan's research direction was known to only a handful of other academics. It might help if you can be more specific regarding the origin of the contamination.
2. If someone did come out and claim that their ideas were used without proper attribution in Tristan's work, then of course, that deserves consideration.
3. What you're suggesting seems purely hypothetical. At present, there is nobody claiming that Tristan's work is "contaminated"
4. Tristan was very willing in his initial statement to give credit to the people who developed the ideas.
This. I don't like people using others' work without giving proper credit. At the same time, people are jumping too quickly to defend these mathematicians without any solid evidence of their claims, while framing this as some sort of "Evil AI Company vs human mathematicians".
In the best case scenario, the mathematicians were standing on the shoulders of the extensive training data from sources that aren't being credited and they may not even have had access to.
I have nothing against these mathematicians because I don't know them. And given that, there's no reason to trust their word any more than OpenAI's.
It's possible the work they were doing, even if related, was a dead end and immaterial to OpenAI's findings. Or maybe they are right and OpenAI stole their work. Who knows the truth right now?
In other words, we need more evidence before making accusations.
If I were a business, I wouldn't want my proprietary business information into a hands of a competitor. But CEOs are racing to do it, and (at best) relying in flimsy contractual guarantees they have no means to verify.
When you want your unpublished work unpublished, don't store them on other peoples computers. Especially not on people their job it is to use data to create money.
Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.
Andrew Wiles gave 3 lectures, and only at the end of the last one he announced that he solved FLT. Imagine someone from the audience announced in between the second and the third lecture that they proved FLT (using his ideas, obviously).
This is a great analogy, and one I'm surprised more people haven't raised. Instead all your hear are comments about mathematicians being sour losers, they should have know the model terms of service etc...
43 comments
[ 0.21 ms ] story [ 8.2 ms ] threadWhat the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!
OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out.
>Did the transcripts of any of the agents include a tool call whose result including user data?
They have already explicitly denied this.
> They have already explicitly denied this.
AFAICT they only denied accessing data through a request targeting a user, not that they accessed (users) data through targeting (an extremely niche) topic, which is the relevant part here.
So, unless I'm mistaken, a "result including user data" is very much still in the air.
I don't know what format they use for storage, but Iceberg would be a reasonable choice. A date in iceberg format is 4 bytes[1]. I checked postgres as well as a reference point. It also uses 4 bytes for a date, so whatever they use it's going to be about that.
Current world population is just shy of 8.3 Billion people [2].
4 bytes times 8.3 billion people gives 30.92 GiB. [3] OpenAI's training data will be in the petabyte range at least.
[1] https://iceberg.apache.org/spec/#schema-evolution
[2] https://worldpopulationreview.com/ and elsewhere, say census.gov if you want a US source https://www.census.gov/popclock/world
[3] https://www.wolframalpha.com/input?i=4+bytes+*+8.3+billion+i...
https://mathstodon.xyz/@tao/117237320796901560
The math community was one of the first to be affected because RLVR makes Math an easier target for AI. We witness AI eating the math community.
Software community was also affected due to similar and other reasons. All other professional communities will face a similar challenge in the near future.
Re-using advances done by others is necessary in maths.
This is how hard sciences does progress.
I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.
Imagine what a genuinely openness-focused organization of this sort could be. Even if we imagined a commercial half, we could imagine a foundation with mass-membership, perhaps with a membership fee equal to 1/2 the typical personal subscription and functioning to set the direction, elect the board, etc., and then a commercial half which might be rough, tricky, deceptive, making deals with anybody.
I think I'd have been fine with the commercial half being a bit of a monster, as long as I'm part of the members and we decide what sort of board it gets and there's a clear "this is basically controlled by the public" and if I were part of a club of this sort, I would, like you absolutely fill up a directory with texts and computer programs and careful annotations to aid training.
and they could have had it. It could have been easy to make an organization like this. I think you still can. An international AI club, the members vote on what sort of training material may be supplied and for what intents, create some committees to review quality, and then everyone starts making their little games and RL environments and annotated stories and programs that ordinary LLMs misunderstand, and then they get together and fine-tune something, and if that works well they then get some staff and better training infrastructure and end up with a commercial half.
How can any of us act surprised, now that OpenAI has gone mask-off? Sam made those Worldcoin orbs. He's Paul Graham's cannibal king. You have been warned at every stop along this road, and yet we still take him seriously when he says empty platitudes to placate investors. AGI is like "China's Final Warning", a plain lie that only teases a potential innovation but never precedes it.
It's just that I, having never seen what that would actually be like, imagine that I'd be fine with a genuinely democratic mass-organization for LLM development which has a cutthroat, even OpenAI-level cutthroat commercial arm.
But later we found that algorithm and math cannot be patented.
But if they are saying this contaminated the model's training data with knowledge of their ideas, who's to say the model they were developing their research ideas with wasn't already contaminated through prior discussions with other researchers about the same topics?
So if contamination is proved, or cannot be disproved, and if these researchers want OpenAI to relinquish its claim to have solved these problems independently, then it would seem they also have to give up their claim to have solved them independently?
2. If someone did come out and claim that their ideas were used without proper attribution in Tristan's work, then of course, that deserves consideration.
3. What you're suggesting seems purely hypothetical. At present, there is nobody claiming that Tristan's work is "contaminated"
4. Tristan was very willing in his initial statement to give credit to the people who developed the ideas.
In the best case scenario, the mathematicians were standing on the shoulders of the extensive training data from sources that aren't being credited and they may not even have had access to.
I have nothing against these mathematicians because I don't know them. And given that, there's no reason to trust their word any more than OpenAI's.
It's possible the work they were doing, even if related, was a dead end and immaterial to OpenAI's findings. Or maybe they are right and OpenAI stole their work. Who knows the truth right now?
In other words, we need more evidence before making accusations.
Dude makes up like 50% of the replies here.
If I were a mathematician I would not my unpublished work to go into the hands of a competitor.
If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise.
I wouldn't want the plot to an unreleased book to be suggested to another author.
Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.
Why is it OK if openAI does it?
OpenAI has a lot of money, a lot of rich investors who want it to go to the moon, and an extensive propaganda operation.