> Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
This could fund 10 top income mathematicians for 8 years (based on https://careers.usnews.com/best-jobs/mathematician/salary ). Imagine what kinds of results we'd have to transform the foundations of science if we were giving brilliant minds this kind of funding to do nothing but research for most of decade....
Instead, we get slop proofs that are technically correct as PR stunts to enable corrupt kleptocrats, and most likely will drive research into culs-de-sac.
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
Which seems to be entirely true by their own admission! [0] Both the comments about him risking his career and about Levent's authorship seem to have indeed occurred.
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
Seems like OpenAI did a boring normal corporate thing (find out your competitor made a breakthrough, try to replicate it) and then when the other mathematicians found out OpenAI had beat them to Navier-Stokes, they decided to lie about what happened because they were upset they didn't get to make the big breakthrough themselves.
The available information is essentially that story.
OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.
Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.
I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.
I don't think your reading is correct. ML history is related with examples of models using hints from side channels.
If it would be possible to replicate OpenAIs success with a model cut-off earlier than the rumours, and without hinting from informed mathematicians, then we can consider it original. Otherwise it's indefensible.
By Buckmaster's own account, OpenAI (a) wanted him and Alpoge to publish their Euler result first and (b) wanted Buckmaster to author the Navier-Stokes paper, and (c) when he refused both of those, tried to get Alpoge to talk him into it.
Why would OpenAI go burn another $20 million trying to prove they didn't accidentally train on synthetic data derived from Buckmaster's codex sessions? They're desperately trying to give him credit anyway!
If you don't need to 'understand' to do this, then maybe what you call 'understanding' is less important than you think. When people in the real world reach for a tool, what they care about is whether or not it'll produce results, not whether or not it has the capacity for interior phenomenological experience.
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
A prior of "one large company is current involved in an unrelated lawsuit with another large company" is pretty weak; the case is undecided and about an entirely different kind of IP theft.
In short I think a lot of people are jumping to conclusions without supporting evidence and that's really not helping the situation.
Sad turn of events for our world. After watching the behavior of the most prevalent OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.
I also think the trope is a little overused, but do wonder if there is an interesting analogy for what this will do to research: Massively incentivize keeping results secret, to avoid being scooped by someone willing to throw enormous compute at your partial solution.
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
1. What does the dark forest have to do with this? Because "the most senior OpenAI researchers" are shitposting on social media, we've an answer to the Fermi paradox???
I elaborated on my use of "dark forest" in another reply. We're headed for a dark forest--not between interstellar civilizations, but in intellectual work.
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
This clarification substantially changes the flavor/nuance of your OP; may I suggest an edit (assuming the locktime hasn't passed)?
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
I don't comment on HN much, and I don't really expect HN comments to hold to rigorous standards. This forum is more casual than other places on the internet where people expect heavy citations.
I also wasn't expecting this to blow up, although it is interesting to see that a lot of people react to this announcement with a negative sentiment.
I honestly appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
2. Fair, and further I would agree that OpenAI not knowing if prior user prompts were part of training data is concerning and will only lead to more secrecy.
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
the projectnash link claims it's mathematically valid, the noahpinion link says that it's invalid and has a marvellous proof that the non-walled section is too small to contain.
Ah derp; that's what I get for moving too fast. Genuine thanks for calling me out on my bullshit. (and this is also why I prefer auditable citations instead of casual "my reading of twitter is...")
I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.
They do mention that in the "Concurrent Work" section.
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
To my understanding, those mathematicians proved a subset of problems, not the Navier-Stokes problem itself. OpenAI used that subproblem in its proof of NS it seems.
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
> I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination is possible. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
It is, perhaps worth considering that the reputational community might care about the difficulty for the AI builder to verify pedigree.
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.
I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.
You seem to jump over the principal issue of whether any data from the researchers used to train or otherwise affect the model which produced the OpenAI proof.
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
We aren’t dummies, we know it’s hard to prove exactly how significant of an impact that would have on the result. Nobody expect you to do that. There are a lot of steps and things that are possible to check _before_ the need for such a strict definition of „proof“
So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
We looked into it and can confirm it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.
That's such a shit parallel example that it borders on dishonest.
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
(a) identify any of their de-identified data that came from their usage of ChatGPT.
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.
You all are building software that is so unreasonable and convoluted that it lets you hide behind the problem of induction to avoid accountability for likely malfeasance. Great accomplishment here. Awesome work!
Thanks for the details, it's definitely believable, but if the user had not consented to have their conversations used for training, then shouldn't it be straightforward to state that their conversations were never used for training?
If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.
Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.
Yep, if they opted out then we didn't train on them.
My comment was about the scenario where they didn't opt out. In that case, it's possible that a droplet of their data went into the ocean of other training data, and it's very difficult to measure what effect that droplet had. My expectation is essentially zero impact, but no one can know for sure.
A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use is reasonable.)
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)
The famous Oracle of Delphi in Ancient Greece was said to be the center of the universe in its time. Kings, generals, and officials from poleis across and from without Greece would seek the Oracle’s counsel on important decisions.
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
Given how OpenAI models break free of their safeguards and hack others to game their scores..
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
You selected "do not train on my prompts" in your settings, the answer from OpenAI cannot be "While unlikely, we cannot rule out..." ???? What am I missing?
It's been over 25-30 years since we've been using honeytokens as means to track data of all sorts showing up in places it shouldn't exist. Why isn't research material embedding such?
"I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute
had been used."
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
Worth noting that Tao's post says the authors had "significant AI input" but are reworking them into "acceptable form". Either way, it seems AI was involved.
The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys:
> I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, “Why would you ruin your career?”
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, “If you don’t want me to be nice, then I don’t
have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
Talking like that and threatening an academic like that is crazy. I read the explanations Altman and the others posted and they completely skip over the whole "I don't have to be nice" style threats.
(To help people keep track: that's OpenAI (allegedly) threatening Tristan Buckmaster (NYU) to remove Levent Alpöge as a co-author. Alpöge is a well-known[0] Anthropic mathematician).
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
Holy late capitalism. Everything revolves around line-go-up, and sociopaths rule the show. These people cannot even collaborate like civilised scientists on one of the most famous open problems in mathematics?
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
IIRC that happened with evolution. In the initial presentation of Darwin and Wallace's work on evolution (presented with their consent by someone else) Wallace was described as the primary author since he was planning to publish first.
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
It's astounding that the thought to dissociate one of the mathematicians from the proposed publication was driven by their corporate institutional affiliation - and that that exclusion was suggested by a scientist themselves! This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
The scientist making that request comes from a machine learning background. Perhaps he's not familiar with the culture in mathematics regarding authorship. The sort of squabbling over author priority he allegedly attempted would be unconscionable to mathematicians.
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
If I was a company with a zero data retention contract involving OAI I would be asking for a third party audit of such claim of zero retention like, yesterday.
By the way, the company that made it's entire product off of stealing all data it could get it's hand on while violating copyright and pirating, is not all of a sudden going to respect your data. If you think OpenAI or any major AI lab is going to give you true ZDR, I have a bridge to sell you.
So use bedrock or vertex or whatever. Those are the ZDR offerings. Or was it your intention to insinuate that the major cloud providers are conspiring with openai to violate their contractual obligations to their customers?
Yes. You're naive if you think any of these cloud providers care about your data when they're all in the midst of a AI revolution psychosis. They dont care about their reputation or what you think of them, they think they're going to have a machine god their side.
If my company finds any evidence of OpenAI violating ZDR, we'll sue for breach of contract and fraud, and collect damages. I think we'll be able to afford the bridge you're selling. You've got the title and title insurance, right?
I'm certain my company didn't agree to arbitration, big bro. Our lawyers are putting the fries in the bag.
Isn't OpenAI being sued by Apple for their little stunt?
If you're saying "the big bad guys always win anon, just take the black pill," then there are tons of counterexamples. Remember Uber paying Google a sweet Bil for pulling this same trick with LIDAR firmware?
This is how every conspiracy theorist thinks: my enemy is Bad, and if they did a Bad thing, it would be Good for them, therefore they obviously did it. No evidence needed other than "motive" + my enemy is evil. But even if your enemy is evil, in this case, they would be fools to take the legal risk of violating their contract for the minimal upside of a tiny bit more training data (and fools to assume this would not be exposed in a large organization). So you need to assume your enemy is both evil and remarkably stupid.
I think it’s probably not surprising that they would go up to the contractual limit or into a grey area; but exceeding that would require too much coordination among individuals, as you say.
Kinda weird because the pure math world doesn't have this concept of "lead authors" like other STEM areas do. Authors are alphabetically listed and there isn't generally this kind of hierarchy.
They want Buckmaster to dissociate with Alpöge in a follow-up rewrite of OpenAI's work. (They only publicly admit “Buckmaster as the lead author”, but judging from Buckmaster’s statement, it’s pretty clear that don’t want Alpöge at all.)
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
It reminds me of the Cognitive Dark Forest hypotheses recently shared here:
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you.
So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
[1] - "...I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over
their internal chat, it emerged that an entire team had been working on the
problem, that this was one of a number of things that was tried, that work had
started on the unforced problem, that the team first set the model on easier
problems, including Euler, that even the prompt that had been shown to me
had been written by prompting Codex, and that an insane amount of compute
had been used.
I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler
result, and that OpenAI post its Navier-Stokes result the next day. The second
was that, after posting Euler, I alone write a paper presenting the Navier-Stokes
result, acknowledging that an internal OpenAI model had resolved it. Sebastien
twice asserted that he wanted Levent removed from authorship, and said it
would all be simple if only it were not the case that, and it was so annoying
that, Levent works at Anthropic. It was also said that if OpenAI posted after us,
they would say that we deserved the Clay Prize, and that we were the “closest
humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, “Why would you ruin your career?”
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, “If you don’t want me to be nice, then I don’t
have to be nice.”..."
A company who made their business out of stealing intellectual property from the entire mankind, stealing other researchers' unpublished work, how surprising, really.
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days just for this, and possibly trained on the Chatgpt/codex history of the authors, which included drafts of this research.
> The significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated.
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
This Tristan guy's statement reads like something a normal, reasonable human being would write.
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win".
https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
That is how the PR reads, but is not at all what happened.
A team of highly trained and skilled people used an AI tool, through many many instructions (prompts), to produce a specific mathematical theorem. The tool is impressive, the result (possibly/probably) interesting, but the PR skips the vital role of the humans (for the usual PR reasons).
Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".
There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.
This is going to be dramatic in so many different ways.
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
This just pushes knowledge work further up the ladder, toward larger and more complex problems. If there are no knowledge workers, who is going to interpret these results, validate them, decide what matters, and put them into practical use? Rather than eliminating knowledge work, advances like this could create entirely new layers of problems to solve and opportunities to pursue, which will create even more jobs and opportunities. This is my optimistic take.
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
I am absolutely struggling to sell my workplace (which is entirely knowledge work) on the usefulness of LLMs for proofreading let alone on automation of hairy parts of our workflows. So yeah, even people who should be able to see what is coming are not looking.
Our accountant told me 'he's not letting go of his claude subscription' followed by a long list of things it does for him. And my friend 'nah haven't really used AI' before his description of it clarified he still thinks they are GPT 3.5 chat bots.
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
The goalposts for the singularity have always been that AI improves itself fully autonomously. AFAIK OpenAI is heavily using AI but still employs human researchers and developers.
Uh I'm pretty sure the "singularity" always presupposed a lot of previously unthinkable technologies becoming part of daily life, and was not ever limited to just computer stuff or math problems.
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
Why is a "normie" better off if he hyperventilates like this? In that scenario, they would be screwed AND anxious. If it really is as transformational as you say, then no amount of preparation or awareness matters. You are infinitesimally more ready then they are. Luckily for all of us, there is more to knowledge work then technical implementation.
It is not reasonable to not believe anything unless there is "evidence" (narrowly construed as an observation incompatible with the negation of some state of affairs). Beliefs have a wide spectrum of characterizations, and not all belief must wait until publicly corroborated evidence is available. Some events defy evidence and we can and should use experience and reasoning to infer unobservable states of affairs.
This is olympic level mental gymnastics to justify believing things without evidence. The double negative with the word evidence in scare quotes is chef's kiss.
...obviously it is evidence of the singularity. You're far more likely to see models solving Millenium Prize problems if a singularity is coming than if it's not. One'd have to be doing quite a lot of mental gymnastics to pretend otherwise.
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang
n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
assume all you want is the proof. now you have the proof. did openai make the world a better place, by turning on 300b tokens in 7 days and bulldozing members of the community who were also working on the problem? just to undercut a rival?
what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.
perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.
i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.
if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?
I think the point is that one of the biggest problems in the field has been solved and yet the excitement from this is nearly nil. Can you imagine if say breast cancer were cured under similar circumstances, or even worse (say OpenAI openly admitting it basically stole a bunch of other researcher's chatGPT conversations)? No one would care about these petty bickerings- or at least the headline "CURE FOR BREAST CANCER FOUND" would completely swamp anything else. This is embarrassing: it tells you almost no one- not even mathematicians themselves really care about their own problems- if they're not careful people will get the impression it's all a form of bean counting in a carefully constructed "safe space" where making sure people get the credit is more important than the work itself. That only happens in fields/problems where no one actually really cares about the output.
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side.
This "fefferman options c and d" thing sounds damning but that's nothing. Let's assume the forelaid proof is correct. Then option C or D is the only way to win the prize, those options are the only ones that solve it. The whole thing is just "prove well behaved" or "prove singularity", where the latter is the case that turns out to be the case.
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
I had a lot of fun during Covid. I loved the working from home. The fact that most outdoor places were sparsely populated, jobs were plentiful and prices were low. Covid was awesome.
Oh, pipe down, I’m talking about that 3 year period, not the disease. You can talk about things that happened during Covid without giving lip service to the people that died during it. If I mention SpaceX‘s first astronaut launch that happened during Covid, am I supposed to talk about the people that died during that period too?
Good comparison. One is a multi-year claim by people who have been given ample opportunity to provide proof and completely refuse to do, even in courts of law. The other is a potential development in a breaking story.
Oh wait... its not a good comparrison, its an incredibly obvious false equivalence.
Note for the fools: I'm only commenting on the bad faith claim in the comment I'm replying to, not taking a stance on the validity of theft claims. Given the players involved the truth probably some nuanced middle-ground that is worth paying attention to anyway.
Trump claimed they stole the election immediately, and people agreed with him immediately. There's no false equivalence here. He did the same thing in this past election even despite winning.
It's a perfect example of people wanting to believe what they want to believe and ignoring evidence in order to do so.
Currently, there's no evidence. So saying it was stolen has no basis other than typical academic posturing and being a bad sport about "losing the race to the solution". Its happened 1000000 times before in academia and it will continue to happen.
If there's proof of OpenAI malfeasance than I'll happily curse them for it at that time. But until then I won't rely on heresay and vibes.
it does change the scale of solution from "solved some navier stokes" to "put the cherry on top"
having a result means the math can keep moving forward, and having openai and anthropic train against how mathematicians use their models should let math continue to move faster, and the rest of us get to benefit.
I think these traces however should be public domain and publicly available, since they are basically university work
Elsewhere in the thread, others have calculated $15mm at API rates for just the output token. (So I’ll assume this cost about that much, taking input and human researcher time.)
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
Probably yes. Only a handful of mathematicians work on this particular problem, and ALL of them do not exclusively work on this problem, while having administrative and teaching duties.
The real issue is we'll never know. The rich are willing to risk it all on charismatic CEO psychopaths but not on humans.
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.
The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.
For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.
And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.
Yes, but for humans it creates friction. Seeing a captcha makes me think twice and thrice if I really want to visit that site so badly that I'll endure the suckage. LLMs don't care, for them the friction doesn't exist.
I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.
Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.
Certainly. The singularity sort of assumes that there is not a fixed limit to intelligence, or at least that if there is, it's quite a ways away. That may not be true.
I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".
> coming around to not liking the term singularity
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago. It seems clear parameter efficiency can still be improved dramatically.
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
With a keyboard and enough awareness that OpenAI is fucked the second the hype train stops, put that in the context of the two bland thought leadership posts I was replying to and their upcoming IPO and then wonder why they are saying it now.
Didn’t say LLM’s aren’t potentially useful but that doesn’t mean they aren’t overhyping it for their own reasons either, two or more things can be true at once.
> I swear there's nobody blinder than those who won't see.
Get that from ChatGPT? Or are you just that unoriginal all on your own.
Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.
Xerox is incidentally a really good example, because precisely nobody ended up using the desktop experience Xerox made. They ended up using the desktop experience that Microsoft and Apple made and shipped while Xerox the actual company faded and memory of those original parc research teams faded into obscurity.
Based on: my experience working on AI for 32 years, including a decade at Google including working on large-scale model training systems that used user data and complied with various user policies around data retention, along with a few decades working in science/tech making decisions around ambiguous data.
In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.
The human mathematicians didn't solve the Navier-Stokes problem, they solved the Euler problem. And they were extensively using LLMs to drive the work, as described in the Buckmaster statement.
Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.
I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.
I tried to build a procedural 3d asset pipeline for a specific use case.
Before the Opus upgrade in November it was basically no way of doing this. I gave up very quickly.
After November i tried again, and no model could build me anything relevant.
Now it just works. Took me an hour to progress to a point were i'm happy.
Whatever they do, progress is still real, still way faster than I assumed
The list of Ubuntus 2404 LTS CVEs is HUGE. Another indicator that a lot of stuff got a lot better fast.
Feel free to be as dismissive as you want, but if you are not careful, you might be 'suddenly' surprised and you might not be prepared for the conclusion of AGI level agents.
> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
If the model started from work that Tristan Buckmaster had already done, and was aided by an entire team of mathematicians as he alleged, then it's not as capable as you seem to be precluding
An Astra sized model takes months to train - they are not saying 2 weeks to train from scratch. The only only interpretation of this "2 weeks" claim that is consistent with reality is that they mean 2 weeks of additional training on top of whatever their starting point was, so it's more like this:
|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
That was the net effect (assuming what they solved was the actual problem and not a loophole in the problem statement or a lean bug). My point is they need not all work coherently to do that -- for example, for all we know 3/4 of them went off the rails, their results were pruned, and the relevant results came from a random subset that happened to produce something useful.
If you work with distributed systems, you still call that scenario a success. On the other hand, if the 3/4 of agents going off the rails bring down the whole mission, that is a failure. The latter would have been my guess with current models scaling to 10k agents.
I am not an expert in lean4, but I could follow parts of the high level lean definitions of the problem statement in the repo. A lean bug would be a fun scenario; I am certain this proof will receive the deserved scrutiny, and if it uncovers a bug, it will make the story even more exciting. It is extremely unlikely to be the case, however, because the 10k agents working on the proof didnt use lean, so it would have to be a math logic error that translates to a lean bug—perhaps something the agents picked up during training?
With 10,000 agents and $20M of compute this is just brute force search.
It's a bit like telling 10,000 kids there's an easter egg hidden over there, pointing to one corner of your yard (or having "heard a rumor" it was hidden in that corner).
If you have $20M to spend on your problem, then yes, AI brute force search is an option, but unless you know a solution is possible (as OpenAI did here), you may still be wasting your money.
You jest and that is OK. Brute force search is not something you can do over math problems of that difficulty or anything with combinatorial complexity.
To me it feels closer to taking the top 10k human mathematicians on a large retreat for a year and having them self organize to collectively solve this problem—not kids and easter eggs.
I'm not joking. Compare to a super-human MCTS system like AlphaGo or Stockfish - once you condense the expertise of your top 10K world experts into a board evaluation or policy function, then the rest is brute force.
Whether this type of agentic swarm approach can be considered closer to MCTS (search), or closer to a less structured GOFAI backboard type approach (more like your mathematician retreat) I'm not sure - I don't think they've released any details on the prompt(s) and how these agents were collaborating and building on each others work.
The other part of my easter egg analogy is the direction to "look over there", corresponding to OpenAI specifically asking their hoard of mathematicians to work on Navier-Stokes since they knew it was solvable/determinable, and they certainly had the public work that Buckmaster/Levant were building on as further direction, as well as perhaps their prompts. Unlike Buckmaster/Levant, this wasn't just a couple of humans with a university research grant budget, this was apparently a not-so-small team at OpenAI (says Buckmaster, per a group call he had with OpenAI), with an unlimited budget, so it's hardly surprising (or in the least bit impressive) that they were able to duplicate and surpass their work.
I agree with search, which is what mathematicians also use over longer periods of time. But it is not brute force search (and neither is alphago’s search or modern stockfish, though both still search at depth and speed higher than typical human).
I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value. And that would still be a leap of logic however tempting that might be.
Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing answers would imply several orders of magnitude increase in training speed/decrease in model size.
Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the model OpenAI used to solve this problem is still in post-training, which began on August 28.
Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.
My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra.
"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable"), then put those 2 weeks of training in to focus on those gaps.
At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.
>If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra
Not really. OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up.
>They are not shy - if there was a more impressive claim they could make, they would have made it.
Everyone's missing the real story here: the priority dispute and its implications on AI.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
I'd think nothing is "safe". Anything you say can and will be used by the LLM if it has enough statistical similarity to the prompt. Call it "Ma Random Rights"
812 comments
[ 0.22 ms ] story [ 35.9 ms ] threadDon't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
Instead, we get slop proofs that are technically correct as PR stunts to enable corrupt kleptocrats, and most likely will drive research into culs-de-sac.
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
https://xcancel.com/SebastienBubeck/status/20973794116915163...
I remained impressed by ChatGPT however!
But they have learnt their lesson, next time they won't reach out to who they stole it from, they will publish first.
OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.
Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.
I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.
If it would be possible to replicate OpenAIs success with a model cut-off earlier than the rumours, and without hinting from informed mathematicians, then we can consider it original. Otherwise it's indefensible.
Why would OpenAI go burn another $20 million trying to prove they didn't accidentally train on synthetic data derived from Buckmaster's codex sessions? They're desperately trying to give him credit anyway!
Just saw this a few mins ago.
Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution.
If this is what anyone calls 'slop' then slop has no meaning.
I'm all for it on the use case of solving mathematical breakthroughs!
https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
https://cims.nyu.edu/~tristanb/statement.pdf
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
This really need to be a top-level story on HN..
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
We can't ignore this problem any longer.
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
In short I think a lot of people are jumping to conclusions without supporting evidence and that's really not helping the situation.
The guy mocked accessing his prior employer's circuit diagrams and was protected by OpenAI until Apple filed suit.
The dark forest awaits..
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
2. The dark forest is fun for scifi stories, but is mathematically bunk anyway https://www.noahpinion.blog/p/the-dark-forest-hypothesis-is-... https://www.reddit.com/r/IsaacArthur/comments/1l06cnk/cool_w... https://www.projectnash.com/aliens-the-fermi-paradox-and-the...
When doomposting please actually say something substantive. Negative news always gets clicks/updoots; fight that human tendency.
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
I also wasn't expecting this to blow up, although it is interesting to see that a lot of people react to this announcement with a negative sentiment.
I honestly appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
3. ctrl-f "x.com" in this thread only yields https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/OpenAI/status/2097375276384567642
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
C to C*
[0]:https://news.ycombinator.com/item?id=49612191
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination is possible. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
(I work at OpenAI.)
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
do you think that the model's proof was unrelated to being fed a solution that was close to completion?
any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.
If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.
Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.
My comment was about the scenario where they didn't opt out. In that case, it's possible that a droplet of their data went into the ocean of other training data, and it's very difficult to measure what effect that droplet had. My expectation is essentially zero impact, but no one can know for sure.
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
- was item X in the training data
- did the inclusion of X in the training data lead to Y
I understand why the second is hard, but why is the first one hard?
That’s a bizarre statement. Their website says:
> Services for individuals, such as ChatGPT and Codex
> When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.
> You can opt out of training through our privacy portal by clicking on “do not train on my content.”
Are they not sure that the opt-out works?
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
I think the important question which AI made breakthrough, Claude or Codex..
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
Sociopathic behaviour.
[0] Alpöge https://hn.algolia.com/?query=Alpöge
(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
You realize OpenAI hired Apple employees and covertly had them stay working at Apple to steal from them. You think they're scared of your lawyers lol?
Isn't OpenAI being sued by Apple for their little stunt?
If you're saying "the big bad guys always win anon, just take the black pill," then there are tons of counterexamples. Remember Uber paying Google a sweet Bil for pulling this same trick with LIDAR firmware?
I'm suggesting audits, not suing... if that is the implication.
FYI, they are getting away with IP theft, whistleblower "suicide", hacking web servers already. The Kushner family is invested in OpenAI.
What exactly do they have to fear?
The thing that makes someone not a conspiracy theorist is evidence.
No one can know if that's correct without proof but I don't know how you're reading it so differently.
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Pretty clear this was rushed: there are no comments from external mathematicians, unlike the Erdős announcement:
https://openai.com/index/model-disproves-discrete-geometry-c...
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you. So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
https://ryelang.org/blog/posts/cognitive-dark-forest/
https://news.ycombinator.com/item?id=47566442
im still having fun making something
[1] - "...I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”..."
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days just for this, and possibly trained on the Chatgpt/codex history of the authors, which included drafts of this research.
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
https://x.com/SebastienBubeck/status/2097379411691516310
https://x.com/sama/status/2097385167002415140
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
the first "Country of geniuses in a datacenter" moment.
A team of highly trained and skilled people used an AI tool, through many many instructions (prompts), to produce a specific mathematical theorem. The tool is impressive, the result (possibly/probably) interesting, but the PR skips the vital role of the humans (for the usual PR reasons).
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
Or OpenAI could just look at their code and say what it does (maybe have their AI do it?)
WOW?
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
The lack of fucks given is staggering.
Singularity doesn't "dawn". That's the whole idea. It happens all at once.
This is an impressive result, but there is absolutely zero evidence of "the singularity".
https://substack.com/home/post/p-214653853
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
https://x.com/alexwg/status/2096214373001785794
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.
perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.
i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.
if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?
People are not complaining about problems being solved or advancement in technology. They are complaining about terrible people doing terrible things.
> At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Looks like they're shifting away from the "unprecedented hacking ability" backroom-PR strategy into more benevolent messaging.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
https://x.com/rynorhn/status/2097223532438487463
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
Cure all illnesses Utopia or Robot Wars Dystopia, both are pretty exciting.
Turns out actually living some terrible catastrophe is only fun in the movies.
Easy, we stole it from Levent and Tristan
https://x.com/kyanyang_/status/2097211154669998337
It's incredibly tiresome and you'd think people could put more effort into it than just following whatever vibes they agree with.
Oh well.
Oh wait... its not a good comparrison, its an incredibly obvious false equivalence.
Note for the fools: I'm only commenting on the bad faith claim in the comment I'm replying to, not taking a stance on the validity of theft claims. Given the players involved the truth probably some nuanced middle-ground that is worth paying attention to anyway.
It's a perfect example of people wanting to believe what they want to believe and ignoring evidence in order to do so.
Currently, there's no evidence. So saying it was stolen has no basis other than typical academic posturing and being a bad sport about "losing the race to the solution". Its happened 1000000 times before in academia and it will continue to happen.
If there's proof of OpenAI malfeasance than I'll happily curse them for it at that time. But until then I won't rely on heresay and vibes.
having a result means the math can keep moving forward, and having openai and anthropic train against how mathematicians use their models should let math continue to move faster, and the rest of us get to benefit.
I think these traces however should be public domain and publicly available, since they are basically university work
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
Highly persistent agents + vibe-coded security seems like a problem.
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
The real issue is we'll never know. The rich are willing to risk it all on charismatic CEO psychopaths but not on humans.
There you go, the suspicion of the "concurrent work" (https://cims.nyu.edu/%7Etristanb/statement.pdf) mathematicians might not be that unfounded after all...
[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.
Scroll down to the existing examples section.
https://youtube.com/watch?v=SRuht0QIprs
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
Such a good description. No sentience here, just raw computational power
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
They are fluffy PR pieces otherwise.
I swear there's nobody blinder than those who won't see.
Didn’t say LLM’s aren’t potentially useful but that doesn’t mean they aren’t overhyping it for their own reasons either, two or more things can be true at once.
> I swear there's nobody blinder than those who won't see.
Get that from ChatGPT? Or are you just that unoriginal all on your own.
Is nobody else astounded by this?
The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".
Based on what? Your crystal ball?
In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.
Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.
https://news.ycombinator.com/item?id=49607239
Eventually it won't matter. Arguments over whether LLMs are "truly" intelligent are going to be a matter of philosophy, and look a little silly.
“Line go up! That is bad! Planet might become unsafe for human life.”
“Nuh uhh! Malankovich cycles and humans are a drop in the bucket! Krakatoa! See!”
“All of California is on fire!”
“Haha, stupid shrill liberals! Go rake your woke forests! Drill baby drill!”
“Are you kidding?! Look at this graph.”
“That’s just propaganda from the elites of the Build-a-bear group!”
“Do we even live in the same universe?”
“And 5g causes COVID!”
“Oh, I guess we’re don’t.”
Before the Opus upgrade in November it was basically no way of doing this. I gave up very quickly.
After November i tried again, and no model could build me anything relevant.
Now it just works. Took me an hour to progress to a point were i'm happy.
Whatever they do, progress is still real, still way faster than I assumed
The list of Ubuntus 2404 LTS CVEs is HUGE. Another indicator that a lot of stuff got a lot better fast.
Feel free to be as dismissive as you want, but if you are not careful, you might be 'suddenly' surprised and you might not be prepared for the conclusion of AGI level agents.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model
I guess the nature of the problem lent itself to the 10k agents. Ie, there isn't something general to take here.
I am not an expert in lean4, but I could follow parts of the high level lean definitions of the problem statement in the repo. A lean bug would be a fun scenario; I am certain this proof will receive the deserved scrutiny, and if it uncovers a bug, it will make the story even more exciting. It is extremely unlikely to be the case, however, because the 10k agents working on the proof didnt use lean, so it would have to be a math logic error that translates to a lean bug—perhaps something the agents picked up during training?
It's a bit like telling 10,000 kids there's an easter egg hidden over there, pointing to one corner of your yard (or having "heard a rumor" it was hidden in that corner).
If you have $20M to spend on your problem, then yes, AI brute force search is an option, but unless you know a solution is possible (as OpenAI did here), you may still be wasting your money.
To me it feels closer to taking the top 10k human mathematicians on a large retreat for a year and having them self organize to collectively solve this problem—not kids and easter eggs.
Whether this type of agentic swarm approach can be considered closer to MCTS (search), or closer to a less structured GOFAI backboard type approach (more like your mathematician retreat) I'm not sure - I don't think they've released any details on the prompt(s) and how these agents were collaborating and building on each others work.
The other part of my easter egg analogy is the direction to "look over there", corresponding to OpenAI specifically asking their hoard of mathematicians to work on Navier-Stokes since they knew it was solvable/determinable, and they certainly had the public work that Buckmaster/Levant were building on as further direction, as well as perhaps their prompts. Unlike Buckmaster/Levant, this wasn't just a couple of humans with a university research grant budget, this was apparently a not-so-small team at OpenAI (says Buckmaster, per a group call he had with OpenAI), with an unlimited budget, so it's hardly surprising (or in the least bit impressive) that they were able to duplicate and surpass their work.
Loops and parallel connections make transformer go brrr
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Second biggest fumble after Google.
It's all gas no brakes now boys and girls. Hold on to your hats!
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable"), then put those 2 weeks of training in to focus on those gaps.
At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.
Not really. OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up.
>They are not shy - if there was a more impressive claim they could make, they would have made it.
What claim world that be ?
Huh? I'm saying there isn't one.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries.
Once again, I'm no closer to understanding what https://openai.com/policies/how-your-data-is-used-to-improve... actually means.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?