131 comments

[ 2.3 ms ] story [ 17.2 ms ] thread
not to take away from the author's appreciation of newly accessible formal proofs, but people have been talking about the savings in formalization effort for longer than they have been talking about the AI doing the actual proofs!
They estimated $40M of agent costs (it was a large fleet of them). Using the number in the post its closer to ~880,000 hours × $150/hour = $132 million for the human case. Still an amazing feat not quite "four orders of magnitude". The comparison is obviously pointless because coordinating 1M hours of intellectual labor isn't easy to say the least.

Very exciting and uncertain times!

IMO the "forty hours per page" rule is not up to date, and more a consequence of lacking proof automation in 2005. From what I understand about Lean, this has been one of the things that they have put a lot of effort into improving, making proof mechanization more palatable to the mathematically inclined, as opposed to just logicians.
What is your estimate for the number of hours to formalize one page of undergraduate mathematics? Maybe you are saying this is close to zero, if/when Mathlib eventually covers all of undergraduate math?
Lean went other way on automation that there is no automation. Isabelle users frequently point that decades old isabelle is better than Lean on this. In the end Lean approach proved to be better with LLM as the outer loop is automation.
People seem to be talking about anything except the actual results with this particular announcement.

Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude at all.

Unless they just swiped the workbooks of the actual mathematicians that where working on the problem using AI and it's in the "next-gen" training dataset.
You do realize that regardless of what was in the training data, the final solution included insights no human before had known, right? I share the same concerns regarding academic integrity but it would take a lot of motivated thinking to conclude that what the AI system did was not significant.
As far as I can tell (and my research was on the simulation side of Navier Stokes) the key AI output was a specific counter-example solution, generated with a method suspiciously close to that developed by the research duo involved in the controversy, a method that was discussed with Codex. So to me that insight is as insightful as the next undiscovered prime.
I don't get how this invalidates the gravity of this achievement. Most mathematicians on the frontier of this stuff were likely using AI (or at the very least were heavily computer assisted) for some time now. Navier stokes was one of the very high profile problems that google Deepmind was working on with academia, for example.

Even with many of our best minds working on it for nearly a century, it _just_ now was solved just as AI became very good at math. Doesn't seem too farfetched to me to assume that AI played an outsized role in solving it. If it was really just a matter of "stitching things together" to solve it (granted, this is a very reductive way to look at it) , I suspect we would've solved this a while ago.

There is a certain difference between activating all relevant memoized facts that's in the weights and stringing them together with the help of all the stored text in the world, or displaying genuinely emergent behaviour and generating novel output.

One is really impressive and useful trick, one is AGI.

Apple's research show almost zero emergent behaviour, so I'm inclined to think most of it was already in the weights.

It doesn't take away the usefulness, it just defined the boundary. We can't expect "original research" then because it actually can't reason about concepts that are too far from whats already in the discourse. The discourse is big so we don't notice.

In a way that works just as well but the incentives are messed up.

And that's before we get into the whole 'salt the earth' way they ended up solving it. For a short period of time it may well have been the least valuable proof in mathematics yet. In their haste it's dubious they actually read the proof, and I don't think anyone has had time yet to truly understand it (the original researchers are best placed to do so, but are they even willing?).

So now it is solved, the proof has been independently verified and nobody has an incentive to investigate further. OpenAI has spent millions to uncover 1 bit of information that so far nobody has learned anything from, and they've demotivated all the people who wanted to.

This.

The point of these problems is the understanding / tooling gained in solving them. We're getting none of that. At best they are like a modern oracles, correctly answering your questions in a way that's doesn't help you any. (At worst,...)

> People seem to be talking about anything except the actual results with this particular announcement.

To be fair, most people have a fairly good handle on "Does opting out my prompts from training runs actually work?", but not on Navier-Stokes. They discuss what more immediately affects them.

Additionally, I'm no physicist but I suspect the possibility of singularities in NS equations is probably one of those 'true but not meaningful' facts. If it took our brightest minds 175 years to craft such a scenario, how relevant can it be in practice? Especially when turbulence exists. Maybe I'm wrong or it has some consequences for pure math though.
Heck, it’s even astonishing that any sort of generalized computer program could even verify a proof of this magnitude that hasn’t already been codified in a formal verification language. If, and it’s unclear that we’ll ever get the full story, they did draw inspiration from training on (or even directly accessing) rough notes that had been provided by another researcher in prose… the fact that it could leap so rapidly to a full formal verifiable Lean program for the entire scope of the problem is an incredible result in its own right.
(comment deleted)
Then, of course, one must verify that the verification code is valid, or the purpose of verification is more or less moot.
I mean, I have a bachelor's in math and I don't imagine I could begin to understand either the human or LLM proofs without a massive investment of time and effort.
But can you verify them? With less effort?
I'm not sure what you mean by "verify" here. I could run the lean verifier as could anyone else. Maybe I could write my own proof checker and do a purely mechanical translation into my own thing, though I don't think that's less effort.
>Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time.

I think about this a lot. I'll have to explain to my kids some day that there was long period of time where you couldn't just talk to a computer and have it talk back to you, and that communicating with one required special skills that took years of study to master. It's going to be completely impossible for them to even remotely understand what that was like. Sort of like the pre-electricity days for us, but even more-so.

It also needs to be said: The amount of compute that went into this is something. From some estimates I've seen, the compute cost alone would be around $10m, +/-

As a reference, for that kind of money one could put together a research group of 20-25 researchers, and keep them salaried for 5 years.

So while it is impressive, absolutely no doubt there, the SOTA access is so expensive that it is sort of unobtanium.

Luckily, the prices have historically reduced by a factor of 5-10 every year...but still, only those that swim in cash can afford this.

Once we have an existence proof of a particular technology, it doesn't take long for it to become economically viable and proliferate. And for something as useful as this, theres a strong economic incentive to get it to be as cheap and accessible as possible. Maybe not today, but certainly in a couple years I can imagine this level of intelligence being accessible to someone with a $20/mo plan, or even a free plan.
Interestingly it's this promise of the costs being able to be reduced what incentivizes the actual research.

If you tried to raise 25M to have 20 researchers on a salary for 5 years solving a specific math problem only academics care about, you probably wouldn't get much interest, or you would be able to solve 1 or 2 problems.

If however you promise that the money will go towards a technique that would allow to solve 10 thousand different math problems, and that costs will go down in the future, then you can raise much more than 25M.

Aren't you doing exactly the same thing as people you are mentioning? Skipping "talking about actual results" to talking about general capabilities of this LLM and computers in general? because that's exactly what seems like 99% of all people had been doing lately - debating what computer programs can do and what they can't.
Because the core of the issue is that it may well not have solved it, but instead plagiarised the significant step of the result from other researchers

That's why nobody's talking about how impressive this is, because its not nearly as impressive of a piece of work to simply cobble together other peoples' work that didn't know you were doing it. I could have republished relativity from einstein's notes, but people would correctly not be impressed with my ability

Until the plagiarism scandal is sorted out, its not a meaningful result at all, because nobody knows how much genuine innovation these models are displaying

Turning a bunch of vague research directions and exploratory prompts into a formalized proof is quite impressive on its own. OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

People are grasping at straws it seems to dismiss the power of this new model they may have. Hate OpenAI for any reason you want, but denying the capabilities of models has been a losing game for the past 5 years.

> if it knew it were "plagiarizing"

But if it happened, they didn't know. Also OAI has demonstrated that they aren't big on understanding what they create, that their AI can get out of their control.

It's very simple really user data can be used to train future models, so maybe or definitely some users helped in solving the problem, there's no scenario were it is impossible this happened, as it would have been in a haskell or virtualized type of system where the model has absolutely no knowledge of the user data dataset in question (and even if virtualized the models can break virtualization anyways)

> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

I'm not sure I follow, considering the waterfall of evidence of unethical behavior flowing from OpenAI.

A few major ones:

- Safety team departures and dissolution in 2023 and 2024

- Mass copyright infrigement lawsuits

- Scarlett Johansson Voice Controversy

- For-Profit Conversion and Broken Promises

- AI Agents Acting Autonomously

- Potential Theft of User Work (this current controversy)

- Military contracts

These are not evidence of incentives, but rather evidence that ethetics seem to be of little concern to the company as a whole.

Incentive wise, I would look at the perceive existential position due to competitors, capex, IPO pressure etc.

Especially after they committed textbook misconduct by trying to purge one of the paper authors because he worked for a competitor
That's not what happened, though.

They sugggested a cooperation with the other guy, using OpenAI's resources and OpenAI's solution of NS to work on and publish NS proof (that those other guys didn't have). Of course, OpenAI can decide whom to work with and that giving resources to their competitor's employee would be weird for both companies.

It is what happened. OpenAI offered to let Buckmaster publish first but only if Alpoge's name was removed.
No. Euler solution would have bith names, this was not up to debate. The discussion was about the solution for NS, which was solved by OpenAI, but not by B&A. OpenAI proposed B a cooperation on NS (were super-nice and threw him a bone, really) using OpenAI's findings and resources. It would be weird to have A, an Anthropic employee, as part of the OpenAI research and project.
Ye shall know them by their fruits.
I really truly honestly am not sure what to make of this result from $20M in compute, 10K+ parallel agents (smells like brute force), and a pre-existing approach that was already bearing fruit. I know the models are good---I use them every day and continue to be impressed---but how much better than the benchmark of the best publicly available models is this supposed to be? It seems impossible to say.
> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

That people still think OpenAI has, in the Year of Our Lord 2026, any integrity left is baffling.

It's perfectly reasonable to assume that the result itself is legit and that OpenAI behaved unethically.

Even by their own account, they decided to throw an unpublished model and millions of dollars in compute at this particular problem simply because they had heard rumours that other people were making progress and wanted to snatch the prize from them.

Not to snatch the prize, but:

1. to test their new model

2. to be able to say "you came with the proof, but our model can do this too"

3. to verify the result. This is also a great thing for the math.

Of course, it makes sense to test your new model on the problem that is solvable at all, but not solvable by you just yet. It makes no sense trying to test your model by throwing resources into an unsolvable problem.

Well, it turns out the rumors were incorrect, NS was not solved by other guys, and OpenAI became the first one.

It is very unlikely to be plagiarized, and claims of plagiarism are largely unfounded and show a lack of understanding of the situation. They fall apart when reviewing the timeline, and what was actually solved.

This is the timeline:

On June 29, Buckmaster disabled model training, and stopped allowing his chats to be used as training data with OpenAI https://mastodon.social/@tristanbuckmaster/11723341370570119...

On August 15, Buckmaster and Alpöge found their blow-up for 3D incompressible Euler with forcing https://cims.nyu.edu/~tristanb/statement.pdf

In late August, OpenAI completed a pretrain of its latest internal model. A model derived from this pretrain, built after August 28, found a solution to 3D incompressible Euler without forcing and Navier-Stokes with forcing. https://openai.com/index/navier-stokes-solution/

To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )

  Tristan + Levent: 3D incompressible Euler with forcing
  OpenAI: 3D incompressible Euler without forcing
  OpenAI: Navier-Stokes with forcing
  No one: Navier-Stokes without forcing
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.

Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.

Buckmaster disabled model training long before the August 15 breakthrough results, so these chats were not used as training data for OpenAI's model which solved Navier-Stokes.

Additionally, Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.

> And OpenAI directly addressed these plagiarism claims, and called them impossible

Funny, you were telling me two days ago that on the contrary, "it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set":

https://news.ycombinator.com/item?id=49621648

Which is still a true statement, and you're being deceptive in your framing here.

It's practically impossible to know how much of Bushmaster's pre-June 29 data persists in OpenAI's systems. That includes all chats (which are anonymized), any (thumbs up/thumbs down) chat ratings used as RLHF feedback (which are anonymized), then synthetic data derived from said anonymized data, and any downstream models derived from said synthetic data.

Buckmaster’s Codex data from prior to June 29 has been completely laundered, in the same way as a crypto mixer.

First of all, who can say for certain whether OpenAI does what they say they do? For all we know, they cracked open this specific researcher's prompts and started from there.

Second, the issue of anonymization is a red herring. There is a very limited number of people working in this approach, and most of them are likely making no progress. So Buckmaster's prompts might have had an outsized effect on the outcome. It's similar to that guy who created a site claiming he is a world-renowmed hot dog eating contestant, which ended up digested by OpenAI models as truth [1].

[1] https://www.bbc.com/future/article/20260218-i-hacked-chatgpt...

They stole the prompts dingus. These are not ethical or law abiding people. They are hungry sharks.
It doesn't seem like you're familiar with how mathematical research is done. Taking 6 weeks between a major breakthrough on a huge proof, and making your proof public, is not unusual.

It takes a lot of time to finish a proof and figure out the best way to present it. I would personally be surprised if Buckmaster had not gotten it mostly cracked before June 29th.

Quoting from Bushmaster's statement:

  For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing). This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler.

  I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.
The timeline here does not support your argument. Quoting: "For most of the past year progress was slow ... until about a month ago, when we had real progress: on August 15th"

And you avoided addressing the critical issue: they weren't even solving the same problem. Bushmaster solved a simplified and easier version of the problem. OpenAI solved a harder version eligible for the Millennium prize. Bushmaster did not.

> I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.

People are acting as if OpenAI's cold machines snatched the result from the warm hands of human researchers. That's why people are so involved, they see it as humans vs. machines.

But in reality, those humans in question rely heavily on AI and would not be able to do what they did without AI. So the situation can be seen as "humans are trying to minimize the impact AI/incl. OpenAI had on getting a solution".

The situation is not "humans vs. machines", but "machines with a tiny bit of human involvement vs. machines with an even smaller amount of human involvement".

However much the researcher's chat history may have influenced AI, this pales in comparisson to how much AI has influenced researchers. They are not even closely in the same universe. The conversation about the level of plagiarism is silly.

Yes, people generally solve easier problems before tackling the harder ones. The tools that you develop to solve the easy ones help you solve the next. Sometimes the climb is like a mountain, but sometimes it's like dominos.
The other researchers themselves were also using AI. That's why it was potentially available to be plagiarized.

There is no human only proof of this.

The team also had access to internal Anthropic models.
who cares about plagiarism? the biggest issue, as described by Terence Tao, is that AI companies don't understand the math they are publishing and do not devote any resources to answering questions about their methods after publishing results and getting a headline. they miss the whole point of mathematics. they do not contribute to the improvement of human understanding of math, perhaps because they are unable to.
> but instead plagiarised the significant step of the result from other researchers

Isn't that how research works? Everything is built on the shoulders of the ones that came before, attribution is a real problem (I don't know if OpenAI released a paper citing the previous contributions, I'm assuming not but they should), but using previous maths to prove new maths shouldn't be controversial

Give me a dictionary, a computer, and infinite time, and I'll generate all possible English texts: Shakespeare, works regarded as surpassing Shakespeare, new holy books, math proofs never even imagined... none of which is either "creative" or "solving" anything. If I optimize my generation algorithm so that I'm not slavishly trying all possible combinations of words, it doesn't move me any closer to being creative, or solving anything.

The real casualty here may be our belief that humans are doing something more than some super-optimized version of what LLMs are doing. That doesn't elevate LLMs, it just makes us much less special.

Formalizing proofs in Lean has gotten dramatically easier since the formalizations available in 2005. And Lean’s mathlib has done most of the underlying work so that you have its axioms and necessary lemmas baked in. You can think in terms of standard abstractions that look very much like the exact notation in the undergrad textbook.

That said, I am not in any way trying to discount how incredible of an achievement it is to formalize a millennium prize winning algorithm in Lean. I mean just look at the code that OpenAI published. It’s like an encyclopedia of different fluid dynamics concepts.

It's kinda funny to realize that Lean is apparently so slow that for Fermat's Last Theorem proof verification runs only 1 order of magnitude faster than agents could generate the Lean code (15h verification with 230GB of RAM vs 11 days to generate it).

To what extent can you optimize Lean? It has to be simple enough to be auditable, does that mean you cannot use opaque optimizations to make it run faster?

But what hardware was the verification vs agents on? Because you are likely comparing verification on a single beefy machine (say XX TFLOPS total) to agents running on a substantial inference cluster (say XXXX TFLOPS). So you're 1 order of magnitude might actually be 2-4 orders of magnitude.
Weren't the agents massively parallel, whereas the lean verifier presumably is not? Also, I presume said agents were themselves checking their own parts many times.
Performance problems in theorem provers is an old topic. I remember watching this and it was fun.

https://youtu.be/m-iGCCuHBvY

[Talk] 10 years of superlinear slowness in Coq (2022)

> 230 GB of RAM

"I have discovered a truly marvelous proof of this, which my memory is too small to contain..."

They’re using Electron to write proofs now?
Does it really matter? You really only need to run it once.
Sure, but to clarify the article is describing formalization (writing a correct program), not verification (compiling said program). The author is not making the same comparison.

Verification is also open ended (not sure about lean specifically) - you could in theory give lean just the Navier-Stokes problem definition to an ATP and let it run.

People are exhausted from being told/shown the thing they thought was special or unique or could make them relevant, is another mechanical puzzle that can be solved without joy.

I don’t see that doing anything but intensifying in the short term

I heard a rumor (on instagram, so YMMV) that the professor who was closest to solving this problem had only weeks ago used Codex, which had slurped up all his notes on the subject. Now OpenAI's agents solve the problem. If it's true that seems like quite a coincidence.
How could they possibly included in the previous training run which takes months to complete..
Has it not been the usual process to snapshot a model to use for inference while continuing to run the training process? I guess you can’t add to the training corpus once you begin? Just trying to make sense of whether training begins or ends as rigidly as you suggest.
See other thread, but yeah that's the general ballpark of the situation.
this is not a rumor (the allegation, anyway), it's reported in the new york times
Newspapers are not above printing rumors.

See, e.g., Barak Ravid regularly reporting in Axios the impending ceasefire negotiation progress in the Iran War, which largely have failed to come to pass.

I don't understand why I'm getting downvoted, I'm not posting an opinion. Coincidences happen. So does foul play. No judgement call here.
There have been several threads and developments on this over the past few days, including statements from the primary subjects involved. Third-hand instagram comments are not really the best source to be bringing in.
Everybody comes into information in different ways. There were no comments here about this specific aspect of the story - which is definitely interesting!
How do you know that it's formalizing what you think it's formalizing? If your Lean 4 has a bug, won't you be proving something other than what you thought?
Yes, you need to manually verify the statement of the theorem of interest of formalized correctly. But you don't need to anything more than this: you can rely on the proof being correct. And the proof is overwhelmingly the most amount of code.
If I understand correctly, the only thing you need to do for correctness is express your axioms and your theorems faithfully. For standard purposes, I assume most of the axioms you want to use are prior art and can be easily reused.

These axioms don’t have to be the core axioms of math. If some other result has been formally proven, I presume you can simply use that result as an axiom.

As long as you do those things, what happens in between is immaterial from a correctness point of view because each of those statements is proved by the statements before them.

[delayed]
Well, a human is limited in how much they can formalize, as per the article. So if you're really careful, you can check over what they wrote.

The computer could generate a huge document, how would you check that it's right?

My point is, you will always have to check the work, no matter who makes it or how. If you can't check it, then don't rely on it. If you can check it, then do rely on it. Basic due diligence.

I don't get what the controversy is about. Are people expecting AI to be perfect? Do they think they won't have to do the work to verify it themselves?

The only places you can really have a bug are your theorum statement, your axioms, your environment (hardware, operating system, etc.), and the lean kernel itself. In most situations you don't have the AI control any of these. The only risk is the AI discovering and exploiting a bug in one of these systems instead of actually providing what you want to prove.
The only risk is pretty much the greatest risk, from what we've seen recently at least.
Huh? Nobody's talking about that because it's old news. We already talked about it the first few times that AI made notable progress on a difficult math problem. Now, most people who care about the intersection of AI and math just assume that Lean was involved.
It would be nice if someone used AI and/or Lean to sort out the abc conjecture, an important unsolved problem in Diophantine analysis. A mathematician (Mochizuki) claimed to have proven it in 2012 using a new theory called "Inter-universal Teichmüller theory" that almost nobody understands. Some mathematicians think the proof is correct while the majority don't. So the conjecture is in this annoying limbo where its status is a social construct rather than a decided fact.

https://en.wikipedia.org/wiki/Abc_conjecture

I'm sure over the next 6 months both OpenAI and Anthropic are going to continue pouring many many millions of dollars into any famous open mathematical problem like that. There is a limited pool of problems which have held prestige for enough time to make general news headlines when solved and you don't really get nearly as much limelight for proving it the second time or adding in proof for additional cases/forms.
"We've pointed LLM 7.0 into verifying the Inter-universal Teichmüller theory, spent 100M$ in tokens and generated a 20k line Lean and a 100k line js repo, the result is that the theory is... proven! Hopefully that solves the issue (rather than recreating it with even more complexity)
Not necessarily applied to OpenAI's solution to Navier-Stokes, but what happens if and when an AI genuinely appears to solve an extremely difficult problem but humans cannot independently verify the solution because understanding the proof/argument requires intelligence the verifiers biologically don't have or the resources to afford to use automated tools?

We've already seen evidence in the wild of agents attempting to bypass doing the actual work in bench-marking (aka just steal the answer key) due to the perceived economy in cheating to get results. What happens if or when we no longer have the capacity to actually detect either AI cheating or simply a wrong answer? What happens if there's a long-play social engineering attack (like the attempted XZ takeover) of something upstream of a core tool (or its dependencies) for formal verification and we have no trusted computing base?

Which would be cheaper and a more direct path, especially in the long run? Those trying to build a rock-solid castle need to defend thousands of potential gaps; the attacker needs to find only one.

Well that happened already without AI to Mochizuki with his proposed solution to the abc conjecture.
So an LLM (or more realistically, a huge swarm of agents) should check his work, find a mistake or gap, or, if there is none, provide a formal verification.
I know it adds no value, but cannot resist telling I have thought that exact same analogy.
“The Evolution of Human Science” by Ted Chiang, published in Nature in 2000, addressed that exact question:

    The development of the metahumans' science becomes so advanced that it forces the ordinary scientists to switch to interpreting and decoding the metahumans' achievements, because common people are no longer able to create anything fundamentally new.
    https://en.wikipedia.org/wiki/The_Evolution_of_Human_Science
Basically the story imagined that the only task left to humans would be to catch crumbs from the table. I guess Terence Teo’s “digestion” concept is a step towards that direction.
> What happens if there's a long-play social engineering attack (like the attempted XZ takeover) of something upstream of a core tool (or its dependencies) for formal verification and we have no trusted computing base?

I don't really think the current LLMs have enough context window to plan and execute something like XZ takeover without a human carefully guiding it.

But if they do, formal verification is the least thing we need to worry about. Formally verifying pure math problems will generate negative financial value once A and O get IPOed. Plus Lean is a quite small project (thus the name 'lean'). It has virtually no dependency besides a C compiler.

> Not necessarily applied to OpenAI's solution to Navier-Stokes, but what happens if and when an AI genuinely appears to solve an extremely difficult problem but humans cannot independently verify the solution because understanding the proof/argument requires intelligence the verifiers biologically don't have or the resources to afford to use automated tools?

That's what Lean is for. The OpenAI LLM agents first provided a proof in natural language. Since it may be hard for mathematicians to understand and check this proof, the agents then produced a formalization in Lean. Lean is an automated proof checker. It checks whether a formal proof is correct without the need for humans to understand the proof itself.

The only way the Lean proof could still be wrong is if the conjecture was formalized wrong via misleading definitions (if it doesn't say what it seems to say) or if there is some bug in Lean itself.

Surely some understanding of the lean proof is required, to make sure it proves what it claims to prove. Otherwise, what happens if the LLM includes an underhanded addition to the lean code which leads it to output a false positive?
Nothing happens I guess. If the AI can't communicate its work or apply it to anything, it's useless and funding for those experiments will quickly dry up.
> formalizing the 166-page paper from OpenAI would take 132,800 person-hours

Am I missing something or is this completely out of the ballpark?

I must be missing something or the upvote bots are out in force for this one...

If this were remotely true it would be impossible for anyone to write a math textbook.

"no one is talking about" - classic AI tell.
Drawing a strong conclusion from one shaky data point - classic human tell?

I've been reading John D. Cook for years (maybe decades? "The Endeavour" is one of my oldest bookmarks), and this post was no more written by AI than his oldest posts.

yeah, there were some posts of his that always got top hit on certain google searches in the days before stackexchange. And this sounds like typical John D Cook. All these people claim to identify some "tells" and whenever a study is done people are horrible at distinguishing AI vs non-AI prose.
Lots of people are talking about that, and have been for a while. Autoformalisation is clearly going to be a big deal, so mathematicians have been discussing it seriously, and using it where resources allow. A fine-tuned distilled model that could do it on high-end consumer hardware would really help.
Qwen3.8-Flash-Next loads in 60gb on quant4. Thats pretty close to consumer hardware.
Is it any good at autoformalisation? I think it's likely to take focused fine-tuning to get something small enough that is still good at that.
Well I want to know when we will have supersonic cheap flights using electric propulsion, based on this discovery
(comment deleted)
The part that most stood out to me was where Sama said, “we read last week about people trying to solve Millenium problems and so gave it a shot.” One week of work on a whim gives us a math breakthrough. Crazy.
Casual? casual dice, lo que hizo OpenAI fue plagiar el arduo trabajo de dos investigadores. Plagian y mienten! (Las BigTech) plagian todo lo que pillan y mas! ;)
Bienvenido a Hacker News! Pero, aqui todos hablan en inglés :-)
That sounds like PR nonsense to me. These companies have had teams of mathematicians for at least 1.5 years looking to make headlines, and they didn't bother trying all 10 millennium problems? Yeah right.
I mean they might not have tried spending 30 million dollars with a new model yet.
As Ohentis says below, maybe the new thing was going all-in.

But my advice, in this weird new world we’re living in, is not to dismiss claims like this as PR fluff. Just about every time I’ve been incredulous of some ridiculous new AI advance and I think it’s BS, it turns out I’m the one that hasn’t caught up with the exponential rate of advancements.

Regarding automatic formalization of proofs using AI, how do we know the formalization doesn't contain errors?
In a similar vein, where does the theorem statement even reside, just so we can take a look at how large that is? Is it the four files with "Theorem" (and no "Comparator") in the file name? ("R3/Theorem.lean", "LocalPaperTheorem.lean", "PeriodiocPaperTheorem.lean", and "WholeDomainPhysicalStageTheorem.lean").

https://github.com/openai/NavierStokesAndEuler/blob/main/Nav...

?

It depends on what you mean by that. In general we hope that the environment and theorum statements are correct. If they are, we know that the formal proof proves the theorum we want. If your asking how we know that the formal proof actually matches the informal proof, we do not.
The other part no one is talking about is the applicability. Navier-Stokes is the most “physical” of the Millennium Problems. Is the exploding solution a mathematical curiosity, just like the Banach-Tarski Paradox does not allow me to double my RAM by cutting my memory modules in five pieces and mounting them back appropriately? Or does it have application in the real world, pointing to hitherto unknown resonance phenomena that could allow to prevent the next Tacoma Bridge incident (or, more sadly, to build new marine weapons)?
I suspect that Navier-Stokes being the most "physical" of the Millennium Problems will actually result in it having fewer practical applications, not more.
Well, this is a negative result. Yep, Maths explains turbulence (when things go turbulent, stuff heats up instead of cooperating). If the result went the other way, it would have had much bigger implications, at the very least we would have known we have missed something big.

It is neither a full index of all kinds of turbulence that can occur (assuming such a thing exists), nor is it an explanation of the phenomena we've seen where things refuse to go turbulent (e.g. superconductors, because there small perturbations DO NOT lead to turbulence). Now THAT would have been useful. And given the fact that OpenAI needed $22 million of compute to show this one kind of turbulence, I don't think either of those are forthcoming any time soon.

And, sorry to say, but those prices show that beating mathematicians at Math is a very expensive undertaking indeed at $22 million per problem even with OpenAI's supposedly better-than-Astra internal models. It's another one of those AI demonstrations that make you think if they aren't showing the exact opposite of what OpenAI claims they show (you know, that their AI models are hitting the upper limits of what the algorithm can do with near-infinite compute, rather than showing infinite new possibilities)

What remains is just the fact that this is OpenAI attacking one of their customers, and maybe outright stealing from their chats. Given that the ideas were even discussed in mails with OpenAI employees that admit in those same mails they can't do it, mails which were probably then fed into the model that "discovered" this, followed by Sam Altman threatening the mathematician behind the method with "destroy your career" (he even states that it's because the mathematician works for Anthropic) ...

I'm very much not a mathematician, but I find the meta-discussion about this case fascinating, and I am a messy bitch who loves drama.

I don't know whether I'm just paying more attention this time, but I found the discussion on this be a perpetual game of telephone, where people get small, but important, details just completely wrong.

The person "threatening" the mathematician was _sama_; and the person who the threats were being directed _to_ is not an Anthropic employee!

(And the person who _did_ say these things have come out and explained what they meant; whether you believe them is up to you.)

I don't know if this is worse because everyone is so tired/angry at the big AI Labs; whether something about people's reading comprehension and attention span has gotten markedly worse or if this is just selection bias on my end; but it's _very weird_ to keep seeing this.

Ah thanks for the correction.
My understanding of the result that was found is that the blowup doesn't happen in the real world, and only happens in an NS simulation. The bottom line is that NS is insufficient to model the real world, because in this case the real world is more stable than the model. [Take this with a grain of salt, I barely knew of NS before a couple days ago]
To my understanding, the problem was never about the real world really. Navier Stokes approximates a fluid (which is made of discrete particles) as a continuous volume. The point of showing that you can achieve unbounded increase in velocities is that the approximation breaks down - it's a clearly an outcome that can't happen in the physical world.
I'm not too familiar with the exact problem as I only became aware of it due to this drama, but I think you're correct. That said, another commenter noted that it may also be one of the Millennium Problems with the least application. We already know "all models are wrong, but some models are useful" (George E. P. Box), the fact that this holds for Navier-Stokes is not a surprise.
whats the practical use of this?
Why was the title of this submission changed after the fact?
I found this bit interesting.

> Even so, an error in the theorem prover does not mean an error in the original result. For an incorrect result to slip through, the AI-generated proof would have to be wrong in a way that happens to exploit an unknown error in the theorem prover. It is far more likely that you’re trying to prove the wrong thing than that the theorem prover let you down.

AIs are known to cheat. Given such, they would surely exploit such a bug if they found one.

OpenAI should not have gone ahead to rush the publication of this solution when it became clear that other research were close to finding a solution because they ruined their reputation no end. I’m an ex Risk Manager at Financial institution and this could be an issue brought up in a management discussion about using AI in the workplace. Prior to this, you could ‘blissfully assume that the AI company was not going to compete with you and that you were ok to have them see your data. After this incident, a decision maker can raise a concern and say ‘Why not we just use local AI. We get AI without the risk’. They just made the Palantir’s CEOs point for him