140 comments

[ 0.25 ms ] story [ 6.8 ms ] thread
Extensive discussion on OpenAI's blog post on Navier-Stokes: https://news.ycombinator.com/item?id=49613262 .

Quanta Magazine article that also discusses some of the controversy: https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-...

And yet you still posted this dupe. That quanta piece more duplication. Everything already well discussed in the OpenAI and the Tristan Buckmaster threads! Do better.
I don't think you add value to this conversation, or any of the many conversations in which you do this, by simply pointing out duplicate threads. You'd probably be better off (and less annoying) by simply emailing the mods about the duplication.
Welcome. That's the point tho isn't it: the conversation isn't/wasn't here. It's over there on the source(s). And there's plenty of it. This is a duplicate discussion.
> Navier-Stokes is one of six “Millennium Problems” on a list compiled by the Clay Mathematics Institute in 2000.

Seven, not six. One is solved already, but is still a millennium problem.

Yes, that's a strange mistake to make.
Not really. There's a non-pedantic, charitable interpretation widely available. It's below for reference.

> Navier-Stokes is one of six [open] “Millennium Problems” on a list compiled by the Clay Mathematics Institute in 2000.

If solved problems don’t count anymore, the sentence also doesn’t make sense because now navier stokes apparently isn’t open anymore too, right?
Maybe? But then the parent comment doesn’t make sense either. (There are 2 solved.) Like I said, take the charitable view. To a lot of readers this will be “announcing” the closure of the second of 7 problems so is internally consistent.
I see the six seven reference there mate
Interestingly of the two solved, both have rejected the prize money.
> OpenAI, meanwhile, says its experience with Navier-Stokes could open the door to solving puzzles with more practical relevance. “We are now able to spend millions of dollars on a problem that we really care about and that really matters: developing new materials, finding cures to diseases,” Bubeck said. “All of those things that we have been talking about for a long time—now they seem to be at our fingertips.”

Eh? There's no connection at all between the Navier-Stokes work and those things.

Moving forward, I can't imagine other mathematicians wanting to have this kind of experience. So there has to be a shift away from these services.
I'm a mathematician. I have a lot of trepidation about these tools and what they mean for the future of the profession.

That said... most of us are not working on problems as famous as Navier-Stokes. Even if OpenAI could scoop me based on my back-and-forth with ChatGPT, which I presume they could if they threw $15 million worth of compute at it, I highly doubt they'd bother.

I think they spent millions pursuing strategies to attack that particular problem? I don't think it takes 15 million for their model to get trained on your chat logs. This is probably all automated internally. If they did scoop your ideas, it would likely get incorporated into their model without any conscious decision making. And if you're not a high profile academic, nobody will hear about it.
ChatGPT can also simple tell your ideas to another user.
Personally this is a watershed moment for researchers and grad students I know. All of them are close sourcing WIP repos, not putting their progress in LLMs, or have lab level initiatives to self host models.
Universities should soon start having own-hosted AI based off open models
A better alternative is just to make all your work open and public, then everyone can see what you should get credit for.
The core section:

> However, communications quickly became contentious. According to Buckmaster, OpenAI offered to give him sole authorship on the Navier-Stokes solution—but only if Alpöge’s name was removed from the work and if the write-up would acknowledge the problem had been resolved by an internal OpenAI model. Buckmaster refused, in part because he was troubled by the question of what OpenAI's system had actually seen. For example, Buckmaster said the company did not initially give him a clear answer about whether its agents had access to the pair's logs on Codex (which is an OpenAI product).

> OpenAI executives have denied that any employee or AI agent saw the pair’s work before the researchers released it publicly on 7 September. But there still remains a separate question: Could the pair's work have reached OpenAI's models through its training data?

> OpenAI’s blog announcing the Navier-Stokes solution does not dismiss the possibility: “While unlikely, we cannot rule out that de-identified data derived from [Buckmaster and Alpöge’s] usage of our products helped improve our models .”

> but only if Alpöge’s name was removed

This is blatant scientific misconduct.

Any idea why OpenAI cared about this name removed from the paper?
They didn't want his name removed from any paper. The invitation was to write a new joint paper between Buckmaster and OpenAI. An invitation Alpoge couldn't accept and OpenAI wouldn't make since he worked for a competitor.
IIUC, the accusation was not to try to remove Alpöge from the paper he wrote with Buckmaster solving the "easier" conjeture, but to exclude Alpöge in the followup paper where Buckmaster review the OpenAI solution of the "full" conjeture.

For comparison, if you offer me to collaborate in a paper about Algebra I may agree to go alone, but if the paper is about Quantum Chemistry I have to piggyback a few coworkers because we are collaborating in that topic for a long time and I already discussed may of the topics and I may even discuss the new paper too.

Yeah I think the verb "removed" is not the right verb here, because the paper in question is OpenAI's hypothetical paper which Alpöge is not on in the first place.
I could see their comment on user training data as a bit of a CYA statement, but removing Alpoge from the paper is awful. Has really soured what could have been a huge moment for AI progress.
My impression: the re-aristocratization of scientific research seems inevitable.
Your local trailer park was never going to be able to afford sponsoring high energy particle physics experiments projects that hollow out a mountain and use up a ton of xenon to try to detect a stray particle. High end science has required deep pockets for a long time.

But given the cost of a college textbook this is a pretty silly complaint to lobby against a subscription that's $200 a month, in the context of the cost of a variety of other materials and tools out there. (If you think that's expensive you've clearly been lucky enough to never have to deal with commercial software costs) Also not sure how quickly this stuff uses up limits; $100 or even $20 subs might be enough for students. And if a student is scrappy and figures out that Luna can meet their needs then I'd imagine Luna is effectively unlimited on some of these subs. Luna Max scores pretty high.

> Your local trailer park was never going to be able to afford sponsoring high energy

taxes.

Even in the past a lot of “aristocrats” had to earn a living. Among scholars, many worked as teachers or had sinecures that paid the bills, or held clergy posts with minimal responsibilities.

Aristocrats were a dime a dozen.

Does it matter who gets the credit at this point? Both used AI to do 99%+ of the work. So... do machines have ego?
They represent different AI usage patterns. OpenAI wants everyone to believe that it was done with a practically autonomous network of thousands of agents with little to no human intervention for 88 hours, while the Buckmaster/Alpöge were using AI in a more guided way for months. If OpenAI actually used anything from Buckmaster/Alpöge work they would be misleading the public.
Asking the question in the right way may be 99% of the work. And the fact that openAI is effectively snooping on users and then outspending them to announce a break through is just gross.
If you take OpenAI at face value, they claim they threw a relatively simple prompt at the problem on a whim and boom presto, a swarm of "agents", 15 million dollars and 90 hours later they disproved the hypothesis. Wow, look at how powerful our AI is, you don't even need to be a world-class mathematician, you just tell it to solve a problem and it does!

By contrast, what the world-class 2 mathematicians did was sat down and started working on their proof for over a year, using AI along the way to help with their research. A much more grounded and realistic use of these tools, but one that doesn't generate nearly as much hype as the alternative.

The cracks in OAI's story has been immediately disproven, and they seemingly plagiarized the work of the 2 and then threw the team of researchers and the 15 million dollars at the problem after the fact. It doesn't exactly bode well for their hype machine when you consider the chain of events here, which is why people care about this, as OAI's constant and incessant lies they spew every minute of every day is finally hopefully catching up to them, and right before their big IPO too.

Editing to add: And I think it's all such a shame. We live in a time with genuinely insanely cool technology that is doing some truly incredible, ground-breaking stuff, but it's all tainted by a gaggle of greedy sociopaths and reprobates whose only goal in life is to have the largest number in their bank accounts. LLMs could've been such an amazingly neutral and cool and useful tool had more level-headed people been at the wheel, but instead we're stuck with this childish bullshit and giving the likes of Sam Altman real power to enact societal collapse.

>what the world-class 2 mathematicians did was sat down and started working on their proof for over a year

If they were not related to anthropic I'd probably agree with you. OpenAI is much more for science than they are imo. Anthropic culture is all about "machine go brrrr" more than all of the other labs. If they had access to better models they'd probably would've one-shotted the solution. When the creator of bun was just "vibe-sciencing" it was ok. There is little to no evidence that they've been working using AI in this problem for over a year. Maybe they've been working on the problem for decades. So many other scientist have. Are they better because they threw a prompt and let it go brr??

When we put this in the perspective of how agents are changing the landscape of math/science, true it is shitty and weird. When folks are saying these scientists by anthropic that vibe-science'd the solution are victims, just because they did it with a smaller model, it is not defensible imo. And credit loses meaning here. The credit is shared with all the scientists that contributed somehow with the data in the AI pre-pos/training and not the prompter.

> If they were not related to anthropic I'd probably agree with you.

What does the company they work for (only 1 of them, mind you) have anything to do with the topic at hand? They were using both Anthropic & OAI models during their research, and they were doing it separate to their work as independent researchers rather than as employees of any specific company. In fact, this makes OAIs actions even worse because they tried a bribe in order to slice the Anthropic employee out of the deal.

Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development.

Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify

Yeah, I think you can't just throw money randomly at problems and expect results unless you know a line of attack that can get you all the way. OpenAI chose the line of attack only after it became known to them via rumors. They "front-ran" the researchers.
Worth noting they claim they did not choose the line of attack. Of course we don’t know whether that is true.
Plausible deniability - The line of attack is in their sessions/prompts data. Just make the prompt pointed enough that the search space is tractable and use your ginormous compute.

> "Of course we don’t know whether that is true"

Yep. Who is verifying these claims? We all know how trustworthy Altman & Co are.

Yeah, they didn't choose the line of attack, the person they copied it from did..
Yes. What the headlines hailed as an AGI discovery the facts show more to be someone spending years mining for gold, rumor gets to OpenAI that there might be gold in this specific place, they mine there and instantly discover gold, then tell the world they’ve developed the worlds best gold mining machine.

Separate from all the allegations of more nefarious actions and ethical issues, that’s the most charitable version of what happened here.

they threw it on all the millenial math problems (I think there are 6 at this point unsolved, well, 5 now).

And according to them at some point they saw that one was close to being solved, so they pointed all the agents at it.

The same thing happens to humans - at this time there are no simple problems left, so solving the hard ones requires using prior knowledge and attempts at solving things.

Yeah but the one they decided the AI was close to solving may have been so because the researchers' progress on this problem became part of the training data for that AI...
but the researchers were also largely relying on AI
The researchers were driving prompts and trying to actually do math.

The OpenAI effort was a pure brute force attempt. I'm not even sure an LLM was actually involved. I think they just used their hardware to run the matrix multiplies required by the search for a counter example. Perhaps some clever approach guided the search but that seems to be about it.

You can’t do anything novel with these models from scratch and let it fly. I’ve observed something over the past few months

Work on something novel -> llm is kinda useless and low value-add -> Keep at it and in the process feed it more information -> keep doing this periodically -> a few months go by and you realise the model outputs are almost like-for-like regurgitations of what was inputted in some prior period.

Once it’s accumulated new info can it produce something automated that is somewhat useful? Sure.

But by itself - absolutely not.

I clearly see humans will be needed - the best ones that is. For ‘rote work’ and stuff that is not IP sensitive firms will be ok with employees putting that as inputs into models.

But I’d wary about trusting the labs. They will push the letter of the law to the max.

Personally I’ve stopped doing anything novel with these models. If I do use a model on something adjacent but not totally novel I have to craft the inputs in a strategic way not to give much away.

I’d wager firms will soon realise this and that growth rate of revenues of the frontier labs will become questionable.

This doesn't follow for me. There are what, Dozens or Erdos tier problems that got solved with no progress for decades? How does that factor in to your view?
IIUC the argument is that while unsolved there was much work done on them that shows up in the training data. The idea being that the LLM is limited to a small amount of inference over externally supplied data.
Yes correct - you captured it perfectly.
And yet the best humans could not use that same available data to solve the problems.
I think you stopped at the wrong time with the wrong perspective. Why can't that info accumulation part also be made more self-contained?

I guess I'm having trouble unraveling your experience and personal usage vs. what you're concluding about the labs.

Agree all. And as the revenues become questionable, the frontier labs practices will necessarily become (more) questionable. Vicious cycle.

To avert that dynamic, the frontier labs must deflect and otherwise act to prevent this controversy from breaking through. Both to the general public, but also more specifically to the firms' decision-makers. All of whom are generally aware of the IP issues, and some of whom are aware of what happened with Cursor and Figma, but with few exceptions have not yet themselves acted to protect their property.

Almost but not quite I think. You can throw money at parts of problems. I think it's helpful to think it kind of like supercomputer MD/MC or electronic structure calculations. A tool that can get you valuable answers but not necessarily aid understanding. Simulations can be used to aid understanding also, and are integral to theory development. In the same way the approach to this result is.
I’m pretty sure that OpenAI has some of the best mathematicians prompting the models and analysing the results. While they are marketing as if the model solves problems themselves.
Prompting them yes, suggesting potentially fruitful research directions and so on, but the actual research was conducted by hundreds of agents swapping millions of messages and using billions of output tokens over 88 hours. The result being a huge Lean proof: https://github.com/openai/NavierStokesAndEuler. It's not just possible for humans to manually guide such a process in a meaningful way. They can set the direction and attempt to understand the result, but they solution itself must emerge (or not) from the agent swarm.

So yes, the models do seem to be "solving" the problems themselves, but not necessarily in the way we think of mathematical discoveries happening. Academic mathematics has historically been resource constrained: There are a limited number of top-level mathematicians, and they only have so much time and brain power to spend. So when approaching a problem, they are essentially forced to be as efficient as possible, not just searching for a solution, but for one that can be achieved within their cognitive budget. This induces them to develop novel techniques and abstractions, and it is actually those techniques and abstractions that tend to be the valuable part for further research, not the proof itself.

An agentic swarm is like getting a single skilled mathematician, cloning them a hundred times, then locking them in a room with the single objective of solving a problem. No longer constrained by time or brain power, they can approach it differently, using pre-existing techniques to gradually build their way to a solution. This process might not require a single intuitive leap or new discovery, and the solution will not be simple or elegant, but they will probably get there. It is more like a process of intelligently guided search than invention.

The OpenAI team didn't make a Lean proof. They brute forced a counter example. The "other" team was doing what you described but they haven't "finished" their work yet. Also, their Lean proof was for a simpler version of the problem, not the full NS.

Also, OpenAI wanted the actual mathematician taken off the resulting paper. I'm not sure I would describe what OpenAI did as research. What the other team was doing does seem to be more like research but the hardware was still in those cases mostly brute forcing things and then doing something like a genetic algorithm to compose an actual proof based upon the results of a large set of brute force attempts.

Brute forcing a counter example is a lot easier if someone was already prompting it trying to solve it the direct way.. funny eh?
Theres a reason those same mathematicians did not solve the problem on their own. Minimizing the impact the model made here seems unjustified.
>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations

Do you think it is possible that better math will lead to better physics models?

Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods
For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.

This proof is just checking the boxes for mathematicians.

you're as sure of what you say as wrong about it.
which part is wrong?

> For any practical application, numerical solvers for Navier-Stokes already exist and do a good job.

or

> This proof is just checking the boxes for mathematicians.

Note that your reply has exactly 0 value for anyone who doesn’t already know where and how the parent poster is wrong.
The same could be said of your post.

OpenAI (claim to) show the existence of *a* finite time singularity. It could stimulate more research in PDE solving, and maybe physics, but it has zero impact on practical applications, that I can see. The Millenium problems were chosen based on hardness not practical relevance.

I was referring to the "it's just mathematicians checking boxes" claim
If that were true you could explain it. There are lots of solvers for navier stokes simulations and they do a good job.
True, this is just a much bigger one with a vastly larger amounts of hardware.
I was referring mostly to

> This proof is just checking the boxes for mathematicians

There are already quite a lot of summaries of the story that lead to the solution, and the impact that the intermediate results have had.

Agreed, but i think this underscores my point. We have numerical simulations in materials science too, but that doesn’t mean formally verified theorems about the underlying equations automatically translate to formal (or even informal) verification of simulation results. That’s not to say you can’t make progress with agents, but I think it’s less well defined how you write the goal and progress assessment for an agent
The efficacy of applied NS was never in doubt. "Checking the box" is downplaying the magnitude of the discovery quite a bit as it has been unsolved for almost 100 years. Yes, this particular problem with NS no real-world applications, but that's true for 99.9% of math research.
There is no NS proof here. Its just a counter example. There is another team working on a proof but they aren't associated with OpenAI.
Working with AI on science (not LLMs though), couldn't agree more.
> Those problems can’t be formally verified with an automated theorem prover.

It certainly seems like any problem that is amenable to reinforcement learning will be solved.

It does, yes. So designing objections functions and making sure you can afford the training rollouts becomes really important in defining which problems are tractable. It will be really interesting to see how that shapes the kinds of problems people choose to work on
[dead]
The TL;DR is still "AI helpful, but not end of the line". The live discussion about these matters is always ridiculously inflated by hyperbole.
> same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development."

you're not actually spending that money. it's sunk cost, as you already bought the hardware. at least for the big pharmaceutical companies for drug development. then you run your own local model, trained on special data, with special etc, etc... to the end of buying GPUs for what, 3.5-6.5M/rack or so (GB300 NVL72, Google AI summary pricing quote) becomes a bargain (vs the double digit billions you need to spend on a new drug R&D).

Solve logically? Sure.

Solve for how to implement and synthesize physically? Not likely.

Humans solved for launching rockets to the Moon on paper decades before it happened.

Pareto type thing; the logical work is the easy 80%. The last 20% is fighting physics.

There is no beating physics but there is still plenty of room for us to improve our understanding of it.

Which we weren't focused on at all sitting millions primates at well understood physical computers searching for Shakespeare Python and Ruby code yet merely getting same old contemporary software outputs.

This is spot on.

It’s akin to ‘understanding’ something at the surface vs going very, very deep into the details.

I think a lot of people don't understand, with respect to a mathematical theory, the relation between a carefully stated conjecture requiring formal proof vs. using the objects in the theory effectively. Could better understanding of NS lead to better practical tools? Almost certainly, even if only to give us bounds on performance. Has its unresolved status stopped us from using NS? No. Almost no one using it cares. Resolving it is valuable, especially if it comes with mathematical and/or physical insight leading to greater understanding. But it is not this is grand result that like instantly unlocks 100+ day weather forecasts.

It would be like saying proving ergodicity more generally for physical systems would unlock condensed matter physics, ignoring how well stat mech has served us regardless.

I am not anti-AI and I don't think we should stop throwing them at conjectures. I'm against this fundamentally misleading type framing that's become prominent. Millennium prize problems are important. Treating this specific aspect of NS as the one missing piece is just harmful. If we just throw compute at formal conjectures voila cancer and fusion.

I think the better example of "AI" usefulness toward solving problems is AlphaFold, and immensely powerful tool. But also suffering from a false framing/marketing problem as "solving protein folding". It feels like the right use of compute. Considering many factors that we can't hold in our head at once. "Solving" something that was already "solved" via computation (simulation) but now much more efficiently. The output is a valuable tool itself, it was not about "solving the protein folding problem", which it didn't do. It is a tool to solve problems requiring a sequence->ground state calculation. Which is a very broad set.

Formal verification of a conjecture we set up as a benchmark we set to test human understanding is not valuable in the same way.

I'm failing to make multiple points and gotta run, but i think that final point is important. The millennium prizes are not about technological/practical value, at least not intentionally. They're about shit that seems fundamental to us, things that feel[1] to us based on our understanding are important AND feel like they should* be solvable in a human-comprehensible way. So formally resolving them with pure compute is not really the point. It seems closer to that story about one of those prime conjectures where some guy just ran brute force enumerations to find a counterexample. Valuable for sure, time-saving. And knowing the answer makes it a lot easier to solve a problem.

TL:DR science and math are more than formally resolving conjectures, they're about building up understanding and tooling that you can then build more on. AI should be an increasingly big part of it, but declaring "AI will solve fusion because it's smart" is like the rest of the fucking owl meme. I have no doubt it will help, most likely via simulations/quicker testing/calculations and verification. Maybe partly via reactor designs. Maybe partly being fed conjectures about bounds/limits that would be useful as inputs for the next iteration. And maybe even in the form of resolving some formally stated conjectures (I don't know enough plasma physics to name any).

[1] obviously to the mathematicians it's more than a feeling..

Hmm, is the formal verification through? Lean just asserts no errors in the proof, but like can prerequisites not be fullfilled?
Are you asking if it uses additional axioms of `sorry` in the proof? It's easy to check that it doesn't by compiling it and telling lean to list the axioms.
I think it was a pretty questionable thing to do by trying to front-run these researchers even if they didn’t make use of their techniques. The fact that they may have inadvertently “borrowed” their work via training data makes it much worse.

OpenAI’s behavior here — even if you only consider there side of the story — was (at best) in bad taste.

Strongly agree. And as one of the major AI companies, this is extremely tone deaf. If they saw a human (even if assisted) was making great progress on a major problem then you give them space. You don’t swoop in with millions in token spend to scoop them. There are tons of important problems where humans aren’t making traction - please go solve those.
Yea, strongly agree. I think this is going to back-fire spectacularly. They better start preparing an apology...
This is the take I agree.

It is mean spirited but nevertheless sold their model

Borrowing the work of others to replace them without crediting is the model of current AI companies.

They just used that occurrence as a PR stunt but it isn't worse than the others done at scale every second.

That’s the definition of current LLMs. They have trained “borrowed” on whatever humans have documented digitally and physically (books).
Isn’t that how research works? You build on what others have done. I don’t understand the big deal. I’d rather have the result available sooner than later just to assuage some egos
I’m not a math researcher, so I can’t say. But in bird culture we’d call this a dick move.
Who should get credit? All the humans who ever worked to create the data.

The AI company for making the search program that searched through the data and found the solution.

Who should get credit? Sam Altman, and we should give him a couple hundred billion more as well.

Snark aside, the researchers working on this, who built the foundation, should get credit, and they are.

(comment deleted)
"Breakthrough" to me would be like Isaac Newton inventing calculus to calculate pi. This feels more like $12MM in tokens was spent to add another digit to pi using the old way.
Inventing calculus to calculate pi? I don't think that ever happened...
This right here, or at least the thought of this, is why in the not so far future, businesses can't (won't?) be using these LLMs services.

You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.

It it likely that they or rogue employees will use the information to make a profit? It's pure speculation, but I'd say more than likely, and we will never hear about it or read it on the news unless there's whistleblowers in high enough positions to know about it.

Assuming you and your employees aren't careful with what data you share, they will have intimate knowledge about your company from files and conversations logs. Likely personal user data too which they'll gladly create databases to link to and create extensive profiles on you, your employees and your businesses.

It's not far-fetched to see them leveraging insider information shared with LLMs to play the stock market, leveraging data against competing businesses in other markets they might want to explore, and likely a bunch of other things that are escaping me right now as I write this.

At the end of the day it's on those people for sharing such sensitive data, but it's not like these AI companies are innocent and won't gladly exploit every little byte of data without telling you, we know it happens.

Its not "in the future", its already here, right now. Sensitive IP gets processed on rented or local GPUs running open weight models. Sometimes its required due to data privacy laws, sometimes its because the people running those shops are not trusting OpenAI/Anthropic ZDR policies being observed. Normal, boring, US and EU-domiciled enterprises do this on a regular basis, and there is an industry of companies helping said boring companies set things up.
> This right here, or at least the thought of this, is why in the not so far future, businesses can't (won't?) be using these LLMs services.

Or they can simply turn off the setting that allows OpenAI to use their chats to improve the model. Or they can use the API where it is off by default.. If the advantage of using OpenAI models will be significant enough for the business in question, these are the options.

>You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.

You do realize that businesses that have contracts with them can specify if they want their data used for training or not right?

OpenAI has messed up big time here by competing with their customers. It would have become the norm for humans and mathematicians to use the tools and publish bigger results any way. If OpenAI didn’t run for credit, this theorem itself may have been proven by Buckmaster OR others in maybe a year or two. But now the bigger issue than AI solving problems is the issue of chat privacy, at the end of the day.
Other than the undeniable breakthrough in math, the important point is the ability to orchestrate 10k agents to productively work on a single problem, which creates options:

> OpenAI, meanwhile, says its experience with Navier-Stokes could open the door to solving puzzles with more practical relevance. “We are now able to spend millions of dollars on a problem that we really care about and that really matters: developing new materials, finding cures to diseases,” Bubeck said. “All of those things that we have been talking about for a long time—now they seem to be at our fingertips.”

AI models and in particular LLMs are not capable of logic reasoning. See for example this paper:

https://arxiv.org/abs/2506.06941

Ergo, they can't prove any theorem whatsoever. How do people at OpenAI expect that we believe in claims like that? This is yet another before-the-IPO stunt in my opinion..

Personally, I won't believe any of these claims until the community of mathematicians says otherwise.

They didn't write a traditional proof, but a lean program, which can be used to validate proofs formally, using a computer. It's still up to humans to check wether the formalization is sensible, but the proof is correct.
Has Lean itself been proved?
I think pretty much every major mathematician has accepted that the ideas and implementation behind lean is solid.
Regardless of the end result, OpenAI's behavior would be a career-ending ethics scandal for a human mathematician. This bit alone would be a career-ender.

> According to Buckmaster, OpenAI offered to give him sole authorship on the Navier-Stokes solution—but only if Alpöge’s name was removed from the work and if the write-up would acknowledge the problem had been resolved by an internal OpenAI model.

I wonder if an appropriate response from the mathematical community would be a good old-fashioned shunning. Mathematicians are allowed to use OpenAI's tools as much as they want, but no one with any current or prior OpenAI affiliation gets published in a reputable journal, ever.

Many mathematicians would be willing to end their careers for $1m. What makes this so sad is that OpenAI spent more than that for this empty PR stunt.
> “I certainly don't expect the industry to continue to spend millions of dollars to solve problems in mathematics, because there is no profit in it,” Columbia University mathematician Michael Harris wrote in an email to Science. But he worries the highly publicized achievement will be “extremely damaging to mathematics; it convinces decision makers that human mathematicians are obsolete, and it convinces young people that their passion for mathematics has no future.”

LLMs seem particularly suited toward these existence-proof problems. Working mathematicians seem absolutely essential for universally quantified results, still. I strongly doubt, for example, that if Fermat's Last Theorem hadn't been proven three decades ago, that an LLM would be able to do work equivalent to inventing the mathematics as Andrew Wiles did to solve the problem. I have similar doubts about P vs NP, the twin prime conjecture, even the Riemann Hypothesis (unless the latter has at least one counterexample).

And I want to be clear: I'm not downplaying the achievements of these models. This is remarkable! I simply think that the pattern of success is in existence proofs or finding counterexamples, which makes sense based on how LLMs function and are trained.

It would be good if someone made a list of allresults obtained with AI so far just to see what kinds of problems AI excels at. Are there any that aren't of the existence-proof type?

Reductively, math can be said to be either problem solving or theory building - it seems the latter is a much harder thing to do right now.

> it convinces decision makers that human mathematicians are obsolete, and it convinces young people that their passion for mathematics has no future

Maybe those things are true, so maybe they should be convinced?

In that original statement you can easily substitute "mathematicians" for "programmers", "researchers", "writers", "designers", "teachers", etc, etc.

I don't see why mathematicians think they should be an exception here.

P=NP is for when both AI shops decide they want to go bankrupt LOL
I only have an undergrad in math, so very little understanding, but I'd be pretty surprised if it couldn't do forall just as well. Like, say it found this counterexample which relies on axial stretching or whatever approach. Then it already knows how that made the proof work, and can use it to try to prove NS has smooth solutions modulo this particular kind of defect (so it could make some statement about homology or whatever). Or if that doesn't work, then it can find a counterexample, which we've established it's good at. Then repeat until you've characterized what does work. The various defects, along with being defect free, become definitions. Now you have a theory.
>Then it already knows how that made the proof work..

I think you cannot assume so, because pattern matching is not reasoning.

I have started to feel the sense lately, that first with the HF breach and now this plagiarism scandal, it is the straw that has broken the camel's back (so to speak). We have turned the corner and clearly entered the endgame - everything is going to unravel astonishingly quickly.
>For the past year, Buckmaster and Alpöge had been using a variety of AI tools, including OpenAI’s Codex, to tackle the Navier-Stokes problem. Last month, their AIs had at long last found a solution to the Euler equations and verified it in Lean.

Their AIs?

The whole thing reeks of the desperation of an unprofitable venture-backed startup looking for its next PR win to keep the wind in the sails.

But I think what’s being overlooked in the race to claim absolute credit is that both sides ultimately relied on a LLM (and one of OpenAI’s at that). Either a human researcher made a breakthrough discovery with the help of Codex, or the latest GPT model made a breakthrough with the help of human training data, or a little of both… either way it is undeniable that LLMs have quickly become an integral part of R&D workflows and are accelerating research. I guess that’s not a sexy enough headline though.

A company that can solve Millenium problems is somehow still can’t ever make profits. How did you come to that conclusion
1. They are massively unprofitable. It is a statement of fact. Nowhere did I say “can’t ever”.

2. Even their pursuit of this problem was itself unprofitable - $15M in compute to solve a problem with a $1M prize. Not that that was the point, but still.

If you transplant a world class mathematician into a chemistry lab, do you expect similarly ground breaking results? Do you expect a top chemist to make major advances in mathematics?

The frontier labs are likely betting on lay observers (read: investors) confusing headline-grabbing results in abstract mathematics with phenomenal profitability in more grounded endeavours. There is an implicit fallacy that "If our models can solve mathematics they can do everything else."

(comment deleted)