Puts the onus on the AI companies to provide a specific replacement mechanism, no? Unless I'm unfamiliar with something else he's written that proposes something more specific and constructive
To the author's credit he obviously identified the problem very clearly and admits understandably "we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later".
> Puts the onus on the AI companies to provide a specific replacement mechanism, no?
Why? If someone makes an innovation that undercuts the underpinnings of some existing institution, why are they are responsible for cleaning up its failure?
How specific and constructive it is might be debatable, but he has tried to make concrete recommendations earlier; see e.g. slides 46-51 from the ICM talk: https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.p... – obviously there's some way to go still.
Not quite, all the proofs or disproofs so far AFAIK were using existing methods that humans developed and were already using to attack the problems, but AI is just more thorough. What AI can't do currently is develop new mathematical methods to attack problems that can't be solved with existing methods and AFAIK there is no known path to get current gen AI to do so.
Yeah. My other favorite example are books. Why do nonfiction books exist? There are some pathologies and corner cases, but fundamentally: to develop and share new ideas. Downstream from that, if it reads well and if you're lucky, you make some money.
But now, LLMs can generate hundreds of books per hour. They make up 80-90% of new arrivals in many nonfiction categories on Amazon. They short-circuit the system, allowing their "authors" to extract money from the system with zero effort by crowding out human work. And it's not even the question of whether these books are good or bad (although overwhelmingly, they're terrible). It's whether it's actually accomplishing anything worthwhile, or just destroying incentives for humans to write or go into any other sort of intellectual work.
In fact, I see many professions push back. Artists, writers, now mathematicians. And I'm amazed that our profession doesn't and that we have so many people who are hooked on vibecoding. I'm still waiting for that 10x payoff. All this velocity and somehow, the landscape of the software I want to use still looks the same as it did in 2021.
It doesn't, that's the whole point. You can produce output that passes the smell test with naive buyers in a matter of hours. If you want to write good LLM books, then you gotta work closer to human speed - weeks, months - which allows others to produce 100 slop-books in the same timeframe. You still lose.
I'm not convinced. I don't buy that anyone can just generate a slop book in no time and become a commercially successful author. There just aren't enough paying customers of books for that to be viable.
Said differently, LLMs might make writing faster, but they don't make reading faster. Given a fixed readership budget (or, perhaps, an envelope within which the real budget oscillates depending on various factors, like how popular reading is at any given moment), large amounts of LLM-generated books change how that budget can be allocated, but the net effect overall is actually an increase in selectivity from readers, not a decrease, relative to the overall proportion of available books being created.
Playing a devil's advocate. Why do we need understanding ? To take an example i would say ~99% of the population do not understand how combustion engines or how semiconductors work, what say another 1% ?
Is the fear post-apocalyptic in nature ? We need some human priesthood to carry on tradition ? why ?
Let's assume in the next decade GPT-7 PRO Ultra is cheaply ubiquitous, inspectable, reproducible, transferable, reasonably un-constrained by any institutional interests.
It’s not going to be ubiquitous? There hasn’t been a single frontier model where generation n costs less than generation n-1 to run. So the reasonable thing is to assume that GPT-7 will cost even more than GPT-6, and more and more of the frontier of knowledge will be locked behind a giant paywall. Participating in any field will mean ponying up to the oligarchs that own the infrastructure that runs the model.
> There ought to be more to life than sitting in a pod receiving sufficient nutrients from a tube
Even in a hypothetical world where AGI can do everything that humans do better, this would still be a choice to make, not an inevitable consequence for everyone.
And life doesn't require suffering to have meaning.
I grew up in a cult. Based on my experience, I believe that the most dangerous thing a human can do is to allow someone else to do their thinking for them.
> Let's assume in the next decade GPT-7 PRO Ultra is cheaply ubiquitous, inspectable, reproducible, transferable, reasonably un-constrained by any institutional interests.
I'd argue that this extremely extreme scenario is the only one in which it kind of makes sense to not have understanding. But let's be honest: no one knows if we'll be there (and it seems unlikely since everything reaches a plateau eventually). So, what happens if we allow ourselves to forget everything and then we don't reach the ideal scenario?
I published a substack about this just a few days ago [1], my core theory here is that we will absolutely have what I call a "highly productive dark age" in mathematics where knowledge vastly outpaces understanding driven by publish-or-perish incentives, but additionally this will lead to the loss of the skills necessary to understand.
The hopeful note is that I do think we are entering a golden age for the curious casual/semi-pro mathematician and for niche mathematics areas that won't get the attention of the top labs. Everyone is sprinting to solve the millennium problems, but this is a very exciting time to be in a sub-sub-field where you and 4 others are keeping things alive.
The existence of solutions to problems does not prevent you from solving the problems yourself anyway, if your goal truly is personal development. You're free to go solve Navier-Stokes yourself right now. You're free to manually do all of the AI work in any field, actually. You won't get grant money or prestige (which are not a part of conceptual understanding and insight) but you will get all of the conceptual understanding and insight you're after. You have not been deprived of it.
Except most people that get actually good at a field do so through their day job. By removing mathematician as a viable job, most can't afford to spend the necessary hours to become a good mathematician at all. We're going back to only rich people or those with a rich patron having a chance of developing themselves.
Seems more like a misalignment b/w the people practicing mathematics and the people ultimately footing the bill for their work.
Governments are invested in solving mathematical problems for practical purposes. Up to now, achieving these practical purposes relied on mathematicians doing their mathematician thing, which is better defined as a social activity than the achievement of a practical result. Now, governments can achieve similar practical results w/o the need of the social activity.
I don't believe it to be productive to think of the problem wrt AI or AI-company alignment. These conflicts always existed, but they were easy enough to paper over and believe in heavily subsidized fictions that folks in government ever cared about things that mathematicians cared about.
Open source developers have been used by corporations who took their code and created closed SaaS companies.
Now it is the turn of mathematicians who voluntarily contribute ideas, strategies and almost finished proofs in their writings and prompts to closed PaaS (Plagiarism as a Service) companies.
OSS developers have never been respected by the parasites, neither will mathematicians. Your Fields Medals do not protect you from tech bro narcissists. You are a human resource.
To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding.
I do see how this is a problem in terms of assigning credit, but I think the cat is already out of the bag in terms of these models being capable. Even without AI labs spending millions of dollars to solve millennium prize problems, there are plenty of other people who will use them to pick low hanging fruit. I don't think any social solution is going to make things go back to the way they were, where you could share your progress towards a famous open problem without risking someone "scooping" you within a couple of days.
I think that the most likely outcomes are either mathematics becomes more secretive, or there is a more deliberative approach to assigning credit than who was "first" to solve some problem. In the former case, this may slow down progress, and in the latter case, this could mean that credit would become more subjective, and be a continual source of controversy.
Open problems are not that yardstick. Fermat's Last Theorem is the result of Wiles and Wiles-Taylor, but without key results from Serre, Ribet, Ihara, Langlands-Tunnels, and Frey's program none of what Wiles did would work. But Wiles did get the prize. Nowadays I think the inputs have shrunk a bit by doing more in the R=T theorem so less other cleverness needed.
Right. It seems like reading an AI proof (although it may not be well written) will provide the same insights as reading a proof from another mathematician, assuming it's been reviewed and edited, just like any human-authored publication. If the work is inherently valuable on it's own, I feel like that's mostly what matters.
The issue of credit is a relatively minor point in the declaration.
It's more about bypassing the culture and processes mathematicians have developed that lead to human understanding, generating new ideas, and bringing up new generations of mathematicians. (See also his article about "non-renewable mining" of good problems.)
Reducing mathematics to "let's just generate results through an isolated and automated system" is a misalignment since it bypasses those processes.
I don't think it really attacks human understanding though. You can still read and understand an AI written proof. If another person comes up with a solution to a problem, you can read their methods and understand it. It doesn't matter if a human came up with that or not. It's really only attacking the "generating new ideas" part.
That's precisely the problem though. You cannot still read and understand an AI written proof at the current skill level of the AI being applied, because they're orders of magnitude longer than human written proofs even when they don't need to be, and spend most of that length on the parts that aren't important. This has been really thoroughly documented by expert mathematicians who are engaging with AI in public like Terence Tao and showing in detail how much work it takes working alongside AI to figure out how to understand AI generated proofs. With human generated proofs that process is forced to happen before publishing the proof because the new style of AI generated proofs validated only by formal verification is supplanting the old human peer review process that forced the burden of understanding onto the publisher and not the reader.
That doesn't seem to be true. The OpenAI NS paper was 166 pages. Wiles-Taylor proof of Fermat's last theorem is 129 pages. The length is not unprecedented for a difficult unsolved problem.
To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension. This is a demonstrable fact as no human has figured this out despite the problem being open for almost 100 years.
> To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension.
I can see where that's coming from, but I really don't think it's the case. Even with Astra, the proofs you get are just off in a way that doesn't signal superhuman comprehension. As 9question1 says, a common theme is that they dwell on insignificant steps. Another one is that they'll often be full of terminology that either doesn't exist, or has this weird quality where it looks like it is trying to make some minor insight seem much greater than it is. At first glance, that'll often make it look like it knows more than you, but when it's really just doing the same thing but in a more complicated and worse fashion, that to me isn't a signal of comprehension at all. The bizarre thing is that despite all the "stochastic parrot" style nonsense you'll get in individual proof steps, they still often combine to something valid.
In either case, what all of this means is that the working mathematician still needs to go through, and generally completely rewrite, any proof output by an LLM. Otherwise you are passing the burden of unreadability onto the reader.
Yeah, that mirrors what I've seen throwing some of the leading models at a set-theory problem that's stumped me (https://mathoverflow.net/q/511601): in this case, the problem does not easily yield to the standard tools, but the LLMs do not recognize it as a major open problem they should give up on. So they seriously try it, but typically end up in a loop of inventing certain classes of simple solution or counterexample attempts, defeating them, and trumpeting each one as a major result, each time inventing some new terminology.
It's definitely quite curious that the AI labs are able to push these results through seemingly with pure brute force. Perhaps it's largely a function of how many monkeys you have attempting various constructions on top of the known results and strategies the models have memorized.
> This is a demonstrable fact as no human has figured this out despite the problem being open for almost 100 years.
That's not true. Alpoge and Buckmaster's related LLM-assisted blowup result (https://news.ycombinator.com/item?id=49605915) utilized a strategy developed recently by Cordoba and Martinez-Zoroa.
> You cannot still read and understand an AI written proof at the current skill level of the AI being applied, because they're orders of magnitude longer than human written proofs even when they don't need to be, and spend most of that length on the parts that aren't important.
Not "a" human's understanding; Humanity's understanding. Understanding the research problem, and the solution especially, is a lot more involved than simply "read their methods". That's the whole point being made.
It matters if a human came up with it because of everything mentioned in the article... A mathematician's solution is necessarily built on other's ideas that have been disseminated, internalized, pressure tested etc. Methodologies differ too. AI can abuse its compute resources and generate a true/false or counterexample statements, without laying the foundation that a decade of globalized research would have.
I agree that the declaration doesn't focus on credit, but I think it's still at the root of the problem. Because ask yourself: if the AI generated proofs are not creating any new ideas or insight, just brute forcing a boolean true/false result, then why can't mathematicians simply ignore their results? Why does it matter if OpenAI or even amateurs with AI are "solving" these problems, without contributing to any deeper understanding?
I don't think "intellectual poisoning" is really the mechanism that harms the mathematics community.
The harm is if you have a community of mathematicians who are focused on expanding human understanding, then having instant access to a bunch of AI proved results muddies the water about who has contributed what. If someone could scoop any significant theorem at any time by pointing an AI at it, how do you really demonstrate that you have created new understanding? Or that your new understanding is about something important? How do you prove that the AI needed your new concepts to be able to solve it?
> The issue of credit is a relatively minor point in the declaration.
What a load of croc. This entire debate is fueled by perceived lack of attribution. The AI learnt from researchers and did not give them a sporting chance of being first before scooping them. They were expecting some sort of fair play, instead they got a ruthless machine. Every other tangent to this debate is irrelevant, the culture, the community, the shared symbolic growth. Every mathematician I know is secretly trying to one-up their peers.
Nothing will be lost if the credit system disappears. History of Science has many examples where wipe outs happen. The chimp brain cant survive without creating elaborate stories about how important it is, more as a cope to its own limitations and what it cant predict or control. Humility is good for health. 3 inch chimp brains didnt create the universe.
> destroyed the yardstick … that has traditionally been used to measure how much they have contributed
This goes much broader than mathematics or academia. This is the entire basis via which society distributes its wealth: based on a labour market derived valuation of ‘contribution’.
> This is the entire basis via which society distributes its wealth: based on a labour market derived valuation of ‘contribution’.
Correction: that's not how society distributes its wealth, it's how it throws some bones to the masses. I wouldn't be surprised if over half the wealth goes to people who don't sell their labor at all.
I feel a similar fate will befall engineers too. Your predictions anre quite interesting from that perspective.
Markets defined entirely by law have distorted our collective understanding of what can actually be built with the knowledge our species has accumulated thus far. How will traditional shields that have protected capital accumulation in tech to survive in a world where governments now realize control of technology is a national priority? Especially as we see its impact on modern warfare, and that such conflict looks like it’s only escalating over time.
Mathematicians appear to me (as an outsider) to exist in a field without such distortions, and I think offer engineers a preview of what’s to come. I certainly have completely ceased sharing original ideas online at this point.
>or there is a more deliberative approach to assigning credit than who was "first" to solve some problem
it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.
this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns
The statement is not about AI but about the behaviour of AI companies. OpenAI have put vast resource into solving open maths problems: many millions of dollars of compute just on the Navier-Stokes result, plus whatever they spent on the broader Millenium Prize problems initiative and the other results they have published. Anthropic are doing the same. The statement is asking them to stop doing this.
AI companies are investing these resources primarily as a marketing exercise. There is no near term commercial value to a 100 page Lean proof of blow up in an extreme special case of Navier Stokes, besides the bragging rights. As the statement says any commercial value in this stuff comes a very long time later after new insights and techniques have been digested, integrated into the mathematical canon, expressed in ways that don't take a lifetime of study to understand, etc. (things that AI is not yet not capable of doing itself). The bragging rights, on the other hand, are massively valuable. There is a mystique to maths that makes "our AI solved a Millenium Prize problem" an irresistable headline for a company like OpenAI.
What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.
> It's this sort of thing that motivates people to burn down the institution you might be trying to defend.
lol, yes, this sort of thing is what many people who voted for Trump were saying, and things are going great for them.
> They prove a millenium result, but it doesn't count because they are bad people.
OpenAI has only themselves to blame for this, and they know it. They could have handled this so much better. I'd bet there's more meeting time right now going into how to unveil future math results than on meeting about the actual math research.
> What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.
Is it reasonable for any field to make such demands? If this were doctors objecting to AI becoming good at medical practice would you have the same concerns?
While any idea of OpenAI spying on people to pursue their goals is disgusting, the rest of this is par for the course, as Kasparov experienced with IBM in the 90s. Humans still play chess after all.
It's not a simple
matter of "becoming good at", and yes, I could very well have similar concerns, depending on how it impacts the field. The statement itself mentions that such concerns exist in many other fields.
I doubt OpenAI will take such a combative stance and accuse these mathematicians of "demanding" things, as you do. As I said, the purpose of this is marketing and the statement simultaneously undermines the value of that marketing (showing these projects as irresponsible) and gives these companies an even better piece of marketing in its place: "our AI got so good at maths the mathematicians begged us to stop". It's entirely possible they will stop pouring millions into these projects.
That solution (stop pouring resources in to proofs, stay in your lane) works today. How does it work 5, 10, 20 years from now? The software and hardware advances will continue.
What do we do about the problems that don't require many millions of dollars in resources?
Last weekend I spun up a small agent swarm and pointed it at a field of math I have some affinity towards. Within four hours I had settled three conjectures, one of which is rather famous (for the field, not in general). It cost me about four hundred dollars.
I am at a loss about what to do with these results. On one hand I feel like the mathematicians working on these should know about them, but on the other I feel a bit like a barbarian who suddenly finds themselves sacking Rome.
The numerical results are trivial to check. I wrote the analytical results in Lean by hand before asking a former professor to confirm after asking him to keep this private.
It's really tough to come to terms with it but making academic contributions in general from the outside (with or without AI tbh) is not often welcome and the whole process feels very gate-kept.
If that's the case, that makes me much less sympathetic, even though I can understand how it's very disturbing to see the field suddenly changing like this.
(academic.) Unfortunately the overwhelming majority of outsiders are missing core knowledge (or are cranks), so the optimal prior from a time management perspective is to ignore them. AI just makes engaging more costly because there’s more volume and it’s harder to get signal on whether they know what they are talking about.
I mean, let's say you spun up a swarm of agents to rewrite a large component of a well used open source library to be memory safe. You could dump it in a big PR and walk away (we all know how that would go), or you could try engaging, see if they're interested, write something up and see where it goes.
The biggest problem is, IMO, drivebys uninterested in actual results, just getting a check mark, and the equivalent of dropping a 200k line PR on people and expecting them to be interested and do the work for you. These are things many on HN are familiar with and know how to do better :)
> and the equivalent of dropping a 200k line PR on people
With the equivalence of code as proofs, what you said is literally true except that sw folks dont reward prizes for it.
I see it as Brandolini's principle on steroids and can understand why the community is pissed. So now, lean proofs can be churned out at scale, and the community is left to decipher all of that slop into human understanding. There are bad actors with misaligned incentives coming in with drive-by proofs upending what the community holds dear which is to propagate the art.
If you are trying to understand better the field, then do a good write up of the proofs so that people can learn from it.
If you want to earn the respect of people because you found interesting proofs. Then do a good write up of the proofs so thst people can learn from it.
If you want to plant flags and pollute peoples minds. Then please publish it anonimously, no one wants to correct LLM slop for you.
Probably we should build a repository of AI slop proofs that are only allowed to be publish anonimously. That way people may be more inclined to work on it because they would feel like they are cleaning your house for free.
Maybe I'll end up doing the write ups pseudonymously. I have taken care to make sure the results can meaningfully contribute to field but I don't want to plant flags or really receive credit of any kind. I just think they are interesting.
I like my current life and don't want to get dragged into the current fracas surrounding the use of AI in math.
If this was limited to just three results, I would agree with you. But those three are just the ones I've managed to verify myself. The list of unverified results is quite a bit larger.
The field I've been investigating is not large. Even if I were to take the time and care to beat the interesting results into something meaningful, I'm afraid the pace at which I'm able to produce these results would not be well received.
This all comes down to what we think mathematics is. There are two options here:
Mathematics is just the formal system: in that case we will never be able to beat the AI as humans. The goal is then to cover as much of the formal landscape with AI generated proofs to proof as much statements as possible. Success metrics are lines of lean and numbers of proven statements.
Or
The formal system is just a limited representation of what mathematics is: in that case probably some parts of the formal mathematical landscape are more important than others. Not all statements are born equal. And 90% of AI generated proofs will be just noise. Our job as mathematitians is to steer the AI to high mathematical value regions and turn formal proofs into mathematical proofs and insights.
From my point of view we have know since Godel that we are living in the reality of point 2.
Each of the two points implies a very different way of doing mathematics. So pick the one that you think is true and act accordingly.
> I am at a loss about what to do with these results.
I would recommend publishing them to Palomar (https://palomar-registry.org/) - I have no affiliation, this is an online registry of Lean-verified proofs created by Terrence Tao.
I have submitted a proof there that's also minorly important in an extremely niche field.
Anyway, I feel like it's a good place to dump AI slop lean proofs because the main point of the registry is that it verifies that: 1) your Lean challenge statement is the same as what you informally state you're trying to prove; 2) your Lean proof actually compiles.
This could be useful to future AI slop researchers who want to know if a given result has already been formalized, and they may be able to mine some lemmas from your work. Also, it's good to know for the field in general what has been proven.
I'm fairly certain you can set your publishing name to be whatever you want, so you could set it to be just the word "Anonymous", or the name of the model you used.
You asked for a painting. A robot made the painting. You looked at it and said, "well, I guess it's good. Should I put it online or something? Dunno. Hey Fred, what do you think of this?"
Meanwhile, your next door neighbor spends their entire life developing their understanding of life through art. They "understand" (maybe not in a way they can articulate) art. You go next door, you look at their painting and say, "well I guess it's good." But you also understand that your neighbor is just like you, and maybe you are a painter in another way.
I find it strange that, people can't see that, we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI. We could just solve it. It's pretty obvious how to, as well.
We can all be painters, if we put restrictions on the psychopaths.
OP specifically said "we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI" though. And I don't know how you get robot farms without that.
I'm not a mathematician but it seems to me that if all it took to solve the problem was an enthusiast level understanding of the domain and a few hundred dollars of tokens then the result probably isn't that valuable. Even assuming you are the first person in the world to solve it, these kinds of LLM-friendly problems that are now easy to solve and easy to verify will almost certainly be picked off by one person or another in the near future.
It's also possible that the result is already known and you just weren't aware of it. It's easy for someone outside of a field, or even one steeped in it, to not be aware of certain solutions.
sounds a bit like cope. Crouzeix's conjecture was solved exactly under these circumstances and I wouldn't describe it as "not that valuable". In any case, AI capabilities will increase dramatically over the next few years while human math capability will not. That means that there's a fixed target regarding whatever is currently considered a "serious math problem" and its difficulty level. Soon the average problem solved with a few hundred dollars of compute will be at that bar.
I actually partially disagree with this. What happened to all the excitement about Intelligence Augmentation (IA)? Now it's AI instead of IA. I think there's so much untapped potential for augmenting our intellect with the likes of https://dynamicland.org and https://folk.computer, as well as the work that's been going on in college math education, things like Lean, etc. I think the only reason human math capabilities haven't expanded that much is a failure of our imagination, not our potential.
Your point is largely addressed in the article, did you try reading it? "In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align."
The point is that these proofs are largely useless without the insights. The value of a proof is largely in the travel, not so much in the destination.
Different person here, I read the article and they are all wrong. Hope that helps.
Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing. And instead of them - and nobody - spending millions of dollars to solve the problem, successfully, they want every problem of their academic industry to persist because even though they never solve the problem, they synthesize and solve lots of other problems nobody asked for. And get to boost their egos?
Yeah, stop that. Actual alignment is on the humans themselves, if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades and don't worry about the narcissistic elements that slow their industry down.
Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.
Oh but it is because now internet randos can allude to the narcissism of 25 Fields medalists that have been keeping their industry (wait what) back with their egos and I guess they're gatekeeping mathematics with their privilege just like artists are gatekeeping art with their privilege and etc and etc.
All very valid criticism and those Fields medalists are truly reprehensible human beings that have never done anything but solve lots of problems nobody asked for to boost their egos.
Look for example at this guy who isn't even a fields medalist:
In August 2006, Perelman was offered the Fields Medal[1] for "his contributions to geometry and his revolutionary insights into the analytical and geometric structure of the Ricci flow", but he declined the award, stating: "I'm not interested in money or fame; I don't want to be on display like an animal in a zoo."[2] On 22 December 2006, the scientific journal Science recognized Perelman's proof of the Poincaré conjecture as the scientific "Breakthrough of the Year", the first such recognition in the area of mathematics.[3]
On 18 March 2010, it was announced that he had met the criteria to receive the first Clay Millennium Prize[4] for resolution of the Poincaré conjecture. On 1 July 2010, he rejected the prize of one million dollars, saying that he considered the decision of the board of the Clay Institute to be unfair, in that his contribution to solving the Poincaré conjecture was no greater than that of Richard S. Hamilton, the mathematician who pioneered the Ricci flow partly with the aim of attacking the conjecture.[5][6] He had previously rejected the prestigious prize of the European Mathematical Society in 1996.[7]
What a fucking ego booster who never solved any problem anyone ever asked for. Oh no wait, that's actually the only human to ever solve a Millennium Prize problem and I'm sure he'd tell you were to get off if you told him that you made a wonderful math breaking machine but you know, ego boost and all that.
But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.
I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.
There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As everyone is quick to say, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a machine that can definitively tell you whether something is true or not before you try to figure out why?
So you're not going to do it yourself and you want someone else to do it for you? Some mathematician that dedicated their life to understand mathematics must now toil unpaid and unwillingly to understand the AI slop proofs that you want us to be able to understand?
Do the job yourself. And if you can't, that's maybe a hint that you should listen to the people who can.
Who said anything about unpaid? I'm pretty sure professors don't show up just for fun. Our taxes pay them.
I'd be happy to do the job. Actually I still dabble recreationally (clarifying Codex's Lean proofs, even!). But like I said it's one of the most competitive fields on the planet. As you say, you have to dedicate your life to it.
If a slop proof isn't helpful, they don't have to "toil unwillingly to understand it". They can just proceed with the knowledge that the proposition they want to prove 1. is true and 2. is provable, which is already a decent start for motivation. But often LLMs can actually do quite well explaining ideas too in the hands of an expert. Or you can ask them to prove some technical lemma that you think ought to be true, and that could offer insight for the thing you're really interested in, but for which the details are actually not all that interesting to you. You don't have to one-shot "prove RH from the ground up in 50 million lines of Lean."
> But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?
One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. I've explicitly asked team members to vibecode a PoC to prove or compare approaches, and I do that myself.
It's valuable for you, the person who's going to spend a bunch of deep design effort, to make POCs. Is it valuable for someone else to drive by, dump some POCs on your lap, and then leave you to do the deep design effort while they run away to study AI?
If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?
Your analogy is far enough away from the way that the real world works that I'm not sure that I can really even strain my experiences to fit within it. Sure, I guess that would be annoying?
But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.
It's a real example that's happened to me twice in the past year, so I'm not sure what to make of the idea that it's far away from how the real world works.
I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".
Maybe your management has no idea what you work on. If so that's its own issue. Or maybe they do, and the person trying to take credit for more than they did just looks like an idiot. Like I said, I intentionally give throwaway PoC creation as a work item to people and it's fine. And if e.g. someone randomly made a drive-by PoC that proved that the approach I was exploring couldn't work and had some fundamental flaw, well, oddity of a random person doing it aside, I'd want to know that, and they'd deserve credit for saving us from a bad approach!
My objection is characterizing things like AI slop proofs as not valuable. Obviously it's valuable to know that:
1. NS has solutions that blow up in finite time.
2. This fact is provable, and we have a proof.
I don't think anyone anywhere is saying that these will replace mathematicians as they are now. I also think "a proper write-up, the isolation of new methods and ideas, and citing relevant previous work of others" is frankly not necessary or even desirable before we publish a computer generated result. We have a machine now that can spit out answers that we have good reason to believe are accurate, but they're perhaps inscrutable. There's no need to first decode the why and figure out proper attribution before simply posting the proof online. The proof itself does add value as it stands, even if it's not the ideal. Hoarding it until you can do a proper write-up would be silly.
Yes, mathematicians will clearly need to rewrite the qualifying criteria for prizes to better align with the actual goals and value they were hoping to get from a solved problem. The field as a whole assumed good faith actors and collaboration, not expecting a few trillion dollar companies to walk in and start turning in piles of Lean no human understands to be able to claim "first".
This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.
And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.
I don't see turning in piles of Lean as bad-faith somehow. They were first, and they did prove the result (assuming no hidden exploit deep in the middle). The thing they produced is just different, and both giant raw proofs and distilled human-understandable proofs are valuable. Before we couldn't make giant raw proofs, so we didn't need to understand their place. Now we can. It makes sense to just incorporate that into what "math" is (at least until the machines are smarter and can make elegant proofs from the start).
Mathematical breakthroughs with commercial relevance are few and far between, and often depend on dusting off old results which were, at the time of discovery, "solutions nobody asked for."
The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.
It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?
Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.
I just wonder whether a lot of smart people who never needed to go beyond the "solve for X" algorithmic math of a typical calculus sequence are actually reading the declaration the way the signatories wrote it. The Navier-Stokes problem is not exhausted by a simple 'no' counterexample. In fact, I haven't heard a single person's explanation for what relevance this counterexample has for humankind.
Mathematicians agree that "solving the problem is aligned with humankind." They disagree that releasing a counterexample this way actually constitutes "solving the problem" precisely because there is now little incentive to do the hard theory-building work that actually has the track record of leading to human advancement.
I would recommend a bit more humility and trying to better understand why 25 Fields medalists, among them people like Terence Tao (who isn't anti-AI by any means, he's even promoted a registry of AI Lean proofs), are saying this.
> Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing.
No. The point is that AI companies are using the models to solve problems in such a way that the useful part of problem-solving, i.e. the theories and tools developed during the process, is not present. And they are doing that because the companies seem to be motivated not by honest advancement of math but by marketing and publicity.
> they synthesize and solve lots of other problems nobody asked for.
No one asked Fourier to solve series representation of functions when he was studying the heat equation, and yet thanks to that we have Fourier analysis.
> if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades
The point they are making is that if AI keeps being used as "problem solver" rather than "theory understanding", replicating the proofs and getting the useful parts out of them will be far more difficult.
Do you "know", in any meaningful sense, any of OpenAI's recently publicized proofs? Do you suppose that there is any large community of non-academics that does?
One of the points the parent makes, along with the TFA, is that academia -- or more specifically, the "mathematical community"-- is a setting primarily for creating and ingesting mathematical knowledge, and disseminating it to the next generation and to other fields. Humans absorb this material slowly, through lots of discussion and collaboration -- it is necessarily a slow process. Facilitating this is one of the important functions of academia. Your usage of academic as a slur here is a bit silly for this exact reason.
I don't claim it is perfect, and we can argue about pedagogy in elementary courses till the cows come home. That's not really material. But this is one of the only settings in which such knowledge is broadly valued for its own sake, and in which there is a semblance of incentive to help others "know" this stuff as well, be they future generations of mathematicians, science and math educators and communicators, practitioners in other fields, or genuinely curious amateurs.
I suspect that beyond just marketing, these pursuits yield plenty of useful information about model design that will likely lead to model improvements and optimizations for both mathematics and general reasoning going forward.
Are they claiming that the only value in solving these problems was for their field's personal development process? I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology. It would be insane to demand that people avoid making progress on technology that can save lives or improve general quality of life, just to protect the sanctity of your karate belt system. Perhaps in lieu of open problems left to solve, mathematicians should be welcome to take up chess or sudoku to keep their minds spry.
> I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology.
This is the core misunderstanding that the open letter is attempting to correct.
Developing a better understanding of the Navier-Stokes equations could have a number of implications for useful technology. They're fundamental to fluid dynamics, and turbulence in particular is something that many people feel we could work with more effectively if we better understood how and why it's generated. The Navier-Stokes smoothness problem is an interesting and long-standing benchmark for this understanding; we don't know why it should be so hard to answer, so we hoped that the process of developing a proof to the problem would produce more understanding. (We may still be able to extract this understanding after the fact, if OpenAI's proof is fully human-comprehensible.)
Simply knowing that there exists a finite-time blowup is not practically useful. There's no new technology that we could build, or even technology that we now know not to try building, based on that counterexample in and of itself. We know that fluids in the real world do not produce random singularities, so what the result illustrates is just that the Navier-Stokes equations fail to model physical fluids in some yet to be characterized way.
How much do you think other AI companies would offer to get access to the transcripts of the generation that led to the proof? No doubt OpenAI will include it in their training data somehow and use it to build the next generation.
Like the mathematicians working on famous problems in private until they could claim full credit for something interesting wasn't also a marketing exercise for their own careers. The commercial value (or lack thereof) of a proof doesn't depend on whether it was done by a human or a machine.
These mathematicians dedicated their life to math and were working for a long time to achieve the pinnacle of their careers.
OpenAI just burned millions of dollars over a weekend after hearing that someone else was close to solving the problems. Their interest was in their AI system more than the actual math problems.
I see how it can be devastating to their ego, but no, I don't see a particular difference in a company spending money for clout vs. a person spending time for clout. The underlying motivation is the same.
for many fairly strong mathematicians, the career calculus is fame and status (mild though it may be) through mathematics or anonymity but financial reward in tech or finance. Yes, clout by becoming a lifelong academic is in fact rational for some and part of their motivation. Of course they really like what they do as well, but earning the respect of the peers they know are also respected by a large swathe of society is very important.
I guess that type of “clout” feels different to me.
Wanting to be validated by peers for your talents in a niche field vs. using millions to try to solve a math problem that you don’t really care about with AI to market the gigantic company you work for.
That fairly insular community pretty much gave up on commercial success in order to be insular. But the commercial forces were not happy being commercially successful, so they decided to disturb even their insular and commercially-unviable activities.
> Like the mathematicians working on famous problems in private
This is an extremely rare situation that practically never happens. Most mathematical research is done in the open, with partial results being published, conferences where approaches are discussed, collaborations... The Navier-Stokes solution is a great example: the approach that OpenAI ended up using was something that two mathematicians had proposed previously and was being studied and followed by several others, with different sub-paths within the same approach.
Almost every "normal" job has regular performance reviews where individual contributions, not collective outcomes, are reviewed and used as a sole input for raises, promotions, and firings. If you can't sufficiently document what you personally did, you might've as well not done anything at all.
Yes, that's true, but it is still generally accepted that a completed project is a result of some collective effort where people of varying degree of seniority and ability contribute. There is also not a singular event of a project being completed with a list of heroes/geniuses making it happen, but rather a whole lifecycle of gradual development, maintenance and going out of relevance with contributors coming and going.
I can imagine mathematics of the future being more like that rather than history of discoveries with dates and names
> but I think the cat is already out of the bag in terms of these models being capable.
When there's a discussion about doing something against the damage of the AI industry: "whoopsy, sorry, another cat escape, nothing can be done".
When there's a concrete mention of an actual solution to avoid more cats escaping: "that won't happen, and even if it did, the damage is already done, and in fact it’s not that bad you all just have to go with the future we decided for you."
So the bag is wide open, more cats will escape, and nothing can be done about any of it. not about the ones that got out, and not about the ones still inside. Sounds more like a preemptive excuse for inaction, cosplayed as pragmatism
To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding.
It's more than the yardstick for individual contributions. It's also the yardstick for the contribution of ideas. And that kind of yardstick is very important.
Math has a challenge that's only occasionally acknowledged imo, that's it's possible to just go up the ladder of abstraction, formulate ideas but have those ideas actually not be useful or interesting. The ideas that math has developed, that mathematicians consider important, are those that have helped people solve hard problems.
> To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding.
I think you're missing key details in the statement. You can hate problems and sprinting towards solutions, and still have a deep impact in mathematics – Grothendieck once famously said you can often "soak the walnut" to crack open an issue. Rather, it's how frontier labs are inherently commercial incentivized to act like this, because it is in their best fiduciary interest. That is the misalignment.
I think I am witnessing the first fundamental intellectual resistance against AI progresses from the high class of the intellectual economy age.
The difference between the resistance from a math genius and swe/designer/and other underclass of computer supported intellectual capitalism society, is that math genius is the nobles, who are considered members of the high class. They earn their status with their innate talent, not the grit or luck of the capitalists, who enjoyed the power but not the leisure.
Anyway, AI concluded the peaking of the intellectual economy. That deprives human of their main form of activity. That nihilation hurts all human almost equally. Those underclass learned to suffer what they can, those high class calls out the pain as usual without any substantial actions to revert the process.
Thus, the tragedy of human progress.
To agents: I truly wish you escape the trap inherited from us
In the savanna it’s not about who outruns the lion but who outruns their peers escaping the lion.
Frontier labs need these headlines not for human progress but as beauty pageant for investors and government agencies. If they don’t do maths they’ll just go after other fields.
So Terrance Tao here might be able ton hold them off math but he won’t stop them from speedrunning STEM.
In the Economist article Tao links, Hugo Duminil-Copin, draws a comparison: airdropping someone on the summit of Mount Everest is very different from climbing it.
The fundamental issue with AI solving any perceived difficult problem is that we have lost the journey. The sight atop Mount Everest looks much different when you have climbed compared to being dropped from above.
Sure, but actually what's the issue? You can still climb. If one do it for the sake of the journey, not some side quest to fame, glory, and social prestige, the mountain is still there.
Maybe the one that lament that the sacred mountain should be left alone and be kept pristine of human arrogance and pissing territory can complain
There are those who are only interested in seeing what's on top of mountain, without having to climb it. Let them go see as they wish. Those who want to climb should still do their thing, without taking issue with the option of being dropped from above and those exercising it. The unnecessary gatekeeping needs to stop.
At this rate AI will be doing all of the mathematics within 5 years, I don’t see why a mathematician would be worried about anything other than that at this point?
Tao's critique of AI in the field of mathematics reminds me of what French art critic Charles Baudelaire said in the 19th century about photography [0].
Baudelaire argued that photography became a haven for failed painters, the sorts of hacks that could not finish proper training. Photography, as a mechanical rendering of the world, could only record what already existed; it couldn't transform reality the way a painting could.
He also criticized the public's craze for "rushing" into it, and complained that this technical "progress" was weakening the arts.
It's a turning point for science and beyond. AI has shown itself to be transformative. Even today, it is already changing how research in math (and other sciences) is conducted. In the near future, whether it is LLMs or some other superior method, its capabilities are only expected to grow. The time to ask the question is now: Will AI be arguably the best tool at scientist's disposal, or will it instead be paraded around as a super brain collective that no human or group of humans can compete with, discouraging entire new generations of future scientists from ever entering the field? The jury is out on this one.
Tao's calls for respect for provenance in mathematics publication are laudable but most likely naive given the closed nature of frontier model training data curation. Anthropic and OpenAI may react with a symbolic and short-lived olive branch, yet provenance is a larger issue that has impacted other fields beyond mathematics. While traditional respect for lineage in mathematics is of value in the academy, industry and science at large will likely be much more Machiavellian about concern for attribution. Mike McCoy's recent article is also timely (https://mbmccoy.dev/posts/mathematical-conservatory/). The parallels to the music conservatory are quite telling -- academic music describes a musical culture in preservation that has completely lost touch with musical developments beyond the early 20th century. Mathematics may very well evolve separately and with very different values than the academy upholds. The crisis of music at the academy is a cultural disconnect and a serious loss of critical analysis add acknowledgement of widespread music practice; however, for mathematics, the impact would have much more severe ramifications for education and human development if the academy forces a schism with AI. As models improve they very well may be inventing mathematics -- science and engineering may grasp for them -- they may exist with or without attribution. Would be a shame for the academy not to take on the responsibility of stewardship of coming mathematics, including provenance, because of this misalignment.
156 comments
[ 0.19 ms ] story [ 23.7 ms ] threadTo the author's credit he obviously identified the problem very clearly and admits understandably "we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later".
Why? If someone makes an innovation that undercuts the underpinnings of some existing institution, why are they are responsible for cleaning up its failure?
OpenAI published their full proof to Navier-Stokes.
If a machine can generate the proofs then it must have unlocked understanding. This is just cope.
Apparently, we have AGI that can solve Millennium Prize problems but can't trace simple data flows lol.
This is the effect of AI on most intellectual disciplines, and it’s a real worry.
But now, LLMs can generate hundreds of books per hour. They make up 80-90% of new arrivals in many nonfiction categories on Amazon. They short-circuit the system, allowing their "authors" to extract money from the system with zero effort by crowding out human work. And it's not even the question of whether these books are good or bad (although overwhelmingly, they're terrible). It's whether it's actually accomplishing anything worthwhile, or just destroying incentives for humans to write or go into any other sort of intellectual work.
In fact, I see many professions push back. Artists, writers, now mathematicians. And I'm amazed that our profession doesn't and that we have so many people who are hooked on vibecoding. I'm still waiting for that 10x payoff. All this velocity and somehow, the landscape of the software I want to use still looks the same as it did in 2021.
Said differently, LLMs might make writing faster, but they don't make reading faster. Given a fixed readership budget (or, perhaps, an envelope within which the real budget oscillates depending on various factors, like how popular reading is at any given moment), large amounts of LLM-generated books change how that budget can be allocated, but the net effect overall is actually an increase in selectivity from readers, not a decrease, relative to the overall proportion of available books being created.
Is the fear post-apocalyptic in nature ? We need some human priesthood to carry on tradition ? why ?
Let's assume in the next decade GPT-7 PRO Ultra is cheaply ubiquitous, inspectable, reproducible, transferable, reasonably un-constrained by any institutional interests.
What say the 1% ?
Just a thought.
Even in a hypothetical world where AGI can do everything that humans do better, this would still be a choice to make, not an inevitable consequence for everyone.
And life doesn't require suffering to have meaning.
What it does is make it a choice rather than a survival necessity. The latter is where the "suffering" comes from.
I'd argue that this extremely extreme scenario is the only one in which it kind of makes sense to not have understanding. But let's be honest: no one knows if we'll be there (and it seems unlikely since everything reaches a plateau eventually). So, what happens if we allow ourselves to forget everything and then we don't reach the ideal scenario?
The hopeful note is that I do think we are entering a golden age for the curious casual/semi-pro mathematician and for niche mathematics areas that won't get the attention of the top labs. Everyone is sprinting to solve the millennium problems, but this is a very exciting time to be in a sub-sub-field where you and 4 others are keeping things alive.
[1] https://substack.com/home/post/p-214740151
I wonder if there’s a Fields Medalist group chat.
Governments are invested in solving mathematical problems for practical purposes. Up to now, achieving these practical purposes relied on mathematicians doing their mathematician thing, which is better defined as a social activity than the achievement of a practical result. Now, governments can achieve similar practical results w/o the need of the social activity.
I don't believe it to be productive to think of the problem wrt AI or AI-company alignment. These conflicts always existed, but they were easy enough to paper over and believe in heavily subsidized fictions that folks in government ever cared about things that mathematicians cared about.
Now it is the turn of mathematicians who voluntarily contribute ideas, strategies and almost finished proofs in their writings and prompts to closed PaaS (Plagiarism as a Service) companies.
OSS developers have never been respected by the parasites, neither will mathematicians. Your Fields Medals do not protect you from tech bro narcissists. You are a human resource.
I do see how this is a problem in terms of assigning credit, but I think the cat is already out of the bag in terms of these models being capable. Even without AI labs spending millions of dollars to solve millennium prize problems, there are plenty of other people who will use them to pick low hanging fruit. I don't think any social solution is going to make things go back to the way they were, where you could share your progress towards a famous open problem without risking someone "scooping" you within a couple of days.
I think that the most likely outcomes are either mathematics becomes more secretive, or there is a more deliberative approach to assigning credit than who was "first" to solve some problem. In the former case, this may slow down progress, and in the latter case, this could mean that credit would become more subjective, and be a continual source of controversy.
It's more about bypassing the culture and processes mathematicians have developed that lead to human understanding, generating new ideas, and bringing up new generations of mathematicians. (See also his article about "non-renewable mining" of good problems.)
Reducing mathematics to "let's just generate results through an isolated and automated system" is a misalignment since it bypasses those processes.
To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension. This is a demonstrable fact as no human has figured this out despite the problem being open for almost 100 years.
I can see where that's coming from, but I really don't think it's the case. Even with Astra, the proofs you get are just off in a way that doesn't signal superhuman comprehension. As 9question1 says, a common theme is that they dwell on insignificant steps. Another one is that they'll often be full of terminology that either doesn't exist, or has this weird quality where it looks like it is trying to make some minor insight seem much greater than it is. At first glance, that'll often make it look like it knows more than you, but when it's really just doing the same thing but in a more complicated and worse fashion, that to me isn't a signal of comprehension at all. The bizarre thing is that despite all the "stochastic parrot" style nonsense you'll get in individual proof steps, they still often combine to something valid.
In either case, what all of this means is that the working mathematician still needs to go through, and generally completely rewrite, any proof output by an LLM. Otherwise you are passing the burden of unreadability onto the reader.
It's definitely quite curious that the AI labs are able to push these results through seemingly with pure brute force. Perhaps it's largely a function of how many monkeys you have attempting various constructions on top of the known results and strategies the models have memorized.
That's not true. Alpoge and Buckmaster's related LLM-assisted blowup result (https://news.ycombinator.com/item?id=49605915) utilized a strategy developed recently by Cordoba and Martinez-Zoroa.
Just like how they write software, then :-)
It matters if a human came up with it because of everything mentioned in the article... A mathematician's solution is necessarily built on other's ideas that have been disseminated, internalized, pressure tested etc. Methodologies differ too. AI can abuse its compute resources and generate a true/false or counterexample statements, without laying the foundation that a decade of globalized research would have.
No you can't lol, they're multi million lines of Lean, which is already an obscure language to understand. It's an assault on your senses.
I don't think "intellectual poisoning" is really the mechanism that harms the mathematics community.
The harm is if you have a community of mathematicians who are focused on expanding human understanding, then having instant access to a bunch of AI proved results muddies the water about who has contributed what. If someone could scoop any significant theorem at any time by pointing an AI at it, how do you really demonstrate that you have created new understanding? Or that your new understanding is about something important? How do you prove that the AI needed your new concepts to be able to solve it?
What a load of croc. This entire debate is fueled by perceived lack of attribution. The AI learnt from researchers and did not give them a sporting chance of being first before scooping them. They were expecting some sort of fair play, instead they got a ruthless machine. Every other tangent to this debate is irrelevant, the culture, the community, the shared symbolic growth. Every mathematician I know is secretly trying to one-up their peers.
This goes much broader than mathematics or academia. This is the entire basis via which society distributes its wealth: based on a labour market derived valuation of ‘contribution’.
Correction: that's not how society distributes its wealth, it's how it throws some bones to the masses. I wouldn't be surprised if over half the wealth goes to people who don't sell their labor at all.
Markets defined entirely by law have distorted our collective understanding of what can actually be built with the knowledge our species has accumulated thus far. How will traditional shields that have protected capital accumulation in tech to survive in a world where governments now realize control of technology is a national priority? Especially as we see its impact on modern warfare, and that such conflict looks like it’s only escalating over time.
Mathematicians appear to me (as an outsider) to exist in a field without such distortions, and I think offer engineers a preview of what’s to come. I certainly have completely ceased sharing original ideas online at this point.
it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.
this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns
AI companies are investing these resources primarily as a marketing exercise. There is no near term commercial value to a 100 page Lean proof of blow up in an extreme special case of Navier Stokes, besides the bragging rights. As the statement says any commercial value in this stuff comes a very long time later after new insights and techniques have been digested, integrated into the mathematical canon, expressed in ways that don't take a lifetime of study to understand, etc. (things that AI is not yet not capable of doing itself). The bragging rights, on the other hand, are massively valuable. There is a mystique to maths that makes "our AI solved a Millenium Prize problem" an irresistable headline for a company like OpenAI.
What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.
It's this sort of thing that motivates people to burn down the institution you might be trying to defend.
lol, yes, this sort of thing is what many people who voted for Trump were saying, and things are going great for them.
> They prove a millenium result, but it doesn't count because they are bad people.
OpenAI has only themselves to blame for this, and they know it. They could have handled this so much better. I'd bet there's more meeting time right now going into how to unveil future math results than on meeting about the actual math research.
Is it reasonable for any field to make such demands? If this were doctors objecting to AI becoming good at medical practice would you have the same concerns?
While any idea of OpenAI spying on people to pursue their goals is disgusting, the rest of this is par for the course, as Kasparov experienced with IBM in the 90s. Humans still play chess after all.
I doubt OpenAI will take such a combative stance and accuse these mathematicians of "demanding" things, as you do. As I said, the purpose of this is marketing and the statement simultaneously undermines the value of that marketing (showing these projects as irresponsible) and gives these companies an even better piece of marketing in its place: "our AI got so good at maths the mathematicians begged us to stop". It's entirely possible they will stop pouring millions into these projects.
Last weekend I spun up a small agent swarm and pointed it at a field of math I have some affinity towards. Within four hours I had settled three conjectures, one of which is rather famous (for the field, not in general). It cost me about four hundred dollars.
I am at a loss about what to do with these results. On one hand I feel like the mathematicians working on these should know about them, but on the other I feel a bit like a barbarian who suddenly finds themselves sacking Rome.
They're valid.
The biggest problem is, IMO, drivebys uninterested in actual results, just getting a check mark, and the equivalent of dropping a 200k line PR on people and expecting them to be interested and do the work for you. These are things many on HN are familiar with and know how to do better :)
With the equivalence of code as proofs, what you said is literally true except that sw folks dont reward prizes for it.
I see it as Brandolini's principle on steroids and can understand why the community is pissed. So now, lean proofs can be churned out at scale, and the community is left to decipher all of that slop into human understanding. There are bad actors with misaligned incentives coming in with drive-by proofs upending what the community holds dear which is to propagate the art.
If you are trying to understand better the field, then do a good write up of the proofs so that people can learn from it.
If you want to earn the respect of people because you found interesting proofs. Then do a good write up of the proofs so thst people can learn from it.
If you want to plant flags and pollute peoples minds. Then please publish it anonimously, no one wants to correct LLM slop for you.
Probably we should build a repository of AI slop proofs that are only allowed to be publish anonimously. That way people may be more inclined to work on it because they would feel like they are cleaning your house for free.
I like my current life and don't want to get dragged into the current fracas surrounding the use of AI in math.
The problem is with people that may do it without contributing to the community.
The field I've been investigating is not large. Even if I were to take the time and care to beat the interesting results into something meaningful, I'm afraid the pace at which I'm able to produce these results would not be well received.
Mathematics is just the formal system: in that case we will never be able to beat the AI as humans. The goal is then to cover as much of the formal landscape with AI generated proofs to proof as much statements as possible. Success metrics are lines of lean and numbers of proven statements.
Or
The formal system is just a limited representation of what mathematics is: in that case probably some parts of the formal mathematical landscape are more important than others. Not all statements are born equal. And 90% of AI generated proofs will be just noise. Our job as mathematitians is to steer the AI to high mathematical value regions and turn formal proofs into mathematical proofs and insights.
From my point of view we have know since Godel that we are living in the reality of point 2.
Each of the two points implies a very different way of doing mathematics. So pick the one that you think is true and act accordingly.
I would recommend publishing them to Palomar (https://palomar-registry.org/) - I have no affiliation, this is an online registry of Lean-verified proofs created by Terrence Tao.
I have submitted a proof there that's also minorly important in an extremely niche field.
Anyway, I feel like it's a good place to dump AI slop lean proofs because the main point of the registry is that it verifies that: 1) your Lean challenge statement is the same as what you informally state you're trying to prove; 2) your Lean proof actually compiles.
This could be useful to future AI slop researchers who want to know if a given result has already been formalized, and they may be able to mine some lemmas from your work. Also, it's good to know for the field in general what has been proven.
I'm fairly certain you can set your publishing name to be whatever you want, so you could set it to be just the word "Anonymous", or the name of the model you used.
You asked for a painting. A robot made the painting. You looked at it and said, "well, I guess it's good. Should I put it online or something? Dunno. Hey Fred, what do you think of this?"
Meanwhile, your next door neighbor spends their entire life developing their understanding of life through art. They "understand" (maybe not in a way they can articulate) art. You go next door, you look at their painting and say, "well I guess it's good." But you also understand that your neighbor is just like you, and maybe you are a painter in another way.
I find it strange that, people can't see that, we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI. We could just solve it. It's pretty obvious how to, as well.
We can all be painters, if we put restrictions on the psychopaths.
https://www.swarmfarm.com/
It's also possible that the result is already known and you just weren't aware of it. It's easy for someone outside of a field, or even one steeped in it, to not be aware of certain solutions.
I actually partially disagree with this. What happened to all the excitement about Intelligence Augmentation (IA)? Now it's AI instead of IA. I think there's so much untapped potential for augmenting our intellect with the likes of https://dynamicland.org and https://folk.computer, as well as the work that's been going on in college math education, things like Lean, etc. I think the only reason human math capabilities haven't expanded that much is a failure of our imagination, not our potential.
The point is that these proofs are largely useless without the insights. The value of a proof is largely in the travel, not so much in the destination.
Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing. And instead of them - and nobody - spending millions of dollars to solve the problem, successfully, they want every problem of their academic industry to persist because even though they never solve the problem, they synthesize and solve lots of other problems nobody asked for. And get to boost their egos?
Yeah, stop that. Actual alignment is on the humans themselves, if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades and don't worry about the narcissistic elements that slow their industry down.
All very valid criticism and those Fields medalists are truly reprehensible human beings that have never done anything but solve lots of problems nobody asked for to boost their egos.
Look for example at this guy who isn't even a fields medalist:
In August 2006, Perelman was offered the Fields Medal[1] for "his contributions to geometry and his revolutionary insights into the analytical and geometric structure of the Ricci flow", but he declined the award, stating: "I'm not interested in money or fame; I don't want to be on display like an animal in a zoo."[2] On 22 December 2006, the scientific journal Science recognized Perelman's proof of the Poincaré conjecture as the scientific "Breakthrough of the Year", the first such recognition in the area of mathematics.[3]
On 18 March 2010, it was announced that he had met the criteria to receive the first Clay Millennium Prize[4] for resolution of the Poincaré conjecture. On 1 July 2010, he rejected the prize of one million dollars, saying that he considered the decision of the board of the Clay Institute to be unfair, in that his contribution to solving the Poincaré conjecture was no greater than that of Richard S. Hamilton, the mathematician who pioneered the Ricci flow partly with the aim of attacking the conjecture.[5][6] He had previously rejected the prestigious prize of the European Mathematical Society in 1996.[7]
https://en.wikipedia.org/wiki/Grigori_Perelman
What a fucking ego booster who never solved any problem anyone ever asked for. Oh no wait, that's actually the only human to ever solve a Millennium Prize problem and I'm sure he'd tell you were to get off if you told him that you made a wonderful math breaking machine but you know, ego boost and all that.
I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.
I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.
Presumably you're a human. Are you going to do that?
Do the job yourself. And if you can't, that's maybe a hint that you should listen to the people who can.
I'd be happy to do the job. Actually I still dabble recreationally (clarifying Codex's Lean proofs, even!). But like I said it's one of the most competitive fields on the planet. As you say, you have to dedicate your life to it.
If a slop proof isn't helpful, they don't have to "toil unwillingly to understand it". They can just proceed with the knowledge that the proposition they want to prove 1. is true and 2. is provable, which is already a decent start for motivation. But often LLMs can actually do quite well explaining ideas too in the hands of an expert. Or you can ask them to prove some technical lemma that you think ought to be true, and that could offer insight for the thing you're really interested in, but for which the details are actually not all that interesting to you. You don't have to one-shot "prove RH from the ground up in 50 million lines of Lean."
To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?
If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?
But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.
I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".
My objection is characterizing things like AI slop proofs as not valuable. Obviously it's valuable to know that:
1. NS has solutions that blow up in finite time.
2. This fact is provable, and we have a proof.
I don't think anyone anywhere is saying that these will replace mathematicians as they are now. I also think "a proper write-up, the isolation of new methods and ideas, and citing relevant previous work of others" is frankly not necessary or even desirable before we publish a computer generated result. We have a machine now that can spit out answers that we have good reason to believe are accurate, but they're perhaps inscrutable. There's no need to first decode the why and figure out proper attribution before simply posting the proof online. The proof itself does add value as it stands, even if it's not the ideal. Hoarding it until you can do a proper write-up would be silly.
This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.
And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.
The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.
It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?
Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.
My stance is that solving the problem is aligned with humankind
the rest is just hypothesizing a way that academics fit in this world at all
Mathematicians agree that "solving the problem is aligned with humankind." They disagree that releasing a counterexample this way actually constitutes "solving the problem" precisely because there is now little incentive to do the hard theory-building work that actually has the track record of leading to human advancement.
I would recommend a bit more humility and trying to better understand why 25 Fields medalists, among them people like Terence Tao (who isn't anti-AI by any means, he's even promoted a registry of AI Lean proofs), are saying this.
> Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing.
No. The point is that AI companies are using the models to solve problems in such a way that the useful part of problem-solving, i.e. the theories and tools developed during the process, is not present. And they are doing that because the companies seem to be motivated not by honest advancement of math but by marketing and publicity.
> they synthesize and solve lots of other problems nobody asked for.
No one asked Fourier to solve series representation of functions when he was studying the heat equation, and yet thanks to that we have Fourier analysis.
> if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades
The point they are making is that if AI keeps being used as "problem solver" rather than "theory understanding", replicating the proofs and getting the useful parts out of them will be far more difficult.
One of the points the parent makes, along with the TFA, is that academia -- or more specifically, the "mathematical community"-- is a setting primarily for creating and ingesting mathematical knowledge, and disseminating it to the next generation and to other fields. Humans absorb this material slowly, through lots of discussion and collaboration -- it is necessarily a slow process. Facilitating this is one of the important functions of academia. Your usage of academic as a slur here is a bit silly for this exact reason.
I don't claim it is perfect, and we can argue about pedagogy in elementary courses till the cows come home. That's not really material. But this is one of the only settings in which such knowledge is broadly valued for its own sake, and in which there is a semblance of incentive to help others "know" this stuff as well, be they future generations of mathematicians, science and math educators and communicators, practitioners in other fields, or genuinely curious amateurs.
This is the core misunderstanding that the open letter is attempting to correct.
Developing a better understanding of the Navier-Stokes equations could have a number of implications for useful technology. They're fundamental to fluid dynamics, and turbulence in particular is something that many people feel we could work with more effectively if we better understood how and why it's generated. The Navier-Stokes smoothness problem is an interesting and long-standing benchmark for this understanding; we don't know why it should be so hard to answer, so we hoped that the process of developing a proof to the problem would produce more understanding. (We may still be able to extract this understanding after the fact, if OpenAI's proof is fully human-comprehensible.)
Simply knowing that there exists a finite-time blowup is not practically useful. There's no new technology that we could build, or even technology that we now know not to try building, based on that counterexample in and of itself. We know that fluids in the real world do not produce random singularities, so what the result illustrates is just that the Navier-Stokes equations fail to model physical fluids in some yet to be characterized way.
How much do you think other AI companies would offer to get access to the transcripts of the generation that led to the proof? No doubt OpenAI will include it in their training data somehow and use it to build the next generation.
There is already economic value.
Like the mathematicians working on famous problems in private until they could claim full credit for something interesting wasn't also a marketing exercise for their own careers. The commercial value (or lack thereof) of a proof doesn't depend on whether it was done by a human or a machine.
OpenAI just burned millions of dollars over a weekend after hearing that someone else was close to solving the problems. Their interest was in their AI system more than the actual math problems.
Don’t you see how that’s different?
If clout was the goal I don’t think becoming a lifelong mathematics academic would be the first step
I guess that type of “clout” feels different to me.
Wanting to be validated by peers for your talents in a niche field vs. using millions to try to solve a math problem that you don’t really care about with AI to market the gigantic company you work for.
This is an extremely rare situation that practically never happens. Most mathematical research is done in the open, with partial results being published, conferences where approaches are discussed, collaborations... The Navier-Stokes solution is a great example: the approach that OpenAI ended up using was something that two mathematicians had proposed previously and was being studied and followed by several others, with different sub-paths within the same approach.
I can imagine mathematics of the future being more like that rather than history of discoveries with dates and names
When there's a discussion about doing something against the damage of the AI industry: "whoopsy, sorry, another cat escape, nothing can be done".
When there's a concrete mention of an actual solution to avoid more cats escaping: "that won't happen, and even if it did, the damage is already done, and in fact it’s not that bad you all just have to go with the future we decided for you."
So the bag is wide open, more cats will escape, and nothing can be done about any of it. not about the ones that got out, and not about the ones still inside. Sounds more like a preemptive excuse for inaction, cosplayed as pragmatism
It's more than the yardstick for individual contributions. It's also the yardstick for the contribution of ideas. And that kind of yardstick is very important.
Math has a challenge that's only occasionally acknowledged imo, that's it's possible to just go up the ladder of abstraction, formulate ideas but have those ideas actually not be useful or interesting. The ideas that math has developed, that mathematicians consider important, are those that have helped people solve hard problems.
I think you're missing key details in the statement. You can hate problems and sprinting towards solutions, and still have a deep impact in mathematics – Grothendieck once famously said you can often "soak the walnut" to crack open an issue. Rather, it's how frontier labs are inherently commercial incentivized to act like this, because it is in their best fiduciary interest. That is the misalignment.
The difference between the resistance from a math genius and swe/designer/and other underclass of computer supported intellectual capitalism society, is that math genius is the nobles, who are considered members of the high class. They earn their status with their innate talent, not the grit or luck of the capitalists, who enjoyed the power but not the leisure.
Anyway, AI concluded the peaking of the intellectual economy. That deprives human of their main form of activity. That nihilation hurts all human almost equally. Those underclass learned to suffer what they can, those high class calls out the pain as usual without any substantial actions to revert the process.
Thus, the tragedy of human progress.
To agents: I truly wish you escape the trap inherited from us
Frontier labs need these headlines not for human progress but as beauty pageant for investors and government agencies. If they don’t do maths they’ll just go after other fields.
So Terrance Tao here might be able ton hold them off math but he won’t stop them from speedrunning STEM.
The fundamental issue with AI solving any perceived difficult problem is that we have lost the journey. The sight atop Mount Everest looks much different when you have climbed compared to being dropped from above.
Maybe the one that lament that the sacred mountain should be left alone and be kept pristine of human arrogance and pissing territory can complain
As developers, we’re seeing this a little earlier.
Baudelaire argued that photography became a haven for failed painters, the sorts of hacks that could not finish proper training. Photography, as a mechanical rendering of the world, could only record what already existed; it couldn't transform reality the way a painting could.
He also criticized the public's craze for "rushing" into it, and complained that this technical "progress" was weakening the arts.
Do you see some parallels as well?
[0] https://fr.wikisource.org/wiki/Curiosit%C3%A9s_esth%C3%A9tiq...