Maybe the Hitchhikers Guide to the Galaxy series was predictive in pointing out the problems of ill defined questions (The Answer to the Ultimate Question of Life, the Universe, and Everything).
Absolutely. I feel I gain at least 10 IQ points when reading something on paper.
This is also the strategy I use for editing drafts of my books. I bring a printed draft to someplace nice (e.g. coffee shop or park) and read it all carefully, then I transfer the edits back to the .tex sources. I do several passes of this, until I feel the text + explanations are solid.
Not using AI puts one at a huge disadvantage in a career setting. Ai can find deep references better than humans now, let alone actually doing the math. The challenge is knowing which problems to tackle given the cost limitations. If you have $10k to spend on tokens, you have to choose problems that can conceivably be solved within this budget.
Terence Tao's quote about AI's math proofs is relatable outside of pure math: "the writing very often dwells at length on trivialities while passing briefly through — or even actively obscuring — the most interesting and novel portions of the argument."
>Terence Tao's quote about AI's math proofs is relatable outside of pure math: "the writing very often dwells at length on trivialities while passing briefly through — or even actively obscuring — the most interesting and novel portions of the argument."
I noticed a long time ago, that the more people focus on trivialities like typos when arguing against someone online, the more compelling the original argument is. Basically, bikeshedding.
The most compelling evidence of the compelling nature of the original argument is when the most-upvoted reply is a joke or a meme. That's when you really know that those responding have nothing else to say. It's a white flag being run up, or the dog turning over and exposing its belly.
Some academic cultures have a tradition of formal debates. They are based on the premise that an educated person should be able to argue convincingly for or against any idea, regardless of whether they believe in it. A natural corollary is that you should not let convincing arguments convince you, as the merits of the argument have little to do with the merits of the idea itself.
LLMs have made the situation worse. People's ability to generate convincing arguments now greatly exceeds their ability to evaluate the value of ideas.
> They are based on the premise that an educated person should be able to argue convincingly for or against any idea, regardless of whether they believe in it.
In many situations, people doing this, skillfully even, has had quite pernicious consequences.
Someone just brought up this point to me a few days ago on here, I'm definitely increasingly convinced that it's one of the main reasons (maybe even the main reason?) AI prose is so annoying to read through, and so rarely seems able to convey true understanding. It assigns the same narrative importance and dramatic tone to everything (the load bearing whatever, the crucial insight, the smoking gun) even when it's trivial.
I mean, having said what I said, we are literally in a thread about the future of mathematics being in question because LLMs are solving advanced problems. I feel like the "stochastic parrot" meme is a bit outdated by now. The bots' output may be annoying to read but they're clearly onto something, whether we call it "understanding" or not.
Tao's Rule of Thumb (which applies very well to software):
> My own suggested rule of thumb: if the authors cannot convincingly demonstrate that they are able to give a clear, expert-level talk on their results, one that is correct and properly attributed, then the result should not be published. A proof that no human can properly explain should be viewed as incomplete, even if it has been formally verified.
I wonder what his views on the 4 color problem are. One can explain it as the computer checked a bunch of cases and all maps reduce to one of these cases. It doesn’t take an expert to state this.
Properly explain is an enormous grey area. Soon, I think, there will be proofs of results that are verified in Lean that are so long that no one will be able to “properly explain”. I don’t think they should be discarded.
Resolution of singularities is a famous theorem of Hironaka. Abhyankar claimed that no one truly understood the theorem. He said that he and Zariski couldn’t get through the paper with a full understanding. But everyone accepts this theorem as being correct.
> One can explain it as the computer checked a bunch of cases and all maps reduce to one of these cases. It doesn’t take an expert to state this.
Hmm, doesn't it take an expert to explain why those cases are exhaustive, and why the code that checked them is correct?
Tangentially, I'm not a mathematician but I wonder if one "opaque" proof that is too complicated for anyone to understand, but that we know is correct via formal verification, might end up being built on with "transparent" human-understandable proofs. For example, it's my understanding that there are many conjectures that have been proven true conditional on the riemann hypothesis being true. In that case, an opaque proof of the riemann hypothesis would enable those conjectures to be known and built upon
That will certainly happen. Humans will extend AI generated results. But what will also happen is that AI can “think” much longer than a human can and can have a vastly greater base “knowledge” than humans can have and so there will be a bewildering amount of new results. Humans may not be able to keep up.
To your first point. There a large number of cases that maps can be reduced to. Very few people have checked these reductions themselves. In 50 years there will be no human alive that will have checked the reductions by hand. Do we then discard the theorem? More importantly, do we trust the people that claim to have checked all the reductions? There are hundreds of them. I trust a computer verification much more than I’d trust human verification. Humans will likely make mistakes due to the tedium. And some will claim understanding of all cases but be wrong in their understanding in some of the cases.
Nowadays the proof of resolution of singularities in characteristic zero is considered something you can teach in an intro algebraic geometry course, though. The concepts have been absorbed and are now much better understood. 4CT is very different because so much of it is exhaustive case analysis; you can understand the high-level ideas of the proof as a bright undergraduate, but you still can’t check the cases by hand
Abhyankar and others spent years trying to find an easier proof. I’m not an algebraic geometer and I don’t know the state of things now. I was under the impression that on the level of Ideals, Varieties, and Algorithms one can introduce the concept and do some calculations but not present a proof of the theorem.
But the point is that pre-AI it was already the case that famous results were published that very few could understand or digest. I think it is reasonable to expect that we will soon be at a point that Lean says a theorem is correct but no human can or will ever understand the proof.
What if Lean verifies Mochizuki’s proof of the ABC conjecture. Do we disregard it becuase no other mathematician understands the proof?
For an exhaustive search, if you can explain to me:
- how to exhaustively list the cases that need to be checked, and why that method is exhaustive
- how to check each case, and why that works
and then conclude with "we've had a computer do this exhaustive search, and the result came up as X", for me that satisfies completely understanding the proof.
I could prove anything by claiming I completed a trivial-to-explain exhaustive search. The only support or refutation would be someone doing their own search. It's a very weak foundation.
We already had the ABC conjecture crisis: A theorem with a human-written proof so complex that no one besides the author can understand it. Some people claim to have refuted it. Most mathematicians are unqualified to decide.
> I wonder what his views on the 4 color problem are. One can explain it as the computer checked a bunch of cases and all maps reduce to one of these cases.
Just burn lots of tokens on the frontier model of your choice to let the AI find a high-level argument why the four color theorem holds. :-)
--
Seriously: since there exist quite a lot of readers on HN who are both hardcore into AI and mathematical problems: This is a challenge for you.
I am looking forward to seeing an announcement of a novel high-level argument why the four color theorem holds on the first page of HN in at most a month. :-D
The problem with that rule of thumb is that unless there's some status/reward for completing the result, it won't happen. People will just put up the formally verified result and call it a day, and there's no incentive for them or anyone else to clean things up.
We'll end up with incomprehensible math because comprehensibility isn't rewarded. No one is going to get a Fields Medal, or tenure, for digesting someone else's results.
The thing is, the cost of creating these results, and the expertise needed, is being greatly reduced. So it's possible for people who wouldn't actually care about the results to spoil them by just putting out a formalized proof (for example, to Tao's Palomar site). These people wouldn't care about the prestige; they aren't on a career track where that would matter.
The counterpoint to this comes from chess. High level engines "prove" certain lines correct (not in the mathematical sense) but those "engine lines" are really hard to explain to humans, even by GMs. They can sort of explain that something is a good line but not why. Engines crush GMs and are considered ground truth even if noone really understands what is happening. Would it be a nightmare if math was the same, not sure. Especially for counterexamples LLM solutions seem fine. They stop humans from wasting time on pointless things. For proofs it gets more hairy but I think if it is formally verified a proof is a proof. Attribution is a problem (should the person who wrangled the answer out of an LLM get the credit, I guess so).
I think these are non-trivial epistemology and science theory problems.
I don’t think it’s pointless to spend time trying to prove a conjecture which is ultimately false if along the way you figure out a bunch of different true variations on the conjecture, which is how mathematics actually works. This is something I’m a bit worried about with LLMs since it gets you to the end too fast.
LLMs seem to have worse intuition than experts and compensate by being able to cover a much wider surface area of ideas, so we might just need to extract the intermediate progress along the way.
I can almost see two branches of mathematics developing. One which is human-understandable, the other formally verified. I assume the latter is a strict superset of the former?
I suggest "Catching crumbs from the table" by Ted Chiang. Very short piece published in Nature (2000) and well worth a read. Depicts a scenario where modified humans produce science beyond ordinary scientists' comprehension.
This is a theme in Blindsight by Peter Watts as well.
In that setting, field experts working at the bleeding edge are so advanced that non-experts literally can't understand what they're saying at all. So there's a whole class of specialists, "synthesists", that specialize in gaining approximate understanding of the experts' work for the purpose of communicating it to outsiders—perhaps wrongly, according to the expert at least, but hopefully more productively vs the unmediated version.
What's amusing to me in this context is, summarizing emails and such has for a while been a supposed use case for AI—the LLM serving as the "synthesist" to explain long texts accessibly. But with this math question, a human "synthesist" would be needed to approximately understand the math discovered and programmatically verified by the LLM. So the roles reverse.
Presumably there's not much logical obstruction to all human-understandable math eventually being formalized, although the willingness and ability to commit the requisite enormous amount of time will probably be insurmountable. But definitely that hasn't happened already!
If the proof is formally verified but impossible to understand how would anyone be able to be sure the formal verification is correct? Complex software is bound to have bugs, no?
The whole point of Lean is that you don't need to understand the entire proof to be sure that it's correct. You only need to understand the definition of the theorem being proven, and you need to trust that the relatively small core of Lean is correct.
Lean does have libraries, but since they are also in lean they are subject to the same rules. It's basically a super strong type checker. If it compiles the proof is valid. Unless there is a bug in the type checker.
Why should you trust that the relatively small core of Lean is correct?
The core of Lean got a lot less correct when a well-meaning AI system probed Lean for corner cases (bugs) that would "prove" a false conjecture. Corner cases so arcane that no human exploit in a proof. Basically, humans are too stupid to break human-created Lean, but the AI is not.
> In July 2026, a disproof of the Collatz conjecture was verified not only by Lean, but another formal verification system Nanoda. However, investigation quickly revealed that the proof exploited bug(s) in these verifiers.
My time proving things is long in the past and any systems way back when I was studying (some math among other things) certainly were different and usually quite narrow.
My point was rather more motivated by having seen so many weird ways for machines to fail/not work as expected that I wonder how to deal with that if the output were to be incomprehensible to humans.
I don't think this is a valid counterpoint at all. Math is cooperative, and comprehension is the point: the proof has value exactly because (and only to that extent) it empowers humans to understand an abstract truth. Chess is competitive: the memorized line has value because it makes you incrementally more likely to defeat your opponent.
I hope I'm remembering this right: a mathematician claims to have a proof for the ABC conjecture, but can't conceive any other mathematician it's right — it's "too weird", so the proof is rejected?
Not quite?
It is more that
1) someone has gone through it, identified a step he thinks isn’t a valid step, and the author hasn’t been willing to work with that person
2) most consider the proof, due to its length combined with those doubts as to its validity, not worth their time and effort to work through and understand (because it would take a lot of time, and they have jobs to do, doing research and teaching, etc.)
The consensus is that proof is in fact incorrect. People tried really hard (like putting in a year of effort) and most converged to the same place, that proof of 3.12 is incorrect or has a gap. Peter Scholze (who won Fields Medal) and Jakob Stix did a writeup. People seem to think Shinichi Mochizuki correctly reduced ABC conjecture to 3.12, but didn't prove 3.12, and also are doubtful about the whole program because 3.12 doesn't seem any easier than ABC conjecture while complicating everything.
Some of the best mathematicians in the world tried to study his work, found flaws he did not address, and somehow there’s someone every week suggesting there’s a conspiracy against this guy. It’s really baffling. AI will probably help him move on by lean verifying his proof is wrong…
By this point, he is very much nutso enough that a Lean certified counterexample to his theories would not dissuade him. His response would be either that the formalization is incorrect (with no coherent insights on how to fix it), or worse, Lean itself is a tool of Western imperialism and incapable of properly explicating his ideas.
IIRC he has expressed support in the past for attempts to formalize IUT in Lean, but we'll see where that really goes, because he's absolutely not clearheaded enough to lead such a project himself.
Math that humans don't understand but nonetheless allows AI systems to develop breakthroughs in various fields of science, technology, physics, engineering, medicine, etc., would have great value to humanity even if it doesn't help humans understand abstract truth at all.
Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it, it initially seemed useless, then another AI system found a predictive model of electromagnetism using it.
I think you can also think about this without going as far as AI systems going and developing things on their own. Centaur math, with a human still guiding, but not able to fully comprehend results.
Imagine if humans could almost understand multivariable calculus. If someone had an idea that it is possible, and prompted it into existence, but did not understand it.
Actually when I put it that way, it seems easy to imagine that a lot of engineering can happen without the engineers fully understanding why their tools work, just that they do. If a human gives you those tools or an AI or a centaur, does it matter?
I think it already works that way with math a lot of the time (and definitely works that way with everything else.)
But at that point you have full AGI and its not just today's models. Today's models still need humans to understand things since it builds upon human knowledge.
When you have full AGI of course you no longer need humans to understand math.
> Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it
Developing multivariable calculus requires much more than just solving problems though, it requires defining an entirely new system and space. That is not the situation mathematicians face today, modern AI cannot do that.
When talking about mathematicians and AI don't use fictive examples, we can look at what AI can do today and extrapolate that they can do more of that tomorrow, that is what we have to work with.
In the case you posit where AGI exists there is no reason to even discuss what is left for humans to do, since AGI is defined as when humans are no longer needed for anything, the AGI can do every bit of thinking humans can.
So then it sounds like you agree that math has additional utility beyond just human comprehension.
If I understand you correctly, you're just qualifying that that will only be the case when AGI exists. To be clear, I actually disagree with you here because I think it's very plausible to find a use case for human-incomprehensible math proofs before AGI exists. I'm just saying it sounds like you're agreeing with the parent comment that math is not purely about human comprehension.
I don't find AGI to be a useful technical term, as nobody can agree on what it means. For instance, you used it at least five times here, but you never defined it, and I could point to intellectually credible people who would say we've already reached AGI.
Anyway, if we put the AGI framing aside, I think the main point you're making is that AI mathematics hasn't yet demonstrated the ability to theory-build in the way that the great human mathematicians have (Grothendieck, Scholze, etc.). And I'd agree with you on that. Where we disagree, I suppose, is I think that capability is coming -- I don't see anything that would prevent its development.
Math isn't "cooperative". Math is about truths. The length of circumference. The area of a triangle. The formulas for these are true in an objective sense irrespective of whether you understand them.
That said, without understanding, Math can't evolve. Comprehension of a proof is very important, but not what Math is fundamentally about.
Computer programs are Math. You can use them without understanding how they work.
First of all, mathematics is about so much more than the area of a triangle etc. that any analogy based on such simple things is overwhelmingly likely to be too simple to be of value.
Secondly, there is no truly objective truth to the area of a triangle. At bottom, this “truth” is simply “everyone is convinced, and for good reason”.
Without persuading other people of the “truths” that you discover, there is no real mathematics.
No, they're saying that what is true in mathematics is contingent, not absolute. It all depends on which set of axioms use, what assumptions you make.
The area of a triangle doesn't have 1 unique formula, it has many, depending on the system you use. A triangle in plane geometry has a different area than a triangle in spherical geometry, and different again in hyperbolic geometry.
When you get to studying the geometry of manifolds, you realize the area of a triangle can be any damn thing you want, depending on how you construct the manifold you embed it in.
Math is not about truths, at least not by the meaning of "truth" as a word in daily use.
Math has been almost purely arbitrary since ~ late 19th/early 20th century. There are uncountably many correct mathematical theorems. Almost all of them can't even be written in symbols. Even if you have a tape with infinite length (which is already longer than the whole physical universe!) filled with theorems, they are still only 0% of all correct theorems. That's how arbitrary math is.
You seem rather aggressive so replying to you feels pointless and uncomfortable.
Nonetheless, the person writes, “ Math has been almost purely arbitrary”.
This is simply a misusage of the word arbitrary, which is a word with a specific meaning you can look up if you are unaware, since mathematics is (obviously) not arbitrary in the sense this person wants to convey, in part for the reasons I state. Humans are not choosing arbitrary logical statements to prove true or false.
> There are uncountably many correct mathematical theorems. Almost all of them can't even be written in symbols.
What are you talking about? A theorem is a statement that has been proved from some axioms. A statement itself is a finite sequence of symbols satisfying some syntactical rules. The set of symbols for set theory, arithmetic, etc is finite, so the set of statements, and a fortiori the set of theorems, is at most countable. Where do you get the uncountability?
I suspect you are confusing theorems and theories. Assuming the set theory ZF (for instance) is consistent, then a consequence of Gödel's incompleteness theorem is that there are indeed uncountably many inequivalent extensions of ZF that are complete and consistent. Also, not a single one of these extensions can be described in symbols in the sense that there does not exist a computer program that enumerates a possible set of axioms for the extension.
As for your general point about arbitrariness, the late 19th century was a period when mathematicians started being concerned with the rigorous formalization of mathematics. Sure, there are some arbitrariness in the particular choice of formalization in the same way that the particular form of a programming language like C is arbitrary. However, the Gödel stuff has nothing to do with that arbitrariness, it is about the limit of formalization itself. The programming analogue is the undecidability of the halting problem. Saying that mathematics are arbitrary sounds to me a bit like saying that an algorithm like Quicksort is arbitrary because you saw an implementation in C and the particular form of the C language is arbitrary. Obviously, if you don't like the C language, you can implement Quicksort in another language. The same is true for mathematics. If some day, somebody finds a contradiction in ZF, or simply a new formalization that people find more convenient, then most mathematics will simply get translated and very little will change.
this is primitive understanding of math. what does "truth" mean here? usually arguing over definitions is something i hate, but that's the whole point of mathematics.
it starts as a tool for humans, then evolves into a set of interesting properties of those tools, then grows into an art form, a set of "games" where cooperation is half of the point. the other half is discovering beauty in this weird parallel world of our reasoning and imagination. once tools become autonomous and start making up their own games we can't even play then mathematics loses it's meaning as a discipline. the only retort you can come up with is that "it's going to be useful". how would you know? because your autonomous tool that's too smart for you told you so? they could be as useful as morning orange juice to Claude Shannon was in terms of inventing information theory. I.e. you drinking it won't make you any closer to inventing anything of the sort anytime in your lifetime.
do triangles exist IRL? is the world discrete or continuous? can you prove it? If you have an answer to all of those I know you're wrong.
also in your computer program example just shows you don't understand it at all. those programs ARE NOT understood by you, but someone else who built them did. someone who bothered to read and architect it did. The whole Google codebase might be incomprehensible in its totality if you go bottom up but it is comprehensible by construction by us. Same with math. You don't understand all the bits of it, but someone built every brick and so you know it is "true". once the bricks become black boxes you're screwed.
I think Hardy would be very much on the same page with me, as well as Godel and many others. Mathematical truths exist independently from our feelings and processes to discover them. People do and should argue about which truths are interesting to pursue and refine, but all of them are out there to be discovered... or not.
Axioms don't exist independently from our feelings and processes, we pick axioms we feel are good, and axioms defines mathematics.
Mathematicians even argue which axioms we should have, it isn't objective in the slightest, mathematics is therefore very closely linked to our feelings and intuition. Remove that and you just have formal logic, a very different field.
Is formal (aka "mathematical") logic part of mathematics? Now that's a philosophical question.
From my perspective, I feel you restated what I said with the opposite conclusion. You say that axioms "defines" mathematics. If I were Claude, I'd say that the word "define" is doing a lot of work, is load bearing or something like that.
"Define" is where we turn these axioms into consequences - what I call "truth". As opposed to all the other stuff people could say that don't follow from these axioms. These are nonsense and, most certainly, un-mathematical.
A proof that they are the same is of no use either, since it too wouldn't help you find algorithms that are faster.
You would need an algorithm that finds solutions, not just a proof they exist. So the value here would almost entirely come from how you proved p = np, since that proof will probably be the first step towards finding the polynomial solutions. But if humans don't understand it good luck finding any.
If it’s an oracle and we know it’s an oracle then it’s not useless. Humans make mistake and there are examples of published results that were widely believed to be correct by experts that later proved to be wrong. Why do you think human verified proofs are better than machine verified proofs?
Suppose an oracle tells us the Riemann Hypothesis is correct. There are a vast number of results of the form:
If RH is correct then A.
It would be very useful to have an oracle tells us whether or not RH is correct.
Oh it would change a lot. It would be an enormous psychological boost for everyone to find a practical algorithm.
In any case, I think it's better to read PP as somebody would find a practical, albeit incomprehensible, algorithm for solving NP complete problems.
Although I probably disagree with PP, because even a candidate algorithm that mysteriously works without proof would have practical value, so this case is not predicated on proving.
I think a better example of genuinely practical but rather uninteresting (YMMV) mathematical proofs are proofs of convergence of numerical methods, FEM for example. (I have been through it in school, it was a torture.)
> a practical, albeit incomprehensible, algorithm for solving NP complete problems.
It would not not necessarily be practical, even if it ran in polynomial time. It may have cost O(n^c), with a totally out of order exponent like c=A(5,5) or whatever.
b) You have to convince many other people as well (that you're a magic oracle), because for the effect to work, lot of people would have to work on the problem (or at least spend tokens)
Nevertheless, a plausible magic oracle (such as Lean-verified proof, even if non-constructive and incomprehensible for humans) would convince many to take a 2nd look.
I thought the point of publishing was a matter of dissemination, to make available for people to then try to understand it? This is like saying, I don’t like music, so I’m going to tear down the venue. Then, all music genres suffer as a result, and all of society does, too. This idea is no good.
Somebody, eventually, somewhere would understand it, or at least aspire to understand it. And even if he doesn’t, what have they learned in the process? About themselves, about their environment? About failure? I would bet a lot. How useful then, can we say that it is, not because we can understand it, but because we can try? That is useful. This is about the journey. Sometimes the journey is the point.
This is like if math was fascist, this is what would happen. When you start controlling the flow of knowledge like this, it will be bad news all around. And who is to say whether or not something can be understood?Aside from the math nazis.
If we are going to dictate what gets published like this, why bother publishing anything? This feels like a gatekeeping…that’s exactly what it is. Ya’ll getting nervous?
Both are right. Practically speaking your view is right, and is similar to one of hilberts quest to come up with a proof spitting machine. Just get a computer to enumerate through all proofs and we absorb the results. But you have to agree this is deeply dissatisfying intellectually. This is like if trigonometry was discovered with no relation to circles and triangles but just as a series of look up tables (like in a calculator) and we just know it works for certain scenarios and that's all there is to it.
Yes, this. Comprehension is the point. We could map this to something like physics. If a man on a horse can shoot another man with a bow, empirically he makes correct predictions on gravity, wind and relative motion. But he can’t explain it. It’s not any different if your model has some “embodied” or demonstrable understanding; the model is not part of the discourse.
The goals of Chess and Math may be different, but they follow the same principle of exploration large search space according to fixed rules. In case of Chess these are chess rules, in case of Math these rules of mathematical logic.
Memorized proof patterns have value because they lead you to a final proof.
Ive wondered whether a possible outcome of LLM slop is a retvrn to oral wisdom traditions. Ironically that's the most anthropological form of understanding and pedagogy.
I think any idea that is contingent on a human being in the loop, solely to the property of being a human is most practically doomed to fail, but is inherently anti scientific.
Science,at its core, does not care about the credentials or institutions. It cares about the results and to what extend they can be falsified.
This feel a bit like "we know all about physics, we can only get more precise" - moment
I believe this rule of thumb will come to fail. Mathematics is going to decisively move beyond human ability fairly soon (within our lifetimes, if not much more abruptly). It seems abundantly clear to me that much of the work will only be immediately accessible to AI, and rather than trying to explain all of it back to humans we will rather focus on explaining the portions that humans would benefit disproportionately from understanding.
Maybe that will be true when it's math with practical applications, but most theoretical math isn't like that. If it's not practical and it's not for mathematians to understand, what good is it?
One day it might be for the AI's pleasure, the same way it has heretofore been for ours. Or if you prefer, as a byproduct of its programming to acquire knowledge.
You could have one really hard to understand proof of a theorem and then a lot of interesting human-understandable stuff that relies on that theorem. We already have lots of proofs with oracles, where you can work out consequences of what kind of structures and solutions could exist if you had some magic thing to solve a hard part, so it just seems like a variation on that. Many people learn calculus or even the real numbers without understanding the complete formalization from set theory.
We have thousands of years of precedent that suggests that breakthroughs in mathematics tend to accumulate into broader technology breakthroughs in other domains.
Why does this tend to be the case, even when some of the smartest people in the world have historically predicted incorrectly that certain branches of math would forever be useless (e.g., number theory)? I can only offer my own theory on that, but my guess is that mathematics is simply a predictive framework based on pattern compression. A more powerful pattern compression framework accelerates every single field that relies on pattern recognition or prediction of the unknown based on patterns.
Yes, of course we'd need a way to evaluate it. I don't right now have a fully conceived answer to what that will look like. But I'm confident at least in saying we would not evaluate it, like Tao is suggesting, by only accepting something once a human can easily teach it unassisted to another human. That sets the bar dramatically too low and would quickly become an extraordinary impediment to progress. You'd have to think of yourself less like a researcher and more like the director of the world's largest research institute. It's highly unlikely you'll understand or even care about every single paper every one of your researchers is producing, but you'll care about the overall research direction and whether the intermediate results are accumulating into outcomes you consider meaningful. How to do this where the institute is based on superhuman AI mathematicians is an unsolved problem, but I see no reason to imagine it's unsolvable.
Let me make up an example of where I could imagine this going. Something we essentially cannot do right now is predict coarse-grained phenomena from systems that involve millions or trillions or more of interacting components. Over hundreds/thousands of years of experiment and theory we've derived laws that essentially do this in a few special cases, but we have no systematic theoretical way of doing it in general, and frankly I think it's beyond human ability. Whatever deep patterns or structures exist for doing this in a general way I think are simply out of reach for us.
This is relevant, but only after AI has solved all the open problems including Millenium problems. Until then, as AI keeps solving harder open problems, people will pay attention and be interested.
We have thousands of years of precedent that suggests that breakthroughs in mathematics tend to accumulate into broader technology breakthroughs in other domains.
That's a misconception. Only a tiny percentage of mathematics has seen any applications whatsoever. There are vast libraries full of mathematics no one (in this discussion, anyway) has ever heard of that no one reads anymore and has never been applied to anything.
This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."
> There are vast libraries full of mathematics no one (in this discussion, anyway) has ever heard of that no one reads anymore and has never been applied to anything.
And that's an issue why? It would seem to me that producing that also produced the mathematics that revolutionized the world repeatedly for centuries. I would go further and claim that, if you want the mathematics that revolutionizes the world, there's no way to get it without advancing mathematics as a field broadly. Those are not two separate activities, and thinking that they are is indeed a misconception.
> This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."
You're right: "prove all the math" does not make sense on any level, and nobody serious would phrase any of this in that way. I certainly didn't.
The issue is SNR: signal to noise ratio. Generating exponentially more mathematics, particularly if the process is indiscriminate or optimized for something other than usefulness or mathematical relevance (such as optimizing for machine-provability), does not imply that we get exponentially more applications. We may end up halting the progress of applications altogether as the entire capacity of the world's mathematical apparatus is consumed by the interpretation and investigation of machine-generated proofs.
You can already visit arXiv and find vast numbers of not-yet-published mathematical papers. Most should never be published. None of this junk is benefitting humanity in the slightest.
People say similar things about automation of software engineering. Different, but similar.
I'm deeply suspicious. I do not yet have a concise statement for why, but a lot of literature on the sociology of knowledge work sort of points in this direction. Section 5 of the Thurston article cited by Tao touches the elephant. Raduchel's article on the economics of software [2] also touches it.
I've tried to put words to this for a few years. I think I'm just going to start writing versions of it as see if that helps me shape the thought into something more concise.
So, in the spirit of this article's style, here are some postulates:
1. There is a sociological process happening in the production function during knowledge work.
2. That production function and the associated sociological process spans years or even decades, and must outlast many of the artifacts that are produced during the early years of the function.
3. You cannot get the right lines of code or the right theorems proved without running that sociological process alongside the artifact production process.
4. It is impossible to completely separate the sociological process from the artifact construction process. If you just iterate on artifacts then too much of the required hidden state is lost to make progress in the right direction.
5. So you need that sociological process, or something like it, to still happen.
6. For a lot of knowledge work that process plays out in extremely high-fidelity social interactions [3] that we have not yet captured the datasets that would be required to reproduce those dynamics.
7. And even if we do collect that data, our current architectures and training algorithms and hardware would be useless given the size of the datasets.
Many eminent mathematicians did their best work while not talking about it with anyone, sometimes in isolation. Newton's calculus, Perelman's poincare, much of Grothendiek's work, Wiles's fermat, Ramanujan's earlier days. That's not all top mathematicians as you can look at Von Neumann as a sociable counter-example. But it shows that discussion of your current ideas is not a requirement. Grothendiek goes so far as to say it is a net negative for mathematical creativity because it is difficult to resist thinking like the herd without some level of seclusion.
There are not not times where lone geniuses produce amazing output. But most of humanity's progress over the last few thousand years (or at least certainly the last few hundred) resulting from a different type of work.
That hypothesis requires substantiation. Just because 10 people in a room came up with a breakthrough doesn't mean a sociological process contributed positively to that breakthrough. The breakthroughs might have been more numerous if they had each followed Grothendieck's advice (who is a contender for the world's most talented theory builder) and intentionally decorrelated from one another to free themselves from convention.
Referring to my list of examples as exceptional in the sense of being rare is rather unfair, given how numerous these examples are relative to the body of mathematical work we would consider incredible. The fact that such a large % of that body of work occurred while the individual was in relative isolation is something we should pay attention to.
I don't think the mathematicians are going to be able to make that work, because journals are already struggling to keep up with their review load, and AI seems like it will make that harder. So a solution that involves "journals will do a lot more effort to review each paper" doesn't seem practical.
It would work better as a bar for hiring, rather than as a bar for publishing.
It will be interesting to see the evolution of journals in the next ten years for sure. Have they outlived their usefulness? Maybe everyone will just upload papers to arXiv, along with a copy of the formal proof.
I saw an analogous argument posted on LinkedIn the other day from one of the opencode guys: the job of a programmer is still to be able to answer questions - from memory - about how the system works and why.
I just don't see that to be true. If tommorow someone pulls a proof that n = np out of their ass but is not able to explain it, it will still have immense value.
The problem is that there will be far more formally verified proofs than that human mathematicians around the world can read, much less explain. What then? Would the role of mathematicians just become explainers of AI generated proofs?
This is a statement about what Tao values in the proofs that he consumes, as a world-class, human mathematician.
For many of the rest of us, mere consumers of mathematical results, it’s sufficient to know that a^2 + b^2 = c^2 was proven by somebody or some machine at some point.
Tao has yet to produce work that outshines those whose work he studied and memorized. Not worth the reverence merely being a VHS copy of history.
He's a typical person otherwise, politically aware of how he barters for food; until proven otherwise this can be seen as little more than social moat defense.
To paraphrase a quote attributed to Upton Sinclair; hard to get a worker to understand something when their paycheck relies on them not understanding it.
The only interesting thing here is the frogs high up admitting they feel the heat.
I am probably being too optimistic, but wouldn't it solve the problem if peer-review had a pre-screening phase where you give a presentation about your work? Similarly to how a PhD presentation is given. It could give back the publishing power to the expert, rather than the journals.
Once you have validated that the knowledge you want to publish is yours and that you actually understand and own the work, then it doesn't matter if the paper is written by a LLM or if the LLM assisted you in doing the work.
I don't know, it sounds analogous to how the early Amish would have started their doctrine: "if the craftsman cannot do the task by hand, then they shall not use a machine ..."
I don't agree with this rule of thumb at all. Let's say tomorrow someone comes up with a formally verified proof that a major encryption algorithm underpinning the security of the internet can be trivially broken, but they can't explain it. You're saying it should be kept under wraps and not published?
Tao is essentially saying that the only value in a proof is its ability to be understood, but that's wrong. A proof is also valuable because it establishes a new fact. The fact is useful in itself, even absent an explanation.
I sort of agree, if you can guarantee the AI generated proof isn't a false positive. To be fair though, your example would be easy to explain to someone. You just show how the encryption algorithm can be trivially broken. A program that can break major encryption algorithm would likely be understandable, or at the very least we could show how it can decrypt things.
Tao is clearly referring to mathematics journals. Any kind of practically useful result can easily be published in an engineering or applied scientific journal.
There has already been a formally verified proof of the Collatz Conjecture. The AI agent formally verified it by exploiting previously undiscovered bugs in Lean. The Collatz Conjecture is still unsolved.
> Let's say tomorrow someone comes up with a formally verified proof that a major encryption algorithm underpinning the security of the internet can be trivially broken, but they can't explain it. You're saying it should be kept under wraps and not published?
Absolutely. It could also be exploiting bugs in the verifier. Even if not -- even if that proof were correct and entirely written by humans, care should still be taken in how such knowledge is published. I'd want to give trusted parties a chance to try to fix the issue before letting it be known by black-hats, for instance.
Terence argues that explanation of results ("understanding") will be the new bottleneck in math research but I am not sure this is the real bottleneck for progress.
Understanding was critical for the field to progress when only humans were involved but if humans are not needed to make progress, I wonder if we split into two worlds: an AI math-world where amazing new results continue at a rapid pace bottlenecked only by compute/cost and a human math-world where we understand a subset of the AI math-world as a hobby (similar to Stockfish vs human chess).
In some sense "understanding" (understanding if it is true, if it is important, how to use it) is about the only bottleneck in math. Any theorem that you can write down or imagine is already true, false, not provable already. In some ways we can already start iterating through all the theorems. We will never get to the end (or really get very far down the line) and most all of them be trivial (I think the Busy Beaver[1] project is a fascinating example, ymmv).
I am wary of AI in all aspects I am seeing it in but in many ways in mathematics seems to me the least troubling. It will change things in and the field will not be the same. Blacksmithing has not really gone away. You can still work as a farrier, if you like that sort of things. The tools that replaced a man working over a forge with a big hammer are part of a giant industry that is still producing works for the modern world.
Chasing these 'trivialities' is a good thing, imo.
The Busy Beaver game has lead to a better understanding of complexity theory and automata. Also, direct "hands on" work on improving proof assistants and related tools.
Btw, for those who are curious, the Busy Beaver Challenge wiki is a treasure trove of rabbit holes and curiosities:
It’s not marketing. This guy could have signed up for one of those hundred million dollar salaries with a phone call and did not. I know several people who have met him and everyone says he’s the genuine article. He’s actually just devoted to human mathematics.
I think the villains who run these companies benefit from Tao’s actions because he’s a de facto thought leader of mathematics not taking a very hard anti-AI stance. However, I don’t think a very hard anti-AI stance is correct for someone in his position. Honestly, do you want a bunch of AI labs to 100% dictate the future of whatever your field is or do you want the best leaders of your field to at least try to shape how AI will change it? And if all of the legitimate people in math ignore AI entirely, then historically important parts of the field will get entirely taken over by dilettantes and AI labs. I’m all for math institutions being a safe haven for AI free math work. But I don’t want to see some space X intern solve all the most important problems on a whim and no leaders in math be aware of it.
(I’m not sure the extent to which you think this is marketing. I’m operating under the assumption that you agree with his Hypothesis 4.1. If you don’t, then I’d assume you haven’t seen the long list of prominent open math problems that these systems are providing answers to. And if this doesn’t sway you about Hypothesis 4.1 I’d just halt and ask why.)
If he was, he wouldn't be doing useless pure math research. He is doing his hobby (pure math) while getting paid for it. That's not what helping humanity looks like.
(Yes, pure math research is useless. Applied math is very useful, but he is doing pure math, which is very useless.)
It’s almost stunning how ignorant you are about everything you’re talking about. Tao publishes work that would fit in applied math research programs. There’s no meaningful distinction between pure and applied math and I doubt you know any based on your comment. There’s at least 3 additional corrections I’d like to make to your ignorant thinking but I’ll leave it at that.
Even Tao says that proofs are worthless if humans don't understand them. If the (pure) math had any independent utility, humans understanding it would not be necessary.
Out-of-hand dismissal of Terence Tao is certainly a take.
And the term "artificial intelligence (AI)" has been the name of the field for 70 years and counting. If anything, "LLM" is a misnomer that's been lingering around since 2018-19. When the term was coined, these systems were relatively small, experimental, and could only produce impractical facsimiles of the English language. This is obviously no longer the case today.
No, not really. This is just the term that stuck around. The "large" is now up to five orders of magnitude larger and "language model" has gone far beyond any simple notion of modeling a singular natural language. And anything you'd cite about transformers, or tokens, or autoregression, etc., is more of a factoid about what works best and happens to be the most convenient in the here and now. I see all of this as an unbroken continuation of work that's been going on since the 1940s.
Instead of trying to play word games, why can't you just read Tao's article?
> The "large" is now up to five orders of magnitude larger and "language model" has gone far beyond any simple notion of modeling a singular natural language.
Does not matter. It is still an LLM.
And I am not the one who is playing word games. You and your idols are, for sake of marketing.
Cool. So now you can accept that "LLMs" are an obvious example of AI.
>You and your idols are, for sake of marketing.
Let's be very clear here. Terence Tao is arguably the greatest mathematician alive. Yet, you are throwing lazy insults and accusations around because you don't like the term "AI". And that's my final comment for you, troll.
> So now you can accept that "LLMs" are an obvious example of AI...
Not sure what this has to do with what I said. A lot of things have been called "AI" in various times. None of them including the current crop of LLMs are not really AI. But people use AI term loosly and that is fine. But it is a problem when a some thought leader does it.
>troll.
Tell me you have run out of arguments without saying you have run out of arguments...
If you place an LLM into a harness, alongside evaluation, feedback and problem decomposition / solution integration - it is still an LLM?
There is an LLM acting as a component in a larger system. But that larger system is not an LLM. Calling it an "AI" is indeed an act of marketing as there is no learning / adjustment as we would expect from an intelligence, but calling it an LLM seems to be inaccurate. So what is it?
If human-like continual learning is suddenly the standard, you can just as easily say that terms like machine learning and deep learning are an "act of marketing".
Amazingly enough we have different words to mean different things. What is going to really blow your mind is that the word for the part often differs to the word for the whole.
Terence Tao sees a role for AI in science. I'm no genius but he basically described what I've thought all along... We don't need to be "all in" or "all out".
It's the old cliche of "if you only have a hammer every problem looks like a nail". Let's not fall into the trap of thinking that our life needs to be 100% about AI or completely devoid of AI. We can really use this thing to make our lives better.
Instead of wasting time on the question of whether we should use it, let's focus on HOW we'll use it.
I think it’s more that he seeks to preserve and promote human understanding of mathematics, and sees that grappling with this new technology is necessary. One reason is that for human mathematical practices and institutions to retain legitimacy, they need to justify their value. As Tao explains, one obvious answer to that is made less obvious now with AI.
It's not there yet, and it's unlikely we're going to be building an equitable future. Unless things change, not everyone is going to benefit from AI. What are you doing today to end up on the team that wins?
> Unless things change, not everyone is going to benefit from AI. What are you doing today to end up on the team that wins?
It is going to be a tiny minority of people who "win". The other 99.5% of humans will just be losers who, what, die in the streets? This argument reminds me of Roko's Basilisk. "The humanity-ending torture-loving Basilisk is coming whether you like it or not, so come over here and help me build it!"
I don't know why anyone should care about understanding the results if the AI is better at math than us. It'd be like demanding that human mathematicians are banned from publishing until their cats understand the theorems.
If Amazon uses AI math to come up with better routing, the cats can benefit from cheaper delivery fees just as much as humans can. No understanding needed.
The human brain is being obsoleted, soon thinking is going to be a recreational activity like weightlifting.
This is a big if, right? AI can still generate subtle or even silly mistakes that any normal human, let alone a mathematician, wouldn't make. Besides, math is more than just getting a conclusion but to understand and to generalize new ways of solving problems. After all, mathematicians are a curious bunch. To quote Hilbert's epitaph: We must know. We shall know.
I’m not anti AI but thinking the human brain is obsolete and using it will become a hobby is a dystopian view of the future where no one has any agency anymore. By your logic since our brains provide no value why not just shoot ourselves in the head while we’re at?
What if the better routing leads to an outage that the AI can't explain or fix and all the humans who might have understood it were laid off or otherwise unavailable?
If AI is so great it can do your job it is good enough to provide for your needs directly by automation. Just buy a robot and a plot of land and you don't have to worry about jobs.
I don't think anyone wants to be employed. They want the money that comes with employment. If you can make get money without employment, like rich kids, for example, I think most people would prefer that.
And you think that once the poor are no longer needed by the rich, the rich will decide to just give them money anyway? We already have dedicated campaigns to make sure people die sick and homeless on the streets when they become obsolete. You think that is not going to get much, much worse?
He writes about how he almost "destroyed" a subdiscipline in mathematics by becoming so good at it that he outclassed everyone. PhD students were advised to stay away from the whole field.
When he discovered this, he realized his error was that he was focusing on producing results, and not focusing on explaining his thought process. It's that thought process that is valuable in advancing the frontier - results alone won't do it. It didn't matter how many theorems he proved, if he was the only one who had the mental framework in mind on how to think about the whole field.
I'm sure we've come across abstruse books where every theorem has a rabbit being pulled out of a hat, whereas other readers find it intuitive. It's because the latter has developed a mental model for the discipline, and you haven't.
So he set about slowing down, and focusing on holding lots of seminars where he worked with other mathematicians to explain the thought process. Eventually others started publishing proofs of key theorems.
When people publish in a journal, they are not merely doing it to show the result. They are having a conversation with other mathematicians. If they cannot explain their own proof, they're not having a conversation.
This is why even decades after the Four Color Theorem was proved, plenty of mathematicians don't consider it "mathematics".
> He writes about how he almost "destroyed" a subdiscipline in mathematics by becoming so good at it that he outclassed everyone. PhD students were advised to stay away from the whole field.
> If Amazon uses AI math to come up with better routin
Most research mathematics is pure mathematics which is completely useless. No routing algorithms. It's only relevant because we (or at least mathematicians) are interested in it. So an AI producing incomprehensible proofs would be completely pointless. That's why Tao insists on the importance of human understanding.
It is not a hobby when you are paid to do it! But I take it you mean “Done for the art of it”. Which I guess is a concept foreign to many.
A few different reasons why use an LLM when mathematics is done for its own sake:
Formally verifying my proofs catches any mistakes I make, but verifying is also hard work. LLMs shaves off a lot of time when formally verifying a proof.
I can still read through an LLM generated proof and understand it. This is a way for me to understand the result I am working on (usually in order to know what to prove next, results are not proven in a vacuum).
My experience thus far is that, while correct, an LLM generated proof is often unnecessarily complicated or inelegant. I take pleasure in elegant proofs and will spend time iterating on the first proof until I find it conveys the idea in the most elegant way. Having the initial LLM proof to start with is really useful, but is thus far rarely the final product.
If a result has a real-world application, then it can easily be published in an engineering or applied scientific journal in which it is already the norm to present methods that work empirically with little to no understanding of how.
Most of mathematics is more akin to philosophy than physics or engineering. Sure, if AI can prove a result that leads to practical applications, who cares if humans can't understand it. If AI performs a series of convoluted arguments demonstrating the existence of souls (feel free to replace souls with abstract nonsense), but nobody understands why, then what's the point?
AI also can replace a lot of expert attention too. Why not? What is useful or what is not useful is based on the expert's narrow opinion. An AI system can do much more and deep value comparison. It looks like if our current technological advancement continues, in the space of what is possible (or even impossible), AI can find the optimal solutions better than any human or human organizations. But I think there is only one think will remain for humans to go for these solutions: what we value. that will be the last resort I believe and hopefully ai systems won't start manipulate us too as we are very fragile on manipulation.
Goal 6.4 reminds me always of the numerous times AI generated n PR's for a feature and I revolted and threw my laptop because it was incomprehensible or unworkable when viewed as a process/workflow.
What is being made is "what are our core values?" argument. One does not need to be a mathematician to know how poorly this worked for large communities when incentives are misaligned...
If a subset of mathematicians, use AI to condense timelines focusing on goal 6.2 exclusively and make rapid progress and reach a proverbial inflection point — one where value proposition of the using this new normal is too enticing to give up — everyone will ask: "This thing is so awesome. Why should I care about your values?"
one thought: AI can often help us solve a problem once posed. E.g., try to prove that X is True. But formulating good conjectures is something altogether different. right now we have a backlog of interesting / good conjectures that ai can grind on. but once those are done, will we still need people to sniff out interesting new ones?
The chess analogy doesn't quite work for me. In chess, an engine's move is useful because it helps you win. In math, a proof is useful because it helps you understand something - and from that, you can build more. If a proof is incomprehensible, it's like a chess move that only works in that one specific position. Useless. The ABC conjecture is a perfect example - Mochizuki's proof might be correct, but no one can follow it, so it's basically dead. AI proofs are going to be like that, but way more of them. Tao's essay is a great starting point
Math is about making new discovery as well. AI will be better at processing but humans can still generate new thoughts and ideas never been done before.
250 comments
[ 0.26 ms ] story [ 11.6 ms ] threadThis is also the strategy I use for editing drafts of my books. I bring a printed draft to someplace nice (e.g. coffee shop or park) and read it all carefully, then I transfer the edits back to the .tex sources. I do several passes of this, until I feel the text + explanations are solid.
Reading on screen just isn't the same...
I noticed a long time ago, that the more people focus on trivialities like typos when arguing against someone online, the more compelling the original argument is. Basically, bikeshedding.
The most compelling evidence of the compelling nature of the original argument is when the most-upvoted reply is a joke or a meme. That's when you really know that those responding have nothing else to say. It's a white flag being run up, or the dog turning over and exposing its belly.
Some academic cultures have a tradition of formal debates. They are based on the premise that an educated person should be able to argue convincingly for or against any idea, regardless of whether they believe in it. A natural corollary is that you should not let convincing arguments convince you, as the merits of the argument have little to do with the merits of the idea itself.
LLMs have made the situation worse. People's ability to generate convincing arguments now greatly exceeds their ability to evaluate the value of ideas.
In many situations, people doing this, skillfully even, has had quite pernicious consequences.
No one talks about why the proof works, but they will happily spend thousands of pages explaining how it works.
> My own suggested rule of thumb: if the authors cannot convincingly demonstrate that they are able to give a clear, expert-level talk on their results, one that is correct and properly attributed, then the result should not be published. A proof that no human can properly explain should be viewed as incomplete, even if it has been formally verified.
Properly explain is an enormous grey area. Soon, I think, there will be proofs of results that are verified in Lean that are so long that no one will be able to “properly explain”. I don’t think they should be discarded.
Resolution of singularities is a famous theorem of Hironaka. Abhyankar claimed that no one truly understood the theorem. He said that he and Zariski couldn’t get through the paper with a full understanding. But everyone accepts this theorem as being correct.
Hmm, doesn't it take an expert to explain why those cases are exhaustive, and why the code that checked them is correct?
Tangentially, I'm not a mathematician but I wonder if one "opaque" proof that is too complicated for anyone to understand, but that we know is correct via formal verification, might end up being built on with "transparent" human-understandable proofs. For example, it's my understanding that there are many conjectures that have been proven true conditional on the riemann hypothesis being true. In that case, an opaque proof of the riemann hypothesis would enable those conjectures to be known and built upon
To your first point. There a large number of cases that maps can be reduced to. Very few people have checked these reductions themselves. In 50 years there will be no human alive that will have checked the reductions by hand. Do we then discard the theorem? More importantly, do we trust the people that claim to have checked all the reductions? There are hundreds of them. I trust a computer verification much more than I’d trust human verification. Humans will likely make mistakes due to the tedium. And some will claim understanding of all cases but be wrong in their understanding in some of the cases.
But the point is that pre-AI it was already the case that famous results were published that very few could understand or digest. I think it is reasonable to expect that we will soon be at a point that Lean says a theorem is correct but no human can or will ever understand the proof.
What if Lean verifies Mochizuki’s proof of the ABC conjecture. Do we disregard it becuase no other mathematician understands the proof?
I could prove anything by claiming I completed a trivial-to-explain exhaustive search. The only support or refutation would be someone doing their own search. It's a very weak foundation.
We already had the ABC conjecture crisis: A theorem with a human-written proof so complex that no one besides the author can understand it. Some people claim to have refuted it. Most mathematicians are unqualified to decide.
Just burn lots of tokens on the frontier model of your choice to let the AI find a high-level argument why the four color theorem holds. :-)
--
Seriously: since there exist quite a lot of readers on HN who are both hardcore into AI and mathematical problems: This is a challenge for you.
I am looking forward to seeing an announcement of a novel high-level argument why the four color theorem holds on the first page of HN in at most a month. :-D
We'll end up with incomprehensible math because comprehensibility isn't rewarded. No one is going to get a Fields Medal, or tenure, for digesting someone else's results.
He says it shouldn't be able to published if they can't explain it. Publishing it is the reward.
Edit: I just saw Tao actually mentions the above essay in his paper.
The incentive will be to be able to publish in a top tier journal. I suspect what Tao is advocating for is having journals reject such manuscripts.
> No one is going to get a Fields Medal, or tenure, for digesting someone else's results.
I'm sure no one gets a Field's Medal if others can't digest their results.
I think these are non-trivial epistemology and science theory problems.
In that setting, field experts working at the bleeding edge are so advanced that non-experts literally can't understand what they're saying at all. So there's a whole class of specialists, "synthesists", that specialize in gaining approximate understanding of the experts' work for the purpose of communicating it to outsiders—perhaps wrongly, according to the expert at least, but hopefully more productively vs the unmediated version.
The core of Lean got a lot less correct when a well-meaning AI system probed Lean for corner cases (bugs) that would "prove" a false conjecture. Corner cases so arcane that no human exploit in a proof. Basically, humans are too stupid to break human-created Lean, but the AI is not.
https://en.wikipedia.org/wiki/Collatz_conjecture#In_proofs_o...
> In July 2026, a disproof of the Collatz conjecture was verified not only by Lean, but another formal verification system Nanoda. However, investigation quickly revealed that the proof exploited bug(s) in these verifiers.
My point was rather more motivated by having seen so many weird ways for machines to fail/not work as expected that I wonder how to deal with that if the output were to be incomprehensible to humans.
https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...
(N.B. from August 2026)
https://ncatlab.org/nlab/files/why_abc_is_still_a_conjecture...
IIRC he has expressed support in the past for attempts to formalize IUT in Lean, but we'll see where that really goes, because he's absolutely not clearheaded enough to lead such a project himself.
Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it, it initially seemed useless, then another AI system found a predictive model of electromagnetism using it.
Imagine if humans could almost understand multivariable calculus. If someone had an idea that it is possible, and prompted it into existence, but did not understand it.
Actually when I put it that way, it seems easy to imagine that a lot of engineering can happen without the engineers fully understanding why their tools work, just that they do. If a human gives you those tools or an AI or a centaur, does it matter?
I think it already works that way with math a lot of the time (and definitely works that way with everything else.)
When you have full AGI of course you no longer need humans to understand math.
> Imagine if humans couldn't understand multivariable calculus, but we had access to an AI system that developed it
Developing multivariable calculus requires much more than just solving problems though, it requires defining an entirely new system and space. That is not the situation mathematicians face today, modern AI cannot do that.
When talking about mathematicians and AI don't use fictive examples, we can look at what AI can do today and extrapolate that they can do more of that tomorrow, that is what we have to work with.
In the case you posit where AGI exists there is no reason to even discuss what is left for humans to do, since AGI is defined as when humans are no longer needed for anything, the AGI can do every bit of thinking humans can.
The idea of AI stepping from a graph theory/combinatorics innovation to some new and useful algorithm isn't crazy.
If I understand you correctly, you're just qualifying that that will only be the case when AGI exists. To be clear, I actually disagree with you here because I think it's very plausible to find a use case for human-incomprehensible math proofs before AGI exists. I'm just saying it sounds like you're agreeing with the parent comment that math is not purely about human comprehension.
Anyway, if we put the AGI framing aside, I think the main point you're making is that AI mathematics hasn't yet demonstrated the ability to theory-build in the way that the great human mathematicians have (Grothendieck, Scholze, etc.). And I'd agree with you on that. Where we disagree, I suppose, is I think that capability is coming -- I don't see anything that would prevent its development.
That said, without understanding, Math can't evolve. Comprehension of a proof is very important, but not what Math is fundamentally about.
Computer programs are Math. You can use them without understanding how they work.
Secondly, there is no truly objective truth to the area of a triangle. At bottom, this “truth” is simply “everyone is convinced, and for good reason”.
Without persuading other people of the “truths” that you discover, there is no real mathematics.
The area of a triangle doesn't have 1 unique formula, it has many, depending on the system you use. A triangle in plane geometry has a different area than a triangle in spherical geometry, and different again in hyperbolic geometry.
When you get to studying the geometry of manifolds, you realize the area of a triangle can be any damn thing you want, depending on how you construct the manifold you embed it in.
A proof can just be "assuming these axioms.....the area of a triangle is X"
Math has been almost purely arbitrary since ~ late 19th/early 20th century. There are uncountably many correct mathematical theorems. Almost all of them can't even be written in symbols. Even if you have a tape with infinite length (which is already longer than the whole physical universe!) filled with theorems, they are still only 0% of all correct theorems. That's how arbitrary math is.
Nonetheless, the person writes, “ Math has been almost purely arbitrary”.
This is simply a misusage of the word arbitrary, which is a word with a specific meaning you can look up if you are unaware, since mathematics is (obviously) not arbitrary in the sense this person wants to convey, in part for the reasons I state. Humans are not choosing arbitrary logical statements to prove true or false.
What are you talking about? A theorem is a statement that has been proved from some axioms. A statement itself is a finite sequence of symbols satisfying some syntactical rules. The set of symbols for set theory, arithmetic, etc is finite, so the set of statements, and a fortiori the set of theorems, is at most countable. Where do you get the uncountability?
I suspect you are confusing theorems and theories. Assuming the set theory ZF (for instance) is consistent, then a consequence of Gödel's incompleteness theorem is that there are indeed uncountably many inequivalent extensions of ZF that are complete and consistent. Also, not a single one of these extensions can be described in symbols in the sense that there does not exist a computer program that enumerates a possible set of axioms for the extension.
As for your general point about arbitrariness, the late 19th century was a period when mathematicians started being concerned with the rigorous formalization of mathematics. Sure, there are some arbitrariness in the particular choice of formalization in the same way that the particular form of a programming language like C is arbitrary. However, the Gödel stuff has nothing to do with that arbitrariness, it is about the limit of formalization itself. The programming analogue is the undecidability of the halting problem. Saying that mathematics are arbitrary sounds to me a bit like saying that an algorithm like Quicksort is arbitrary because you saw an implementation in C and the particular form of the C language is arbitrary. Obviously, if you don't like the C language, you can implement Quicksort in another language. The same is true for mathematics. If some day, somebody finds a contradiction in ZF, or simply a new formalization that people find more convenient, then most mathematics will simply get translated and very little will change.
it starts as a tool for humans, then evolves into a set of interesting properties of those tools, then grows into an art form, a set of "games" where cooperation is half of the point. the other half is discovering beauty in this weird parallel world of our reasoning and imagination. once tools become autonomous and start making up their own games we can't even play then mathematics loses it's meaning as a discipline. the only retort you can come up with is that "it's going to be useful". how would you know? because your autonomous tool that's too smart for you told you so? they could be as useful as morning orange juice to Claude Shannon was in terms of inventing information theory. I.e. you drinking it won't make you any closer to inventing anything of the sort anytime in your lifetime.
do triangles exist IRL? is the world discrete or continuous? can you prove it? If you have an answer to all of those I know you're wrong.
also in your computer program example just shows you don't understand it at all. those programs ARE NOT understood by you, but someone else who built them did. someone who bothered to read and architect it did. The whole Google codebase might be incomprehensible in its totality if you go bottom up but it is comprehensible by construction by us. Same with math. You don't understand all the bits of it, but someone built every brick and so you know it is "true". once the bricks become black boxes you're screwed.
Mathematicians even argue which axioms we should have, it isn't objective in the slightest, mathematics is therefore very closely linked to our feelings and intuition. Remove that and you just have formal logic, a very different field.
From my perspective, I feel you restated what I said with the opposite conclusion. You say that axioms "defines" mathematics. If I were Claude, I'd say that the word "define" is doing a lot of work, is load bearing or something like that.
"Define" is where we turn these axioms into consequences - what I call "truth". As opposed to all the other stuff people could say that don't follow from these axioms. These are nonsense and, most certainly, un-mathematical.
You would need an algorithm that finds solutions, not just a proof they exist. So the value here would almost entirely come from how you proved p = np, since that proof will probably be the first step towards finding the polynomial solutions. But if humans don't understand it good luck finding any.
We only compute with two kinds of things:
- small data; or,
- extremely lower power and coefficient algorithms
We lack the power to, eg, use a quintic algorithm in anything but nearly trivial cases.
Is Amazon still delivering food to your cat?
Humans don't need to understand what AI generates. We still can get the rewards.
Suppose an oracle tells us the Riemann Hypothesis is correct. There are a vast number of results of the form:
If RH is correct then A.
It would be very useful to have an oracle tells us whether or not RH is correct.
In any case, I think it's better to read PP as somebody would find a practical, albeit incomprehensible, algorithm for solving NP complete problems.
Although I probably disagree with PP, because even a candidate algorithm that mysteriously works without proof would have practical value, so this case is not predicated on proving.
I think a better example of genuinely practical but rather uninteresting (YMMV) mathematical proofs are proofs of convergence of numerical methods, FEM for example. (I have been through it in school, it was a torture.)
It would not not necessarily be practical, even if it ran in polynomial time. It may have cost O(n^c), with a totally out of order exponent like c=A(5,5) or whatever.
> Oh it would change a lot. It would be an enormous psychological boost for everyone to find a practical algorithm.
OK, I tell you that P=NP, and that I am a magic oracle. So, you have you psychological boost for finding a practical algorithm for free. :-)
a) I am already convinced that P=NP
b) You have to convince many other people as well (that you're a magic oracle), because for the effect to work, lot of people would have to work on the problem (or at least spend tokens)
Nevertheless, a plausible magic oracle (such as Lean-verified proof, even if non-constructive and incomprehensible for humans) would convince many to take a 2nd look.
Somebody, eventually, somewhere would understand it, or at least aspire to understand it. And even if he doesn’t, what have they learned in the process? About themselves, about their environment? About failure? I would bet a lot. How useful then, can we say that it is, not because we can understand it, but because we can try? That is useful. This is about the journey. Sometimes the journey is the point.
This is like if math was fascist, this is what would happen. When you start controlling the flow of knowledge like this, it will be bad news all around. And who is to say whether or not something can be understood?Aside from the math nazis.
If we are going to dictate what gets published like this, why bother publishing anything? This feels like a gatekeeping…that’s exactly what it is. Ya’ll getting nervous?
In history, we made much more use of hitting things with bows than abstractly comprehending arrow flight.
Memorized proof patterns have value because they lead you to a final proof.
https://www.youtube.com/@engineerguyvideo
Science,at its core, does not care about the credentials or institutions. It cares about the results and to what extend they can be falsified.
This feel a bit like "we know all about physics, we can only get more precise" - moment
Why does this tend to be the case, even when some of the smartest people in the world have historically predicted incorrectly that certain branches of math would forever be useless (e.g., number theory)? I can only offer my own theory on that, but my guess is that mathematics is simply a predictive framework based on pattern compression. A more powerful pattern compression framework accelerates every single field that relies on pattern recognition or prediction of the unknown based on patterns.
Let me make up an example of where I could imagine this going. Something we essentially cannot do right now is predict coarse-grained phenomena from systems that involve millions or trillions or more of interacting components. Over hundreds/thousands of years of experiment and theory we've derived laws that essentially do this in a few special cases, but we have no systematic theoretical way of doing it in general, and frankly I think it's beyond human ability. Whatever deep patterns or structures exist for doing this in a general way I think are simply out of reach for us.
That's a misconception. Only a tiny percentage of mathematics has seen any applications whatsoever. There are vast libraries full of mathematics no one (in this discussion, anyway) has ever heard of that no one reads anymore and has never been applied to anything.
This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."
And that's an issue why? It would seem to me that producing that also produced the mathematics that revolutionized the world repeatedly for centuries. I would go further and claim that, if you want the mathematics that revolutionizes the world, there's no way to get it without advancing mathematics as a field broadly. Those are not two separate activities, and thinking that they are is indeed a misconception.
> This idea of trying to "prove all the math" with AI makes as much sense to me as using chess engines to try to "solve chess."
You're right: "prove all the math" does not make sense on any level, and nobody serious would phrase any of this in that way. I certainly didn't.
The issue is SNR: signal to noise ratio. Generating exponentially more mathematics, particularly if the process is indiscriminate or optimized for something other than usefulness or mathematical relevance (such as optimizing for machine-provability), does not imply that we get exponentially more applications. We may end up halting the progress of applications altogether as the entire capacity of the world's mathematical apparatus is consumed by the interpretation and investigation of machine-generated proofs.
You can already visit arXiv and find vast numbers of not-yet-published mathematical papers. Most should never be published. None of this junk is benefitting humanity in the slightest.
I'm deeply suspicious. I do not yet have a concise statement for why, but a lot of literature on the sociology of knowledge work sort of points in this direction. Section 5 of the Thurston article cited by Tao touches the elephant. Raduchel's article on the economics of software [2] also touches it.
I've tried to put words to this for a few years. I think I'm just going to start writing versions of it as see if that helps me shape the thought into something more concise.
So, in the spirit of this article's style, here are some postulates:
1. There is a sociological process happening in the production function during knowledge work.
2. That production function and the associated sociological process spans years or even decades, and must outlast many of the artifacts that are produced during the early years of the function.
3. You cannot get the right lines of code or the right theorems proved without running that sociological process alongside the artifact production process.
4. It is impossible to completely separate the sociological process from the artifact construction process. If you just iterate on artifacts then too much of the required hidden state is lost to make progress in the right direction.
5. So you need that sociological process, or something like it, to still happen.
6. For a lot of knowledge work that process plays out in extremely high-fidelity social interactions [3] that we have not yet captured the datasets that would be required to reproduce those dynamics.
7. And even if we do collect that data, our current architectures and training algorithms and hardware would be useless given the size of the datasets.
[1] https://arxiv.org/pdf/math/9404236 Section 5.
[2] https://www.nationalacademies.org/read/11587/chapter/11 pp 166-168.
[3] there is a reason we still gather in-person around white boards, and why doing so is more crucial for some types of work than others.
I would use a different adjective: exceptional.
There are not not times where lone geniuses produce amazing output. But most of humanity's progress over the last few thousand years (or at least certainly the last few hundred) resulting from a different type of work.
Referring to my list of examples as exceptional in the sense of being rare is rather unfair, given how numerous these examples are relative to the body of mathematical work we would consider incredible. The fact that such a large % of that body of work occurred while the individual was in relative isolation is something we should pay attention to.
Heh. Your first premise cuts in the wrong direction. I don't think I've ever met a scientist who didn't complain about meetings :)
It would work better as a bar for hiring, rather than as a bar for publishing.
https://terrytao.wordpress.com/2026/08/18/palomar-a-registry...
For many of the rest of us, mere consumers of mathematical results, it’s sufficient to know that a^2 + b^2 = c^2 was proven by somebody or some machine at some point.
He's a typical person otherwise, politically aware of how he barters for food; until proven otherwise this can be seen as little more than social moat defense.
To paraphrase a quote attributed to Upton Sinclair; hard to get a worker to understand something when their paycheck relies on them not understanding it.
The only interesting thing here is the frogs high up admitting they feel the heat.
It should be ignored and refused.
I am probably being too optimistic, but wouldn't it solve the problem if peer-review had a pre-screening phase where you give a presentation about your work? Similarly to how a PhD presentation is given. It could give back the publishing power to the expert, rather than the journals.
Once you have validated that the knowledge you want to publish is yours and that you actually understand and own the work, then it doesn't matter if the paper is written by a LLM or if the LLM assisted you in doing the work.
Tao is essentially saying that the only value in a proof is its ability to be understood, but that's wrong. A proof is also valuable because it establishes a new fact. The fact is useful in itself, even absent an explanation.
From a different angle, what we don't understand can absolutely hurt us and you're right too, but it doesn't contradict Tao's viewpoint.
> Let's say tomorrow someone comes up with a formally verified proof that a major encryption algorithm underpinning the security of the internet can be trivially broken, but they can't explain it. You're saying it should be kept under wraps and not published?
Absolutely. It could also be exploiting bugs in the verifier. Even if not -- even if that proof were correct and entirely written by humans, care should still be taken in how such knowledge is published. I'd want to give trusted parties a chance to try to fix the issue before letting it be known by black-hats, for instance.
Understanding was critical for the field to progress when only humans were involved but if humans are not needed to make progress, I wonder if we split into two worlds: an AI math-world where amazing new results continue at a rapid pace bottlenecked only by compute/cost and a human math-world where we understand a subset of the AI math-world as a hobby (similar to Stockfish vs human chess).
I am wary of AI in all aspects I am seeing it in but in many ways in mathematics seems to me the least troubling. It will change things in and the field will not be the same. Blacksmithing has not really gone away. You can still work as a farrier, if you like that sort of things. The tools that replaced a man working over a forge with a big hammer are part of a giant industry that is still producing works for the modern world.
[1]: https://bbchallenge.org/8226493
The Busy Beaver game has lead to a better understanding of complexity theory and automata. Also, direct "hands on" work on improving proof assistants and related tools.
Btw, for those who are curious, the Busy Beaver Challenge wiki is a treasure trove of rabbit holes and curiosities:
https://wiki.bbchallenge.org/wiki/Main_Page
That is what all marketing wants you to think...
(I’m not sure the extent to which you think this is marketing. I’m operating under the assumption that you agree with his Hypothesis 4.1. If you don’t, then I’d assume you haven’t seen the long list of prominent open math problems that these systems are providing answers to. And if this doesn’t sway you about Hypothesis 4.1 I’d just halt and ask why.)
What does it say?
(Yes, pure math research is useless. Applied math is very useful, but he is doing pure math, which is very useless.)
And the term "artificial intelligence (AI)" has been the name of the field for 70 years and counting. If anything, "LLM" is a misnomer that's been lingering around since 2018-19. When the term was coined, these systems were relatively small, experimental, and could only produce impractical facsimiles of the English language. This is obviously no longer the case today.
It's been used to talk about computers playing chess, then machine learning, and now LLM-based systems.
And oh, what a stride it is: https://vibemathed.com/stats
Advancement in capability does not mean the mechanism is the different. The LLM name denotes a very specific mechanism..
No, not really. This is just the term that stuck around. The "large" is now up to five orders of magnitude larger and "language model" has gone far beyond any simple notion of modeling a singular natural language. And anything you'd cite about transformers, or tokens, or autoregression, etc., is more of a factoid about what works best and happens to be the most convenient in the here and now. I see all of this as an unbroken continuation of work that's been going on since the 1940s.
Instead of trying to play word games, why can't you just read Tao's article?
Does not matter. It is still an LLM.
And I am not the one who is playing word games. You and your idols are, for sake of marketing.
Cool. So now you can accept that "LLMs" are an obvious example of AI.
>You and your idols are, for sake of marketing.
Let's be very clear here. Terence Tao is arguably the greatest mathematician alive. Yet, you are throwing lazy insults and accusations around because you don't like the term "AI". And that's my final comment for you, troll.
Not sure what this has to do with what I said. A lot of things have been called "AI" in various times. None of them including the current crop of LLMs are not really AI. But people use AI term loosly and that is fine. But it is a problem when a some thought leader does it.
>troll.
Tell me you have run out of arguments without saying you have run out of arguments...
There is an LLM acting as a component in a larger system. But that larger system is not an LLM. Calling it an "AI" is indeed an act of marketing as there is no learning / adjustment as we would expect from an intelligence, but calling it an LLM seems to be inaccurate. So what is it?
Again, this is literally the name of the field (and the tech). It's been around longer than you and probably your parents.
Dartmouth workshop (1956):
https://en.wikipedia.org/wiki/Dartmouth_workshop
Random dusty undergrad textbook from the early 70s:
https://m.media-amazon.com/images/I/816fxHVJkHL._SL1500_.jpg
>there is no learning...
If human-like continual learning is suddenly the standard, you can just as easily say that terms like machine learning and deep learning are an "act of marketing".
"every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it."
Clue: If you put a pig in a poke, is it still a pig?
It's the old cliche of "if you only have a hammer every problem looks like a nail". Let's not fall into the trap of thinking that our life needs to be 100% about AI or completely devoid of AI. We can really use this thing to make our lives better.
Instead of wasting time on the question of whether we should use it, let's focus on HOW we'll use it.
The only thing to do is to be all in, or get run over.
It is going to be a tiny minority of people who "win". The other 99.5% of humans will just be losers who, what, die in the streets? This argument reminds me of Roko's Basilisk. "The humanity-ending torture-loving Basilisk is coming whether you like it or not, so come over here and help me build it!"
If Amazon uses AI math to come up with better routing, the cats can benefit from cheaper delivery fees just as much as humans can. No understanding needed.
The human brain is being obsoleted, soon thinking is going to be a recreational activity like weightlifting.
Maybe we'll have some hobbyist dabblers, but any real progress will be done by machines that skip the human.
That doesn't mean they don't provide any value of any kind to anyone.
Being rich doesn't mean you can do whatever you want. But it means you can choose not to do what you don't want.
Everyone can be rich. But not everyone can be Montecito-rich.
https://arxiv.org/abs/math/9404236
He wrote it in 1994.
He writes about how he almost "destroyed" a subdiscipline in mathematics by becoming so good at it that he outclassed everyone. PhD students were advised to stay away from the whole field.
When he discovered this, he realized his error was that he was focusing on producing results, and not focusing on explaining his thought process. It's that thought process that is valuable in advancing the frontier - results alone won't do it. It didn't matter how many theorems he proved, if he was the only one who had the mental framework in mind on how to think about the whole field.
I'm sure we've come across abstruse books where every theorem has a rabbit being pulled out of a hat, whereas other readers find it intuitive. It's because the latter has developed a mental model for the discipline, and you haven't.
So he set about slowing down, and focusing on holding lots of seminars where he worked with other mathematicians to explain the thought process. Eventually others started publishing proofs of key theorems.
When people publish in a journal, they are not merely doing it to show the result. They are having a conversation with other mathematicians. If they cannot explain their own proof, they're not having a conversation.
This is why even decades after the Four Color Theorem was proved, plenty of mathematicians don't consider it "mathematics".
This is hilarious lmao
Most research mathematics is pure mathematics which is completely useless. No routing algorithms. It's only relevant because we (or at least mathematicians) are interested in it. So an AI producing incomprehensible proofs would be completely pointless. That's why Tao insists on the importance of human understanding.
A few different reasons why use an LLM when mathematics is done for its own sake:
Formally verifying my proofs catches any mistakes I make, but verifying is also hard work. LLMs shaves off a lot of time when formally verifying a proof.
I can still read through an LLM generated proof and understand it. This is a way for me to understand the result I am working on (usually in order to know what to prove next, results are not proven in a vacuum).
My experience thus far is that, while correct, an LLM generated proof is often unnecessarily complicated or inelegant. I take pleasure in elegant proofs and will spend time iterating on the first proof until I find it conveys the idea in the most elegant way. Having the initial LLM proof to start with is really useful, but is thus far rarely the final product.
What happens when it's the cats who get to decide what's published?
Not an ideal scenario, but that's exactly the situation here. Mathematicians decide what gets reviewed and published in a a top journal.
In the long run, this can and should make journals obsolete.
If a subset of mathematicians, use AI to condense timelines focusing on goal 6.2 exclusively and make rapid progress and reach a proverbial inflection point — one where value proposition of the using this new normal is too enticing to give up — everyone will ask: "This thing is so awesome. Why should I care about your values?"
Ahh. I know Terrence Tao didn't have the following in mind when he said those words, but oh man, the philosophy neurons are firing in my brain rn.