> I propose the following reconceptualization of the goal of a mathematics PhD: to become a world expert on some interesting, deep topic, and to be able to convey that interest and understanding to others. Part of operationalizing this might be a thesis, but the degree would be awarded primarily on the basis of a rigorous defense, in which the student explains the topic to their examiners until they are satisfied.
I think this is a refreshingly forward looking idea and I agree with it 100%, especially the the "rigorous defense" part. That is a good measure of how well the topic has been researched and understood by the researcher. This is where the humans can be "in the loop".
> How different would this look from current PhDs? I think students would still meet with an advisor, who might suggest a topic. That topic could be explored with AI assistance, or not, but the student would be responsible for understanding it; it might be much more open-ended and larger than the typical PhD is currently.
Interesting point about "more open-ended" and "...larger than the typical PhD". I think the author has a point. Earlier, the bottleneck was the candidate's/researcher's understanding and knowledge. Now with AI tools, it is so much easier to zero in to relevant knowledge, get your questions answered quickly which might lead to understanding more quickly.
For e.g., before the advent of public libraries and printing press, the knowledge was inaccessible and guarded. So that was the bottleneck.
Then books became ubiquitous and the bottleneck to knowledge and understanding was people's motivation AND knowledge of WHAT books and topics to research.
Then came the internet and free PDFs of books and research articles. Now, the bottleneck was still people's motivation and a mild version of what books and topics to research. I say "mild" because one can lookup articles and newsletters, and book reviews and come up with a list of reading.
Now comes AI and it looks like the only bottleneck is people's motivation.
I believe there was also a silent, yet potent, bottleneck all along which is also removed by AI: personal tutor/coach/teacher/professor etc. Let's say if I am reading a textbook on manifolds or some research paper and I have a question about a specific theorem or even a mathematical operator being used. Before AI my only way to get my questions answered was to read more books (PDFs or print), or ask on math exchange or math overflow and wait for someone to answer, or to ask a professor. This could take up to a week.
Now all of that has been cut down to 1 hour or less with an interactive chatting session.
!!!!!
So....the only bottleneck is people's motivation! QED
I'm not sure why an advanced degree is necessary for that at all, besides the pride of a vanity title. It's pretty much what Bill Nye does for science.
The idea that most any modern "interesting" aspect of mathematics is going to be understood (or often even explained in enough detail to reveal what is interesting) in an hour is pretty rare. Either the student's aptitude, the tutorial, or the mathematics are unique.
I think we often delude ourselves as to how well we understand problems and their solutions. Some instructors even make you feel that you understand better than you do by pointing to a few approximations or simple solution spaces that obscure the larger complexity. Just looking in wonder at the many categories of three-body solutions (currently on hnews) is enough to remind me of this.
Another bottleneck is money required to pay corporations who own the technology to "do mathematics". In that future, there will never be another Ramajunan.
I am not implying that this is a new idea and I don't believe the author did either. If you read the full article, the context is clear.
The author is re-emphasizing the importance of thesis defense old school style and highlighting that being familiar with one's own material rigorously should still be a requirement.
Everything that leads up to the thesis defense perhaps changes signficantly.
The more I see these posts about mathematics institutions reforms and challenges from AI advancements the more it looks like they may need to go through a death. Or to put it another way they may need to start again from first principles.
If math is truly about spreading intuition and understanding then our institutions have dropped the ball decades ago and have not been able to grab hold of it since (if they ever had it to begin with)
Not sure why you're down the bottom when the current top post says pretty much the same thing. I agree that the reevaluation and refocusing that is being forced by AI is one the maths establishment could fruitfully have had a long time ago.
Excellent optimistic post in a sea of negativity, and with actual suggestions, too.
After reading, my mental image is this: think of Olympiads in Ancient Greece.
* A weightlifter was only awarded a laureate if he were able to lift a heavy stone (have no idea what they were lifting, for illustrative purposes only :-)
* Along comes Archimedes who invents what we would call an exoskeleton. Now any regular guy can lift twice as much as last year’s athlete.
* What to do? You can cancel the Olympiads, but they are actually useful as training, motivation, etc So now you have to give the prize on other factors, eg how well he can lift, has he opened a gym in the city, etc
BTW, physics and bio are not exempt, so those researchers better read and try to stay ahead.
I don´t think so. For programming agents can run code, check compiler output, etc. For mathematics, it is almost the same once you factor in the usage of lean.
For the reality, you can´t close the loop that fast, or with that precision. You will have to slow down by several orders of magnitude.
Doesn’t have to. There are petabytes of experimental physics data that can be fed to AI to extract additional insights. The only holdup is this is slightly harder to than math.
With bio, you’re right, generally designing and conducting an experiment goes hand in hand and theoretical biologist is not a common label.
I am not a mathematician, but I can’t see how we are going to address the problem which we already see in coding:
Impossibility to independently validate all AI results
And in math it goes even worse. In coding code reviews are typically still the form of action you do within days. In math, historically, the lifecycle of proof is months if not years. Take as an example Millennium problems. They require at least two years of validity after publishing. Two years! In modern times with amount of output AI can produce, it feels like infinity.
We are inches close if not at the moment already when humans can’t reliable validate proofs and mathematics produced by AI. Then next research will be based on this AI-written-no-human-in-the-loop results. And we will end up in just few years in a world where novel and frontier problems will be articulated by AI and proven by AI based on AI results and humans will be incapable of understating the mere nature of the solution.
I don't think that's actually the real problem. Along with the progress in answering mathematical questions, recent progress on AI-powered autoformalisation has been astonishing. All the recent AI discoveries have been accompanied by Lean proofs.
And, yes: that doesn't absolutely guarantee correctness. The Lean kernel has had soundness bugs, and may have some still. But it's pretty strong evidence of correctness nevertheless.
The concern among mathematicians is not mainly that they doubt the correctness of any of these discoveries, but that human understanding may be devalued.
I am not that worried, but rather just observing. Humanity is about to enter the phase when we will be using things based on ideas no human ever properly understands. This thought … disturbing, somehow?
It is perfectly valid counterpoint to say that we already do it. We everyday use myriad of things, tools, and software we have 0 clue how it operates. But for us as humans it was reassuring that we know that at least there are a few other alive humans who know it, who create it and who can explain it.
>Humanity is about to enter the phase when we will be using things based on ideas no human ever properly understands. This thought … disturbing, somehow?
This is just normal though. We were building sophisticated bronze and steel tools long before any complex understanding of metallurgy or chemistry. Medicine is still the wild west.
I first wanted to say fire, even though it's a cliche, but then I thought that in antiquity we used like everything without anything that would qualify today as understanding. Also now we have a lot of stuff that we "know" it works based on complicated numerical simulation.
I think the most "understanding" we ever had was in the 40s-50s designing nuclear bombs with slide rules. It was the culture that produced the idea of psychohistory.
> Medicine is still the wild west.
Reminder that we have no idea how anesthesia works.
But one of the things AI also excels in is summarizing and can do so hierarchically. One of my favorite things to do with a concept I'm new at is "ELI5" then "explain like I'm a high school student" then "explain like a bright undergrad in XYZ" then "explain to a working professional in this domain". It's a lot of steps, but I've found it very effective (for me) to learn with -- and I've done something similar with code and math (although not math proofs -- I'm not a mathematician). But my point is that I think we can use AI to also teach us these proofs they're building in a way that I don't understand today about human proofs.
If we use AI well here we could actually understand math much better than we do now.
With formalized math, you only need to validate the problem statement (in theory, in practice agents have already managed to exploit Lean compiler bugs, but the incidence of those should decrease enough to be practically lusable for 'blind' validation of AI proofs in the foreseeable future).
I actually think it's the opposite: Lean proofs and autoformalization make it very easy to announce proofs alongside proofs of the correctness of those proofs (Lean certificates). It's not an absolutely fool-proof combination (the Lean kernel could still contain bugs), but it does immediately attach a very substantial degree of credibility to the result.
And that I think is essential to why some of the world's leading mathematicians are taking this so hard. In a world where we "merely" have AI systems capable of superhuman informal reasoning, verification, correctness, and acceptance could still only be conferred or anointed by human mathematicians. But a world that combines superhuman informal reasoning with superhuman autoformalization is a fundamental shakeup in the institutional order.
The recent proof of Fermats Last Theorem is interesting: it is (iirc) 13 million lines of lean code. And type-checking takes 5 hours or so on a pretty beefy machine. I cannot independently verify the proof, and I have to take Anthropics word for it that it actually type-checks.
That seems like a red herring. Have you independently verified the human generated proof of FLT? Surely someone else will try to verify Anthropic's formalization on different hardware. Plus, it seems likely that FLT formalizations will improve / get shorter over time, requiring less compute. And computers (and type-checkers) will continue to get faster over time as well. So maybe in 5 years you could own a computer fast enough to verify a/the proof in say a week, instead of 5 hours.
> I think so. This machine might produce answers we value, but it would not, in itself, produce human understanding of those answers.
It's nice that the author is optimistic, but won't the AI be best placed to dumb down its increasingly complex proofs into a language us lowly humans can understand? To keep thinking until it can refactor complex proofs into ones from 'the book'?
As they go on to explain, a human understandable proof is different than a human actually understanding the proof. That actual human understanding (like, in a brain of a human) is one of their stated goals.
Producing human understandable proofs is possibly a job best for humans today, but the author appears to agree with you that this is probably fleeting (and argues that even if you disagree, it should probably be treated as if it is fleeting when planning for the future):
> Right now AI systems arguably underperform us at theory-building, asking questions, exposition, … so we could prioritize and reward those skills. I think this is unwise: compare the speed at which the academy adapts to the speed at which model capabilities improve. We need to consider the endgame. If the models remain incapable in some domain, we can adjust later.
I love how this is (without slighting the problems entailed) coming at it from a perspective of infinitude and abundance (we will always have more problems to solve) - which is the correct framing, particularly when dealing with ideas, or fields in the which ideas are the driver/product/output/material, and ideas themselves, the field itself, are infinite.-
As someone who has a degree in math, I still can't help but think mathematicians are getting a little bit of a comeuppance. In a lot of areas of mathematics there had been little effort to make the work understandable and leaves numerous folks who could benefit from the knowledge on the outside looking in. Now AI comes along and do the same to mathematicians. Makes me chuckle a little bit.
It is not lowering the bar to find better ways to demystify and explain things. In fact, I would say those who can explain it well understand it the best. Richard Feynman would be my best example.
Do you have any examples of places you felt like there was a lot of gatekeeping? Perhaps having studied mathematics I am a bit blind to the issue here and would like to learn more.
I agree making simple things sound complicated to appear more impressive is bad but there are limits. Even with Feynman he could only go so far, e.g. his interview about why questions and magnetism.
I am not thinking of intentional gatekeeping, but more of the kind where mathematicians build themselves their own island of concepts and notation with no thought of building a passageway for others to engage and make use of the theory.
I would find modern algebraic geometry highly useful as someone who works in computer graphics and computer vision, but much of the theory is akin to learning a new language and I don't get the sense that those who publish their work in this field care if I enter their world.
I do think in a lot of fields there is a lot of impact/influence to be had by people who are willing to do work to bridge different fields. I wonder to what degree this is because practitioners don't always see how their work could be used elsewhere.
I would say that it is lowering the bar to "find better ways to demystify and explain things." But there is no harm to this bar falling, because this is a bar for entry. There is a separate bar for making research discoveries (being credited for them, specifically), and that one should not be lowered, because it is a bar of standards.
This feels backwards. I frequently joke "physics is the subset of mathematics that reflects the observable world". Math can be as abstract as it wants, but physics has a constraint. It must model an observable world (related, this us part of why people say String Theory is math and not physics)
I think you're a little naive if you think this is because of gate-keeping by mathematicians rather than the essential complexity of mathematics. Mathematicians individually and as a group would love nothing more than a world which is capable of understanding their work more deeply.
In the exact same way that LLMs allow anyone to vibe code an app but do not replace real understanding of system design due to its essential complexity, non-mathematicians will quickly learn that asking an LLM to pump out advanced mathematical statements to you, even if they are correct (and even if you could verify them) does not constitute understanding, and that the human brain is the bottleneck either way.
It is only if the LLM is super-human at simplification and explaining that a difference will be noted. This would be excellent for mathematics but its not a foregone conclusion (and the argument of most mathematicians, such as Terence Tao, is that this distillation process is one of the key parts of doing mathematics, and that LLMs so far seem to be going in the opposite direction. I suspect its probably user error and leveraging the tools better will produce different outcomes, but mathematicians are only just starting the journey that software developers have been going through, so patience is needed).
Folks keep using the word "gate keeping" -- I never used that term nor implied intentional gate keeping. It's simply not caring if those outside the club understand -- there is no one guarding the door its just folks don't care if anyone finds it.
Gatekeeping is a human universal. Even those who preach maximum inclusion do exclude plenty, of course there is always explanation why that doesn't count. If you try to run any community you quickly learn the importance of gatekeeping. Eternal September etc. See any site or fandom that gets mainstream etc. I don't think anyone has any duty to actively pull in as many as possible new people. Zen masters used to chase away students or make them sit on their doorstep for days and send them away anyway. Now, the west is essentially Christian and so the missionary impulse is real, but it doesn't work without the other stuff in the package.
I too have a degree in math and it's easy to see some of what's happening as too much of an emphasis on intellectual chest thumping and to little emphasis on showing why understanding the concepts is useful for everyone. I will say the problem may well increased here by the American education system which denigrates such understanding and isn't influenced by mathematicians.
But also, with things rapidly changing, perhaps in ten years chat programs will not only do superhuman but make their proofs marvelously accessible and provide incredible tutoring sufficient to bring any curious up to a super high level quickly. Then what can you say and what can you complain of.
But I think "do everything machines" are necessarily inevitable but the situation does make it uncertain where the limits are.
I would say the research university is getting comeuppance.
Pedagogy is the primary purpose of educational institutions. That includes universities. However, when research is placed first, you are often left with mediocre teachers, because "who cares?" They were hired to do research and lend the university that kind of prestige; teaching is an afterthought squeezed into the spaces that remain.
And frankly, we haven't the faintest clue today what education is even for. That doesn't bode well.
Add to this the other troubles facing the university, like declining attendance, skyrocketing costs, and AI cheating and we're in for a very interesting transition indeed.
Why should we expect everyone to be both a great researcher and communicator? The obvious result is that it is as effective as engineering managers. Sure, there's some amazing ones, but most aren't. Though that also doesn't mean no expertise in the field (i.e. non-engineering manager) is any better. It is just that managing/communicating is a different and orthogonal skill.
What needs to happen is we need to make it okay for people to specialize in more things. More nuance to this rather than trying to throw everyone into nice easy to manage buckets. Those buckets are just unrealistic abstractions filled with hope, denial, and laziness. Reality is surprisingly complex. Math can do a really good job helping you understand that, but it's a sufficient condition, not a necessary one
Never been my experience of maths, which is one of the most accessible disciplines to have ever existed.
I also have a degree, but mostly taught myself the important undergrad-level concepts as a kid. I just went to the library and borrowed any of the hundreds of books written to clearly communicate maths to beginners or downloaded any of the free ebooks / lecture notes.
Name a single field of endeavor that is more open in 2026. Software certainly isn't one of them - the best stuff has always been gatekept.
> resulting in the production of an abundance of PDFs. The contents of some of those PDFs may even have important applications.
I hope, from the depths of my soul, that the static typeset report format for transmitting knowledge and understanding will finally die and be laid to rest.
As a researcher in theoretical computer science, I love PDFs more than any other format when it comes to mathematics.
There simply is no contender to LaTeX and PDFs.
Lucky for you, almost all research in math, cs, and physics, are put on arxiv, where you can download the source code (.tex) as well as get an HTML render.
The author argues for evaluating Ph.D. candidates based more on the oral thesis defense than on the actual thesis.
By essentially the same reasoning, I’ve been arguing for prioritizing in-person design/code reviews over code-only async PR comments.
The important thing is to verify that the human has a coherent design in mind and can demonstrate that it got implemented, regardless of who or what was at the keyboard. “I dunno, I guess Claude thought this was a good idea” is not a coherent design.
I think upweighting the live components of academia is inevitable in the age of automatically produced writing, but I also find it depressing that people think so little of writing that they imagine it obsolete because of AI.
AI writing is garbage. If you're indifferent to how much better good human writing is than AI writing, you are not qualified to evaluate writing.
Person B: “key pieces are delayed due to the developer not understanding all of the LLM doesn’t and implementation.”
Manager: “does it work? What are the risks?”
Person B: “well yes it works for now but we’re accumulating tech debt due to a lack of understanding and potential flaws that haven’t been thought out yet”
Manager: “they want feature X, ship it, we can deal with it later, I don’t care if it’s not coherent as long as it works.”
How many decades at this point has these been a push for functionality over everything at all costs? And you have a mechanical snow plow now. Most businesses don’t care about later risk or any future planning beyond the quartet horizon, they’re not concerned about how it will effect their performance in 3 quarters or lead to instability or issues, those are future problems for a future person and we’re here for money now.
The simulated dialog you have provided is very much representative of what many of us have heard first-hand. However, I disagree with:
Most businesses don’t care about later risk or any future
planning beyond the quarter horizon, they’re not concerned
about how it will effect their performance in 3 quarters or
lead to instability or issues, those are future problems
for a future person and we’re here for money now.
Businesses care about "later risk" and what it implies.
Individuals within an organization do not unless it will specifically affect their bonuses/promotions.
An oral defense is ultimately a meeting, with the same weaknesses as other types of meetings. In particular, if new information comes up in a meeting, you cannot reasonably expect to get useful responses from the other participants. First impressions maybe, and plausible-sounding bullshit from those prone to generating that. If you want anything more, you need to provide the information in advance or schedule another meeting later.
Oral defenses in academia are largely rituals. If a student fails their defense, it's almost always the supervisor's fault. The supervisor is supposed to be the primary quality control. With their regular meetings with the student over the years, they should be able to tell whether the student has achieved sufficient understanding and contributed enough to graduate. If the supervisor thinks the student is ready to defend, the reasonable expectation is that the student will pass and graduate. The defense is mostly there to let the other examiners validate the supervisor's judgment.
Here's another example of a similar but slightly different failure mode of the live defense: The department where I got my degree had a practice around PhD defense where there would be a prof each representing the four research pillars + the dean. It was well known that certain professors disliked each other enough that they would snipe each other's candidates. You'd just hope not to work for one whose enemies got picked. If you were a superstar, you'd pass, but if you were an okay candidate or someone strong but with stage fright, you be could toast even if they did good research.
The defense was only briefly about the actual thesis, then switched over to whatever research interest the committee members had, and they'd drill into their pet subjects. This was in physics, hardly politics forward normally, and at a very well respected university.
You are missing the entire point of the oral defense while almost stating it. The existence of the oral defense is what extrinsically motivates the supervisor to do their job well. If the student fails, the supervisor loses faces in front of their peers, which obviously they try to avoid. Removing the oral defense would massively reduce the quality of PhDs because supervisors will not do their job.
Also in my observational experience in physics, a fair number of students (probably like 10%) get some sort of major corrections to their thesis. After writing this, I did a search, and in the UK, the number across all fields is 16% [1].
And even then, the other examiner will only fail the candidate under egregiously extreme conditions. Rather, if the supervisor wants to pass someone who is unfit but not egregiously, this will just erode the supervisor's reputation among peers. Because there is a grapevine of course. So it can't be regular. But academics also really care about face and embarrassment, so they won't make a scene, u less there is preexisting drama around the supervisor. The only fails during the oral I've heard of happened when the candidate insisted to defend against their advisor's wish.
I do this already if there’s any kind of issue. Just talk the person if they’re next to you or zoom them if not (don’t set up a meeting). A five minute conversation can save hours of back and forth.
The tricky part is that live review is much more expensive than async review. But maybe that just means we were getting away with a cheap proxy because it happened to work reasonably well.
>“I dunno, I guess Claude thought this was a good idea” is not a coherent design.
I am sorry, but the writing on the wall is that this is where software industry wants to go. As a software developer, I cannot but help notice how unrealistic this is.
We are entering the next level of bloated, inefficient and unreliable software....
1. It will take much longer to understand the output of the machine that it takes to prompt and create it.
2. The only? best? one? way to /verify/ that you /in fact/ understand the output of the machine is to explain it to someone else.
So there will be a machine generating koans which need to be meditated upon and discussed with human social back-pressure validating understanding. I think this could be much more cooperative and at a minimum this will be a way different math social construct.
I like the quote from Hilbert that was brought up in the article: "we must know, we will know".
With AI, it might be the case that we don't know, we won't know, but the machine does.
The central question, namely whether humans should be in the loop, will be repeated again and again in the years to come for all industries, starting with mathematics.
Better plan is to shut down PhD and make students take a oral thesis at bachelor and master level and help them become a productive economic participant as soon as possible.
People had the same complaints that the code produced by early coding models was messy, lazy, poorly commented, and so forth. The central complaint was that it was just too difficult for humans to review. The answer is just to improve the models and move on.
Similarly now we're getting AI doing math. The math is a giant vibe-coded ball of wax. So just make the models better at explaining what they're doing to humans - that's the end of it.
Rather than just go on and on about how it's the end of the world if we don't do this, why don't we just do it?
There's a fundamental difference between the goal of code and math. The goal of code is to produce software that does something useful. As long as the code does what it's supposed to do, arguably, it's good to ship. (As you imply, we want the code to be good enough to also be reasonably certain there are not too many bugs, that it is maintainable and can be extended etc. This is what early models failed at but now seems broadly fine.)
But for math: what is the point of a proof if no one will read it and no one uses its result? To quote the article, "AI systems will [...] result in the production of an abundance of PDFs. The contents of some of those PDFs may even have important applications." But if there's no one reading the PDFs, what's the point - no matter how good your AI model.
The point of math is understanding. So mathematicians should feel free to use AI as much as you want, but in the end, they should've gained some understanding on what happened.
If the point of math is understanding then maybe our incentive structure is wrong. Maybe it should be teaching the concepts to as many people as possible rather than just continuing to write papers that 10 people in the world understand, which is the current state of a lot of math.
The primary goal of all basic sciences is human understanding. "Truth" is no more a goal for mathematicians than the physical laws are a goal to physicists; they simply exist in nature. The goal is rather to develop useful language and conceptual frameworks for reasoning and communicating, which also underpin all practical applications.
Sciences don't have goals, people have goals and they differ. Some are fans of pure math as a kind of religious or almost erotic activity in elegance and beauty, others are application minded. Some are in it for the community and outreach and conferences, some are in it to just sit in an office alone and be left alone to do it in a zen like flow state all day and night. Some treat it as a 9-5 to pay the bills with a skill they happen to be fit for but aren't especially passionate about.
I won't understand it, but can benefit from it in better data structures with proven invariants, faster algorithms, convergence guarantees for some iterative computations, better practical linear algebra and matrix factorizations, deriving things in statistical hypothesis testing, tightening upper and lower bounds etc. I don't care if some mathematician somewhere who isn't me "understands" it or not. It has practical use. I know that practical is dirty peasant word for many ivory tower mathematicians but that's their problem and I don't interact with them much, except when they throw a hissy fit like this and try to ban restrict matchmaking to their guild and ban the plebs from getting math from anywhere but them on their terms.
One funny thing is that in order to tune the models to make what they're doing explainable to humans, you need to have humans involved in the RL pipeline to indicate which explanations are good.
You can understand this as learning a mapping between the model's internal "world" (i.e., 'meaning,' which is hopefully coherent and consistent -- but definitely not always! see, e.g., https://arxiv.org/html/2505.11581v1) and language (i.e. 'form') that reflects that world.
For this to work, you need both coherent / consistent internal model worlds, and also good mappings onto human language. Supervision by mathematicians has provided the signal for both internal coherence (though this can also come from interacting with a proof oracle) and for good explanations. If models exceed human capacities, you could imagine that aligning their explanations potentially becomes harder (though not necessarily). Also, humans naturally have to do the same thing: as researchers we must find analogies to make our work legible to collaborators or laypeople. Often in doing this, we further clarify our own understanding!
More deeply I think the "end of the world" vibe arises not only from the practical need to have models that explain, but also Litt's (and many other fields' researchers) grappling with being relegating to not mattering.
I think article is more about the goal than the how. Perhaps it will be better to have ai produce human accessible explanations. No one is saying not to do what you propose. The goal remains the same.
I think it's pretty analogous to coding, honestly. When it comes to gaining true understanding of a piece of software, I still find today's models almost as useless as they were a year ago.
They can write great code, they can gain understanding of something for themselves, but they still kinda suck at explaining it to people. As the article says, if you want understanding, it seems like there is no replacement for getting in the weeds yourself. A model can help, but it's not going to magically download knowledge into your brain.
> Similarly now we're getting AI doing math. The proofs compile but are a mess. So just make the models better at writing clean proofs and explaining what they're doing to humans. That's the end of it.
The assumption here is that the true/false of the theorem is the important outcome. While it is certainly part of it, a big part of maths is the understanding you gain from a proof. Many of the best proofs elegantly explain some aspect of the maths which was previously unclear and expand our understanding of the world.
To use a programming related example, imagine that an LLM spits out a solution to the travelling salesman problem which works in O(n) time. On the one hand that's very convenient for whatever problem you happen to be trying to solve at the time...but there's also an answer to P=NP in there! The former means your delivery drivers app works a bit faster on their busy days, the latter fundamentally shifts how humanity thinks about certain problems.
Going back to the maths, there have been theorems that were proved (by people) where the proof is broadly seen as 'unsatisfactory' in that it doesn't really expand our understanding. I assume some of these LLM proofs are a bit like that: we now know that the thing is true, but we really want to know why it's true, and how that changes our understanding.
> We already interview faculty hires; we must now do the same for graduate admissions
They hire PhD students without hearing them give a talk and then doing interviews? In Germany, the applicant gives a talk (30-40 min) to the research group they want to join, usually presenting their master's thesis, engage in discussion, often share lunch with the group, then do 1 on 1s with individual members of the group and a longer one with the PI. Obviously this can vary within Germany too, but I couldn't imagine hiring someone without something like this.
1. US a lot of international PhD applicants, so traveling before even being accepted is difficult, and 2. lots of people don't do a masters.
when I applied to PhD programs (not in math) it was basically CV + personal statement + recommendation letters + short chats with interested faculty :shrug: Maybe it was because my CV was "strong" but the chats were more see if interests were aligned, rather than actually interviewing me.
Unlike Europe, in the US, most PhD applicants do not have a masters. 10 years ago, most had not even conducted substantial senior projects (e.g a semester or year). Though increasingly senior projects / undergraduate "theses" have become more common.
In the UK I had a panel interview with 4 people, in person. But for international students who often brought their own funding sponsored by their government that didn’t happen.
That's an interesting wrinkle. If anything, externally funded international students seem like exactly the case where you'd still want the interview, since the funding answers "who pays for this?" but not "does this person have the background and understanding to do the work?"
As with the rest of the domains, AI/LLMs will do syntax and search better than any human.
In code, any developer whose differentiaton was clean code and knowledge of different technologies is now average.
In math, any mathematitian whose differentiation was to manipulate formal systems and know tricks of different domains will be average.
Fortunately, humans do more than syntax and search.
The bad news for developers is that if you know what the output of your program should be (which happens most of the time), almost all of the job is syntax and search to build the code that reproduces the output.
The good news for mathematitians is that for the majority of problems you never know the output, or you just know the output is either "True" or "False". There are some cases where you need something else, for example "a solution that blows up in finite time". For those cases AI will outperform you easily (see new Navier-Stokes solution)
So if as a mathematitian you were doing more than syntax and search, then keep doing that and use AI just for what its best.
Another mathematician waxing poetically about a future that will not come to happen. Call me a pessimist if you'd like, but capital has no incentive to ensure mathematicians maintain their current status in society.
If you're a mathematician you are in the same boat as the software engineer, and the Dodo.
"A computer or monkey could easily start at the axioms of ZFC and iteratively apply deduction rules....simply conjecture all mathematical propositions in alphabetical order...The prospect of automating mathematics by enumerating all conjectures, and all proofs of ZFC, is probably not so disturbing to you."
I thought we were going to get at least some brief comment on Godel here?
Gödel effectively says ZFC must be incomplete, otherwise it would not be sound, but does that stop you from listing all mathematical propositions it can generate in some well-defined order?
I think this goes in the right direction. You have to rethink the role of human work in math, can't put your head in the sand and cling to your comfy institutions and system just because you got to know it's ins and outs and just want it to be like that forever.
But the bigger picture is: while I understand the author know his field best and wants to keep the post focused, the same issue will hit many more fields. We need to also have a broader discussion that involves more fields of knowledge work, largely academic scholarship but also regular office work, then it will come to engineering design, medicine, it's already coming for 3d modeling and vfx, software dev, it will come for a lot of middleman services. Not at the same rate, but we have to understand that it's not just that math will change.
Change will be the default. Everything will change. It will be much easier to make math change because all things will change. You shouldn't worry and imagine that funding criteria will be like today or that journals or academia or politicians expectations will be like today. No, everything will adjust with some timing differences of course but it won't be a static world and then math changing and having to justify and fight to explain the change to other actors who are baffled. They won't be baffled they will themselves have to change.
The world will transform as much as it did when society moved from feudal agrarian to urban capitalist industrial, or from the vast majority doing physical labor to a service economy with a huge amount of desk jobs. I can't tell how it will change exactly but it will be bigger than what we have seen in the last couple of generations or maybe more.
While I broadly agree with the premise of re-directing the “purpose” of math, I quite detest the idea that judgement might be primarily based upon some in person discussion, or
oral presentation, and the claim that written mathematics that is not orally communicated might be less worthwhile in some sense (i know this isn’t the exact statement of the authors proposition).
There are a good deal of people, whom, falter much more in oral discussions, whether this be for a psychological thing, stage fright, or difficulty explaining things on the spot. There are also certainly brilliant people, who can’t give an informative, discussion inviting talk to save their lives, but given enough time, can formalize their thoughts in writing at the highest levels of their field, and that writing is likewise very enlightening (sometimes).
It’s not clear to me, that, AI as is, could not pose successfully in an oral discussion of a topic. I mention this because it seems that one implication of the article is that AI might write things that are logically correct, but devoid of understanding. I suggest rather that 1) it is not extremely improbably that AI is incapable of generating mathematics that furthers human understanding and if 2) it is indeed highly likely that they cannot generate mathematics that furthers human understanding in a textual format, then surely one could also differentiate between human and AI on a textual level, and judge the contribution of a human, without the need of oral discussion?
I suppose another aside is, one might claim that the existence of AI means people have much much more text to filter for, and so, it becomes difficult to find one person’s good writing amidst a sea of, logically correct, yet understanding devoid textual content. But by and large much or mathematical academia certainly operates off of some reputation/vouching system presently anyways, that already serves as a “filter” in some sense. Perhaps the existence of such a system/culture is not a good thing, but oral discussions/seminars certainly aren’t immune from such predilections.
Perhaps I’m babbling like an idiot, but the entire and sole purpose of this comment is just to say: for the love of god please don’t let the standard be judged by oral presentation
What is the incentive for a person to sit though seminars and evaluations? People already hate redundant meetings. What is the incentive to change the system from the existing one to one that rewards this verification somehow? Who benefits from this change?
One of the things that struck me was that most people have an area of their workflow that they would be happy to hand off to an AI, so that they can focus on the thing they love. However, that area is different for each person. One person's grind-work is another person's love-work.
This is why it's so hard to build consensus around where the "red lines" are in this space.
115 comments
[ 2.0 ms ] story [ 40.6 ms ] threadI think this is a refreshingly forward looking idea and I agree with it 100%, especially the the "rigorous defense" part. That is a good measure of how well the topic has been researched and understood by the researcher. This is where the humans can be "in the loop".
> How different would this look from current PhDs? I think students would still meet with an advisor, who might suggest a topic. That topic could be explored with AI assistance, or not, but the student would be responsible for understanding it; it might be much more open-ended and larger than the typical PhD is currently.
Interesting point about "more open-ended" and "...larger than the typical PhD". I think the author has a point. Earlier, the bottleneck was the candidate's/researcher's understanding and knowledge. Now with AI tools, it is so much easier to zero in to relevant knowledge, get your questions answered quickly which might lead to understanding more quickly.
For e.g., before the advent of public libraries and printing press, the knowledge was inaccessible and guarded. So that was the bottleneck.
Then books became ubiquitous and the bottleneck to knowledge and understanding was people's motivation AND knowledge of WHAT books and topics to research.
Then came the internet and free PDFs of books and research articles. Now, the bottleneck was still people's motivation and a mild version of what books and topics to research. I say "mild" because one can lookup articles and newsletters, and book reviews and come up with a list of reading.
Now comes AI and it looks like the only bottleneck is people's motivation.
I believe there was also a silent, yet potent, bottleneck all along which is also removed by AI: personal tutor/coach/teacher/professor etc. Let's say if I am reading a textbook on manifolds or some research paper and I have a question about a specific theorem or even a mathematical operator being used. Before AI my only way to get my questions answered was to read more books (PDFs or print), or ask on math exchange or math overflow and wait for someone to answer, or to ask a professor. This could take up to a week.
Now all of that has been cut down to 1 hour or less with an interactive chatting session.
!!!!!
So....the only bottleneck is people's motivation! QED
Exciting time!
In my country, that's exactly how it is.
Yes, you need to have a thesis to defend, but ultimately it all comes down to the (oral and live) defense/disputation.
That's a radical departure from "PhD" being a certificate that someone is qualified to produce new research.
What you describe is more like a Masters Degree.
PhD has nothing to do with expertness.
If you have a PhD, you have completed some kind of research training.
That's all there is. Says nothing about knowledge or whether or not you're a genius.
You cannot conclude anything else, and nobody claims that you can.
If someone has a PhD, they have some training in doing research.
I think we often delude ourselves as to how well we understand problems and their solutions. Some instructors even make you feel that you understand better than you do by pointing to a few approximations or simple solution spaces that obscure the larger complexity. Just looking in wonder at the many categories of three-body solutions (currently on hnews) is enough to remind me of this.
Perhaps a parallel to this is what's happening with youth sports in the US. It is becoming increasingly inaccessible.
What?
"a rigorous defense, in which the student explains the topic to their examiners until they are satisfied."
is exactly how PhDs were awarded for hundreds of years.
Even my BSc in Applied Physics (1977) had a viva voce that was a substantial fraction of the final exam.
The author is re-emphasizing the importance of thesis defense old school style and highlighting that being familiar with one's own material rigorously should still be a requirement.
Everything that leads up to the thesis defense perhaps changes signficantly.
Who pays the salary of the mathematicians and the cost of accessing the AI?
If math is truly about spreading intuition and understanding then our institutions have dropped the ball decades ago and have not been able to grab hold of it since (if they ever had it to begin with)
* A weightlifter was only awarded a laureate if he were able to lift a heavy stone (have no idea what they were lifting, for illustrative purposes only :-)
* Along comes Archimedes who invents what we would call an exoskeleton. Now any regular guy can lift twice as much as last year’s athlete.
* What to do? You can cancel the Olympiads, but they are actually useful as training, motivation, etc So now you have to give the prize on other factors, eg how well he can lift, has he opened a gym in the city, etc
BTW, physics and bio are not exempt, so those researchers better read and try to stay ahead.
For the reality, you can´t close the loop that fast, or with that precision. You will have to slow down by several orders of magnitude.
There is one thing you are missing. Math is precise. Physical measurements are arbitrary imprecise...
Impossibility to independently validate all AI results
And in math it goes even worse. In coding code reviews are typically still the form of action you do within days. In math, historically, the lifecycle of proof is months if not years. Take as an example Millennium problems. They require at least two years of validity after publishing. Two years! In modern times with amount of output AI can produce, it feels like infinity.
We are inches close if not at the moment already when humans can’t reliable validate proofs and mathematics produced by AI. Then next research will be based on this AI-written-no-human-in-the-loop results. And we will end up in just few years in a world where novel and frontier problems will be articulated by AI and proven by AI based on AI results and humans will be incapable of understating the mere nature of the solution.
And, yes: that doesn't absolutely guarantee correctness. The Lean kernel has had soundness bugs, and may have some still. But it's pretty strong evidence of correctness nevertheless.
The concern among mathematicians is not mainly that they doubt the correctness of any of these discoveries, but that human understanding may be devalued.
It is perfectly valid counterpoint to say that we already do it. We everyday use myriad of things, tools, and software we have 0 clue how it operates. But for us as humans it was reassuring that we know that at least there are a few other alive humans who know it, who create it and who can explain it.
With AI soon that comforting zone will be gone.
This is just normal though. We were building sophisticated bronze and steel tools long before any complex understanding of metallurgy or chemistry. Medicine is still the wild west.
I think the most "understanding" we ever had was in the 40s-50s designing nuclear bombs with slide rules. It was the culture that produced the idea of psychohistory.
> Medicine is still the wild west.
Reminder that we have no idea how anesthesia works.
Do any one person even understand the humble pencil?
https://dn790006.ca.archive.org/0/items/i-pencil-pdf-2019/I%...
If we use AI well here we could actually understand math much better than we do now.
And that I think is essential to why some of the world's leading mathematicians are taking this so hard. In a world where we "merely" have AI systems capable of superhuman informal reasoning, verification, correctness, and acceptance could still only be conferred or anointed by human mathematicians. But a world that combines superhuman informal reasoning with superhuman autoformalization is a fundamental shakeup in the institutional order.
It's nice that the author is optimistic, but won't the AI be best placed to dumb down its increasingly complex proofs into a language us lowly humans can understand? To keep thinking until it can refactor complex proofs into ones from 'the book'?
Producing human understandable proofs is possibly a job best for humans today, but the author appears to agree with you that this is probably fleeting (and argues that even if you disagree, it should probably be treated as if it is fleeting when planning for the future):
> Right now AI systems arguably underperform us at theory-building, asking questions, exposition, … so we could prioritize and reward those skills. I think this is unwise: compare the speed at which the academy adapts to the speed at which model capabilities improve. We need to consider the endgame. If the models remain incapable in some domain, we can adjust later.
PS. The validation problem, being one.-
I agree making simple things sound complicated to appear more impressive is bad but there are limits. Even with Feynman he could only go so far, e.g. his interview about why questions and magnetism.
I do think in a lot of fields there is a lot of impact/influence to be had by people who are willing to do work to bridge different fields. I wonder to what degree this is because practitioners don't always see how their work could be used elsewhere.
https://www.maths.tcd.ie/pub/Maths/Courseware/ProblemSolving...
In the exact same way that LLMs allow anyone to vibe code an app but do not replace real understanding of system design due to its essential complexity, non-mathematicians will quickly learn that asking an LLM to pump out advanced mathematical statements to you, even if they are correct (and even if you could verify them) does not constitute understanding, and that the human brain is the bottleneck either way.
It is only if the LLM is super-human at simplification and explaining that a difference will be noted. This would be excellent for mathematics but its not a foregone conclusion (and the argument of most mathematicians, such as Terence Tao, is that this distillation process is one of the key parts of doing mathematics, and that LLMs so far seem to be going in the opposite direction. I suspect its probably user error and leveraging the tools better will produce different outcomes, but mathematicians are only just starting the journey that software developers have been going through, so patience is needed).
But also, with things rapidly changing, perhaps in ten years chat programs will not only do superhuman but make their proofs marvelously accessible and provide incredible tutoring sufficient to bring any curious up to a super high level quickly. Then what can you say and what can you complain of.
But I think "do everything machines" are necessarily inevitable but the situation does make it uncertain where the limits are.
Pedagogy is the primary purpose of educational institutions. That includes universities. However, when research is placed first, you are often left with mediocre teachers, because "who cares?" They were hired to do research and lend the university that kind of prestige; teaching is an afterthought squeezed into the spaces that remain.
And frankly, we haven't the faintest clue today what education is even for. That doesn't bode well.
Add to this the other troubles facing the university, like declining attendance, skyrocketing costs, and AI cheating and we're in for a very interesting transition indeed.
Why should we expect everyone to be both a great researcher and communicator? The obvious result is that it is as effective as engineering managers. Sure, there's some amazing ones, but most aren't. Though that also doesn't mean no expertise in the field (i.e. non-engineering manager) is any better. It is just that managing/communicating is a different and orthogonal skill.
What needs to happen is we need to make it okay for people to specialize in more things. More nuance to this rather than trying to throw everyone into nice easy to manage buckets. Those buckets are just unrealistic abstractions filled with hope, denial, and laziness. Reality is surprisingly complex. Math can do a really good job helping you understand that, but it's a sufficient condition, not a necessary one
I also have a degree, but mostly taught myself the important undergrad-level concepts as a kid. I just went to the library and borrowed any of the hundreds of books written to clearly communicate maths to beginners or downloaded any of the free ebooks / lecture notes.
Name a single field of endeavor that is more open in 2026. Software certainly isn't one of them - the best stuff has always been gatekept.
I hope, from the depths of my soul, that the static typeset report format for transmitting knowledge and understanding will finally die and be laid to rest.
There simply is no contender to LaTeX and PDFs.
Lucky for you, almost all research in math, cs, and physics, are put on arxiv, where you can download the source code (.tex) as well as get an HTML render.
https://ciechanow.ski/archives/
...for starters?
By essentially the same reasoning, I’ve been arguing for prioritizing in-person design/code reviews over code-only async PR comments.
The important thing is to verify that the human has a coherent design in mind and can demonstrate that it got implemented, regardless of who or what was at the keyboard. “I dunno, I guess Claude thought this was a good idea” is not a coherent design.
AI writing is garbage. If you're indifferent to how much better good human writing is than AI writing, you are not qualified to evaluate writing.
Manager: “Why isn’t feature X available?”
Person B: “key pieces are delayed due to the developer not understanding all of the LLM doesn’t and implementation.”
Manager: “does it work? What are the risks?”
Person B: “well yes it works for now but we’re accumulating tech debt due to a lack of understanding and potential flaws that haven’t been thought out yet”
Manager: “they want feature X, ship it, we can deal with it later, I don’t care if it’s not coherent as long as it works.”
How many decades at this point has these been a push for functionality over everything at all costs? And you have a mechanical snow plow now. Most businesses don’t care about later risk or any future planning beyond the quartet horizon, they’re not concerned about how it will effect their performance in 3 quarters or lead to instability or issues, those are future problems for a future person and we’re here for money now.
Individuals within an organization do not unless it will specifically affect their bonuses/promotions.
You can make a perfectly rational business choice to take on tech debt to get a feature sooner.
Oral defenses in academia are largely rituals. If a student fails their defense, it's almost always the supervisor's fault. The supervisor is supposed to be the primary quality control. With their regular meetings with the student over the years, they should be able to tell whether the student has achieved sufficient understanding and contributed enough to graduate. If the supervisor thinks the student is ready to defend, the reasonable expectation is that the student will pass and graduate. The defense is mostly there to let the other examiners validate the supervisor's judgment.
The defense was only briefly about the actual thesis, then switched over to whatever research interest the committee members had, and they'd drill into their pet subjects. This was in physics, hardly politics forward normally, and at a very well respected university.
Also in my observational experience in physics, a fair number of students (probably like 10%) get some sort of major corrections to their thesis. After writing this, I did a search, and in the UK, the number across all fields is 16% [1].
[1] https://www.lexacademic.com/blog/how-common-is-passing-with-...
I am sorry, but the writing on the wall is that this is where software industry wants to go. As a software developer, I cannot but help notice how unrealistic this is.
We are entering the next level of bloated, inefficient and unreliable software....
1. It will take much longer to understand the output of the machine that it takes to prompt and create it. 2. The only? best? one? way to /verify/ that you /in fact/ understand the output of the machine is to explain it to someone else.
So there will be a machine generating koans which need to be meditated upon and discussed with human social back-pressure validating understanding. I think this could be much more cooperative and at a minimum this will be a way different math social construct.
With AI, it might be the case that we don't know, we won't know, but the machine does.
The central question, namely whether humans should be in the loop, will be repeated again and again in the years to come for all industries, starting with mathematics.
People had the same complaints that the code produced by early coding models was messy, lazy, poorly commented, and so forth. The central complaint was that it was just too difficult for humans to review. The answer is just to improve the models and move on.
Similarly now we're getting AI doing math. The math is a giant vibe-coded ball of wax. So just make the models better at explaining what they're doing to humans - that's the end of it.
Rather than just go on and on about how it's the end of the world if we don't do this, why don't we just do it?
But for math: what is the point of a proof if no one will read it and no one uses its result? To quote the article, "AI systems will [...] result in the production of an abundance of PDFs. The contents of some of those PDFs may even have important applications." But if there's no one reading the PDFs, what's the point - no matter how good your AI model.
The point of math is understanding. So mathematicians should feel free to use AI as much as you want, but in the end, they should've gained some understanding on what happened.
What is the point of writing software if nobody will run it?
> So mathematicians should feel free to use AI as much as you want, but in the end, they should've gained some understanding on what happened.
So they ask the AI to explain the proof.
I do think that is one goal of math but I don't think it's the only one.
I think an additional goal is simply "truth", which can be found without understanding as we've seen with these human-incomprehensible proofs.
Yet another is practical applications. While there's less of these in pure mathematics than in most domains, they do still exist.
You can understand this as learning a mapping between the model's internal "world" (i.e., 'meaning,' which is hopefully coherent and consistent -- but definitely not always! see, e.g., https://arxiv.org/html/2505.11581v1) and language (i.e. 'form') that reflects that world.
For this to work, you need both coherent / consistent internal model worlds, and also good mappings onto human language. Supervision by mathematicians has provided the signal for both internal coherence (though this can also come from interacting with a proof oracle) and for good explanations. If models exceed human capacities, you could imagine that aligning their explanations potentially becomes harder (though not necessarily). Also, humans naturally have to do the same thing: as researchers we must find analogies to make our work legible to collaborators or laypeople. Often in doing this, we further clarify our own understanding!
More deeply I think the "end of the world" vibe arises not only from the practical need to have models that explain, but also Litt's (and many other fields' researchers) grappling with being relegating to not mattering.
They can write great code, they can gain understanding of something for themselves, but they still kinda suck at explaining it to people. As the article says, if you want understanding, it seems like there is no replacement for getting in the weeds yourself. A model can help, but it's not going to magically download knowledge into your brain.
The assumption here is that the true/false of the theorem is the important outcome. While it is certainly part of it, a big part of maths is the understanding you gain from a proof. Many of the best proofs elegantly explain some aspect of the maths which was previously unclear and expand our understanding of the world.
To use a programming related example, imagine that an LLM spits out a solution to the travelling salesman problem which works in O(n) time. On the one hand that's very convenient for whatever problem you happen to be trying to solve at the time...but there's also an answer to P=NP in there! The former means your delivery drivers app works a bit faster on their busy days, the latter fundamentally shifts how humanity thinks about certain problems.
Going back to the maths, there have been theorems that were proved (by people) where the proof is broadly seen as 'unsatisfactory' in that it doesn't really expand our understanding. I assume some of these LLM proofs are a bit like that: we now know that the thing is true, but we really want to know why it's true, and how that changes our understanding.
> The assumption here is that the true/false of the theorem is the important outcome
You're replying to a comment about "clean proofs" and explaining to humans. You talk about true/false anyway. Re-read please.
They hire PhD students without hearing them give a talk and then doing interviews? In Germany, the applicant gives a talk (30-40 min) to the research group they want to join, usually presenting their master's thesis, engage in discussion, often share lunch with the group, then do 1 on 1s with individual members of the group and a longer one with the PI. Obviously this can vary within Germany too, but I couldn't imagine hiring someone without something like this.
when I applied to PhD programs (not in math) it was basically CV + personal statement + recommendation letters + short chats with interested faculty :shrug: Maybe it was because my CV was "strong" but the chats were more see if interests were aligned, rather than actually interviewing me.
In code, any developer whose differentiaton was clean code and knowledge of different technologies is now average.
In math, any mathematitian whose differentiation was to manipulate formal systems and know tricks of different domains will be average.
Fortunately, humans do more than syntax and search.
The bad news for developers is that if you know what the output of your program should be (which happens most of the time), almost all of the job is syntax and search to build the code that reproduces the output.
The good news for mathematitians is that for the majority of problems you never know the output, or you just know the output is either "True" or "False". There are some cases where you need something else, for example "a solution that blows up in finite time". For those cases AI will outperform you easily (see new Navier-Stokes solution)
So if as a mathematitian you were doing more than syntax and search, then keep doing that and use AI just for what its best.
If you're a mathematician you are in the same boat as the software engineer, and the Dodo.
Better learn a trade buddy /s
I thought we were going to get at least some brief comment on Godel here?
But the bigger picture is: while I understand the author know his field best and wants to keep the post focused, the same issue will hit many more fields. We need to also have a broader discussion that involves more fields of knowledge work, largely academic scholarship but also regular office work, then it will come to engineering design, medicine, it's already coming for 3d modeling and vfx, software dev, it will come for a lot of middleman services. Not at the same rate, but we have to understand that it's not just that math will change.
Change will be the default. Everything will change. It will be much easier to make math change because all things will change. You shouldn't worry and imagine that funding criteria will be like today or that journals or academia or politicians expectations will be like today. No, everything will adjust with some timing differences of course but it won't be a static world and then math changing and having to justify and fight to explain the change to other actors who are baffled. They won't be baffled they will themselves have to change.
The world will transform as much as it did when society moved from feudal agrarian to urban capitalist industrial, or from the vast majority doing physical labor to a service economy with a huge amount of desk jobs. I can't tell how it will change exactly but it will be bigger than what we have seen in the last couple of generations or maybe more.
There are a good deal of people, whom, falter much more in oral discussions, whether this be for a psychological thing, stage fright, or difficulty explaining things on the spot. There are also certainly brilliant people, who can’t give an informative, discussion inviting talk to save their lives, but given enough time, can formalize their thoughts in writing at the highest levels of their field, and that writing is likewise very enlightening (sometimes).
It’s not clear to me, that, AI as is, could not pose successfully in an oral discussion of a topic. I mention this because it seems that one implication of the article is that AI might write things that are logically correct, but devoid of understanding. I suggest rather that 1) it is not extremely improbably that AI is incapable of generating mathematics that furthers human understanding and if 2) it is indeed highly likely that they cannot generate mathematics that furthers human understanding in a textual format, then surely one could also differentiate between human and AI on a textual level, and judge the contribution of a human, without the need of oral discussion?
I suppose another aside is, one might claim that the existence of AI means people have much much more text to filter for, and so, it becomes difficult to find one person’s good writing amidst a sea of, logically correct, yet understanding devoid textual content. But by and large much or mathematical academia certainly operates off of some reputation/vouching system presently anyways, that already serves as a “filter” in some sense. Perhaps the existence of such a system/culture is not a good thing, but oral discussions/seminars certainly aren’t immune from such predilections.
Perhaps I’m babbling like an idiot, but the entire and sole purpose of this comment is just to say: for the love of god please don’t let the standard be judged by oral presentation
One of the things that struck me was that most people have an area of their workflow that they would be happy to hand off to an AI, so that they can focus on the thing they love. However, that area is different for each person. One person's grind-work is another person's love-work.
This is why it's so hard to build consensus around where the "red lines" are in this space.
They are leeching off the work of mathematicians and pretending that new math was invented.