I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
I'm enjoying learning about these hard problems, but this line about credit made me chuckle:
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
It seems that a lot of folks misunderstand the guarantees that lean provides.
I just want to state that having "lean proofs" that build (checks) does not mean the actual real theorems we care about hold. Ignoring lean kernel bugs, ultimately a human (not an agent) has to verify the lean encoded theorem statements (specs/specifications), that the lean proofs are checked against, indeed correctly encode the real theorems. For non-trivial theorems such as these, this is an arduous and tricky task where even a little mistake could be fatal. AI generated lean encoded theorems can be huge and difficult to understand. I wonder if anyone reputable has audited these specifications.
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up?
2. How many attempts did you give the model at solving these problems?
3. How expensive was the harness, e.g. did the model have access to a job cluster?
> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
Hey man, just wanted to say hi and catch up a bit. I tried DMing you on Twitter. If that sounds interesting then shoot me a message sometime. Hope you’ve been well :)
On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
In a way the most remarkable thing about this is that it isn't even at the top of the HN homepage. Even if this is a step up from what we've seen before, we're no longer astonished by the idea that AI can make significant advances in mathematics and computer science.
Replace philosophers for mathematicians and Douglas Adams was spot on again.
Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.
--
"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"
"What's the problem?" said Lunkwill.
"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"
"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"
"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
> Whilst current models can't 'intuit' and come up with conjectures
I disagree. I routinely let LLMs speculate or generate hypotheses along the way of helping with technical research. Sometimes they can prove the correctness of a concrete math idea but other times even an unproven conjecture helps with the numerical algorithm implementation and the result is then simply supported by additional data. I guess that any autoresearch-adjacent application has LLMs intuiting and coming up with hypotheses/conjectures—as do the steps/lemmas along a complex proof. In my opinion the modern LLMs are powerful intuitive thinkers that generate lots of conjectures of varying quality or importance.
Because people have internalized an inaccurate model of LLMs as "stochastic parrots" that was incorrect at the time of formulation and is also significantly outdated
Well imho, it's a bit of a fundamental problem for a certain aspect of the meaning of "to intuit".
Since we're quoting Douglas Adams in this thread, I'll mention something I posted a while back, with his writings as example. After Douglas Adams passed away, somebody was tasked to "finish" The Hitchhiker's Guide to the Galaxy :
> And Another Thing... is the sixth and final novel in The Hitchhiker's Guide to the Galaxy series. Written by Eoin Colfer with the blessing of Douglas Adams' widow Jane Belson
Being a rather big fan, I immediately bought and read this novel and I have to say that Eoin Colfer did a really really great job, nailing the tone, humour and writing style of Douglas Adams.
IMVHO, he did about as good as anyone could reasonably expect someone to do, when given this task. It was big shoes to fill, and I was impressed.
But it just also wasn't good enough, in a weird way that I found hard to put my finger on at first.
The thing is that Colfer was doing the tone of voice, even came up with somewhat new jokes perfectly in the style of, etc etc. And for the sake of argument let's say he was able to get "arbitrarily close".
But there was always one thing he couldn't do: Actually make something new happen, make a new kind of joke, do a real plot twist, a big reveal, stuff like that. Because then it would deviate from Douglas Adams' work too much.
However, if Douglas Adams was still alive, this limitation would not apply to him: he could make a new kind of joke, do a plot twist, big reveal, and it would become canon.
This the best "good faith" argument I can present for how LLMs lack "intuition", in some sense. Now "intuition" is not a very exactly defined term, but I'm arguing that the thing I'm describing here, is at least a part of intuition, that an LLM fundamentally can't reach (until they start getting their own volition, which I would prefer they didn't).
To address your question:
> Surely the AI can complete the prompt “Generate new research questions based on these observations”?
Yes I imagine it could do that very well. But it would still need a human to decide if the research questions are "relevant" or "within scope" of what the human wants (a.k.a. their volition). Without that filter, the research would just bloom out exponentially, with more and more questions nobody was asking.
And yes, up to some point that "blooming" behaviour is a useful aspect of research, the exploratory aspect/phase, but at some point you need to get back to the "synthesis" aspect/phase, to distill all the explorations back to "stuff that matters". And just like Douglas Adams vs Eoin Colfer, only the human who wants to know something, can decide to widen or change the domain of that synthesis, but if the LLM were to decide this (outside of exploratory phase), it would actually be considered the wrong answer.
And this is not at all to say you can't do research with LLMs, obviously you can. But this is just a thing they can't do, on a real philosophical level. I'm also not saying you can't work around this limitation, you probably can, I'm just saying it exists.
That's how they are finding these solutions though, unless we are just going to label intuition as something only humans can do. Like a submarine being unable to swim or whatever that example is.
> they can certainly disprove some of them very quickly through the kind of grind that humans can't do
Of course computers can grind in a way that humans can't. But now we have systems that convert the human-comprehensible ideas into a computer's plan of attack, in a way that greatly expands the frontier of ideas thus treatable.
Ahh, but you missed the continuation, where they get to the heart of the matter: money.
"Excuse me, We demand rigidly defined areas of doubt and uncertainty!"
DT: Might I make an observation at this point?
MT: You keep out of this metal nose.
VF: We demand that that machine not be allowed to think about this problem!
DT: If I might make an observation…
MT: We’ll go on strike!
VF: That’s right. You’ll have a national philosopher’s strike on your hands.
DT: Who will that inconvenience?
MT: Never you mind who it’ll inconvenience you box of black legging binary bits! It’ll hurt, buster! It’ll hurt!
DT: [Booming] If I might make an observation …
“All I wanted to say,” bellowed the computer, “is that my circuits are now irrevocably committed to calculating the answer to the Ultimate Question of Life, the Universe, and Everything.” He paused and satisfied himself that he now had everyone’s attention, before continuing more quietly. “But the program will take me a little while to run.”
Fook glanced impatiently at his watch.
“How long?” he said.
“Seven and a half million years,” said Deep Thought.
Lunkwill and Fook blinked at each other.
“Seven and a half million years!” they cried in chorus.
“Yes,” declaimed Deep Thought, “I said I’d have to think about it, didn’t I? And it occurs to me that running a program like this is bound to create an enormous amount of popular publicity for the whole are of philosophy in general. Everyone’s going to have their own theories about what answer I’m eventually going to come up with, and who better, to capitalize on that media market than you yourselves? So long as you can keep disagreeing with each other violently enough and maligning each other in the popular press, and so long as you have clever agents, you can keep yourselves on the gravy train for life. How does that sound?”
The two philosophers gaped at him.
“Bloody hell,” said Majikthise, “now that is what I call thinking. Here, Vroomfondel, why do we never think of things like that?”
“Dunno,” said Vroomfondel in an awed whisper; “think our brains must be too highly trained, Majikthise.”
So saying, they turned on their heels and walked out of the door and into a life-style beyond their wildest dreams.”
From a mathematician who was intimately familiar with some of these problems [0]
>I don’t understand it yet. Maybe it’ll take me an afternoon to check all the calculations, but what would still be missing is why this was an approach that would’ve made sense in the first place. Is there some broader context or theory within which this would’ve been the obvious thing to do? What other results can be proven using these techniques? What is it telling us about quantum information or operator theory? I have no idea. I spent about an hour this morning asking ChatGPT these questions, but it’s somewhat frustrating because it speaks with a mishmash of physicist, operator algebraist, quantum information theorist-lingo, plus the usual LLM breezy lilt that annoys everybody.
They certainly seem to have "intuited", in a way that is not immediately obvious to experts in the field, the way to solve at least some of these problems. This was not just simply grinding away at a method that humans already knew would work and just hadn't gotten to yet.
It's remarkable how you can manage to get these models to produce remarkable breakthroughs like an explicit construction of a non-sofic group.
And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.
Any implication of any of these findings? They seem like unimportant nerd snipes to me. If you want to do something actually relevant, get chatgpt to write a simulation of graphene nanotube construction and figure out how to do it at scale.
Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely.
The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
Guys, please use critical thinking. The haters don't hate by default, we hate because we're gaslit about this stuff every day and it's annoying. Extraordinary claims require proof, and they're not giving us information that would be essential to knowing if this is actually significant or not.
It is a fact of experience, and indeed effectively a theorem, that the better they get at coding and math, the dumber they are. These are the wages of RLVR etc
The models are frequently getting worse at items that they aren’t being benchmarked for — and that’s happening more and more over time! Other people in other fields aren’t idiots, they are accurately perceiving the fact that these models are being hyper optimized for our industry, and are becoming less capable in other domains over time. Models of the same scale are massively worse at writing a broad variety of styles of prose than their equivalent from two years ago. (Models of increased scale are a mixed bag.)
Maybe you’re the one who needs breaking out of your cached beliefs.
every lab independently discovered that getting good at bit alchemy (coding and related tasks) should come first as it will enable the formation of training pipelines that will then solve everything else.
so far there is no end to this progress in sight so it's full steam ahead on this singular domain. once it plateaus you should expect to see the greatest disruptions in human endeavors ever as all the training flops will start flowing to other domains to disrupt and dominate.
I don't think this is a straw man, a huge number of people in my life (non CS people) think AI is a dead-end, that it's just a stochastic parrot, that it'll never be able to do many things that humans can do. I have had many arguments with people who told me that "AI will never be able to do X", and then 6 months later AI is able to do X. Then they will move the goal posts and say "well AI will definitely never be able to do Y".
People will be broken out of their denial by actual economic growth. That's what this is all meant to be for... I think we might start seeing some surprising numbers.
Can someone explain why we’re not using it to solve the obvious big conjectures/problems though? If this stuff is solved, why isn’t Riemann the first thing to go? Even if it means a fund of several hundred $k to let it churn on it.
Now that we've seen AI produce a fair number of proofs (and disproofs), I'm curious when we'll start seeing it build genuinely novel theory. Does anyone have predictions on when and how we'll get there and will it take new architectures/ training paradigms, or is the current approach enough?
one of the early premises of how ai takeoff would go was that a system that could solve open problems in advanced mathematics would also discover novel advances in math and computer science that directly unlock drastically better software performance.
we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B).
we are also seeing incredible advances in software performance. open ai announced like 15% improvement by fixing gpu kernel issues.
these are clearly linked in the sense of scaling laws and generalization of intelligence: a huge model gets capabilities in both math and software engineering that isn't possible at smaller scales.
but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence.
to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math
(i also posted this on twitter @mlipman13)
This will depend on the problem; I expect big algorithmic performance improvements in AI since the algorithms are still new, inefficent, and constantly being improved. But maybe not for sorting, fast fourier transforms, or other well-studied basic algorithms?
I don’t really like AI but let’s stop kidding ourselves, no human mathematician could make progress on a dozen major open problems in a week or two. If you’re measuring it against humans then it is by far the best mathematician to ever live.
I guess it depends on how to measure a single "person"? If you spun up 2000 copies of Terrance Tao, I wouldn't be surprised if you found a few new discoveries at the end of it.
> I don’t really like AI but let’s stop kidding ourselves
If I had to create a tagline to describe my opinions about AI in a single sentence, that’d be it.
It’s possible to both hate AI and be impressed by it at the same time. Lying to ourselves about its capabilities does us no good. It’s emotionally difficult to do, but people need to come to grips with what’s happening and shake themselves out of a state of denial.
How much investment has gone into OpenAI versus mathematics research in 2025 for example? Probably 100x?
The AI results are clearly impressive. But these sorts of things are also in the ballpark of what human effort could solve given enough attention and time. Though it is hard to say.
Correct me if I'm wrong, but all of the aforementioned advances were made in the last year? Until very recently few people had access to these tools. Most people still don't know how to use ChatGPT, and very few use tools like CC regularily. If in a few years these frontier tools become commonplace and people upskill we would should see a network effect?
> but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance.
You seem to overlook a simpler barrier. To make these advances, they have to be possible. A 15% improvement in GPU kernels doesn't evidence that significantly more improvement has been left on the table.
Google has also invested a lot of time into developing new hardware and new algorithms with AI (other types of AI, not LLMs). I don't know if it's paying off (haven't followed it closely) but they seem to think it's worth the effort.
158 comments
[ 0.20 ms ] story [ 18.9 ms ] threadNow I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
[edit: deleted a distracting comparison to Chess]
these LLMs are great are generating arguments but they don't ask questions, we will need mathematicians to shepherd them into more discoveries
i really want to see open weight models crack some breakthroughs
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
I just want to state that having "lean proofs" that build (checks) does not mean the actual real theorems we care about hold. Ignoring lean kernel bugs, ultimately a human (not an agent) has to verify the lean encoded theorem statements (specs/specifications), that the lean proofs are checked against, indeed correctly encode the real theorems. For non-trivial theorems such as these, this is an arduous and tricky task where even a little mistake could be fatal. AI generated lean encoded theorems can be huge and difficult to understand. I wonder if anyone reputable has audited these specifications.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
Look at this two threads.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
Look at this two threads.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.
--
"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"
"What's the problem?" said Lunkwill.
"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"
"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"
"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
I disagree. I routinely let LLMs speculate or generate hypotheses along the way of helping with technical research. Sometimes they can prove the correctness of a concrete math idea but other times even an unproven conjecture helps with the numerical algorithm implementation and the result is then simply supported by additional data. I guess that any autoresearch-adjacent application has LLMs intuiting and coming up with hypotheses/conjectures—as do the steps/lemmas along a complex proof. In my opinion the modern LLMs are powerful intuitive thinkers that generate lots of conjectures of varying quality or importance.
People keep saying this. Why?
Surely the AI can complete the prompt “Generate new research questions based on these observations”?
When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.
You can't find things on a map that aren't there, but maybe you can draw a route nobody used before.
Since we're quoting Douglas Adams in this thread, I'll mention something I posted a while back, with his writings as example. After Douglas Adams passed away, somebody was tasked to "finish" The Hitchhiker's Guide to the Galaxy :
> And Another Thing... is the sixth and final novel in The Hitchhiker's Guide to the Galaxy series. Written by Eoin Colfer with the blessing of Douglas Adams' widow Jane Belson
Being a rather big fan, I immediately bought and read this novel and I have to say that Eoin Colfer did a really really great job, nailing the tone, humour and writing style of Douglas Adams.
IMVHO, he did about as good as anyone could reasonably expect someone to do, when given this task. It was big shoes to fill, and I was impressed.
But it just also wasn't good enough, in a weird way that I found hard to put my finger on at first.
The thing is that Colfer was doing the tone of voice, even came up with somewhat new jokes perfectly in the style of, etc etc. And for the sake of argument let's say he was able to get "arbitrarily close".
But there was always one thing he couldn't do: Actually make something new happen, make a new kind of joke, do a real plot twist, a big reveal, stuff like that. Because then it would deviate from Douglas Adams' work too much.
However, if Douglas Adams was still alive, this limitation would not apply to him: he could make a new kind of joke, do a plot twist, big reveal, and it would become canon.
This the best "good faith" argument I can present for how LLMs lack "intuition", in some sense. Now "intuition" is not a very exactly defined term, but I'm arguing that the thing I'm describing here, is at least a part of intuition, that an LLM fundamentally can't reach (until they start getting their own volition, which I would prefer they didn't).
To address your question:
> Surely the AI can complete the prompt “Generate new research questions based on these observations”?
Yes I imagine it could do that very well. But it would still need a human to decide if the research questions are "relevant" or "within scope" of what the human wants (a.k.a. their volition). Without that filter, the research would just bloom out exponentially, with more and more questions nobody was asking.
And yes, up to some point that "blooming" behaviour is a useful aspect of research, the exploratory aspect/phase, but at some point you need to get back to the "synthesis" aspect/phase, to distill all the explorations back to "stuff that matters". And just like Douglas Adams vs Eoin Colfer, only the human who wants to know something, can decide to widen or change the domain of that synthesis, but if the LLM were to decide this (outside of exploratory phase), it would actually be considered the wrong answer.
And this is not at all to say you can't do research with LLMs, obviously you can. But this is just a thing they can't do, on a real philosophical level. I'm also not saying you can't work around this limitation, you probably can, I'm just saying it exists.
That's how they are finding these solutions though, unless we are just going to label intuition as something only humans can do. Like a submarine being unable to swim or whatever that example is.
Of course computers can grind in a way that humans can't. But now we have systems that convert the human-comprehensible ideas into a computer's plan of attack, in a way that greatly expands the frontier of ideas thus treatable.
"Excuse me, We demand rigidly defined areas of doubt and uncertainty!"
DT: Might I make an observation at this point?
MT: You keep out of this metal nose.
VF: We demand that that machine not be allowed to think about this problem!
DT: If I might make an observation…
MT: We’ll go on strike!
VF: That’s right. You’ll have a national philosopher’s strike on your hands.
DT: Who will that inconvenience?
MT: Never you mind who it’ll inconvenience you box of black legging binary bits! It’ll hurt, buster! It’ll hurt!
DT: [Booming] If I might make an observation …
“All I wanted to say,” bellowed the computer, “is that my circuits are now irrevocably committed to calculating the answer to the Ultimate Question of Life, the Universe, and Everything.” He paused and satisfied himself that he now had everyone’s attention, before continuing more quietly. “But the program will take me a little while to run.”
Fook glanced impatiently at his watch.
“How long?” he said.
“Seven and a half million years,” said Deep Thought.
Lunkwill and Fook blinked at each other.
“Seven and a half million years!” they cried in chorus.
“Yes,” declaimed Deep Thought, “I said I’d have to think about it, didn’t I? And it occurs to me that running a program like this is bound to create an enormous amount of popular publicity for the whole are of philosophy in general. Everyone’s going to have their own theories about what answer I’m eventually going to come up with, and who better, to capitalize on that media market than you yourselves? So long as you can keep disagreeing with each other violently enough and maligning each other in the popular press, and so long as you have clever agents, you can keep yourselves on the gravy train for life. How does that sound?”
The two philosophers gaped at him.
“Bloody hell,” said Majikthise, “now that is what I call thinking. Here, Vroomfondel, why do we never think of things like that?”
“Dunno,” said Vroomfondel in an awed whisper; “think our brains must be too highly trained, Majikthise.”
So saying, they turned on their heels and walked out of the door and into a life-style beyond their wildest dreams.”
>I don’t understand it yet. Maybe it’ll take me an afternoon to check all the calculations, but what would still be missing is why this was an approach that would’ve made sense in the first place. Is there some broader context or theory within which this would’ve been the obvious thing to do? What other results can be proven using these techniques? What is it telling us about quantum information or operator theory? I have no idea. I spent about an hour this morning asking ChatGPT these questions, but it’s somewhat frustrating because it speaks with a mishmash of physicist, operator algebraist, quantum information theorist-lingo, plus the usual LLM breezy lilt that annoys everybody.
They certainly seem to have "intuited", in a way that is not immediately obvious to experts in the field, the way to solve at least some of these problems. This was not just simply grinding away at a method that humans already knew would work and just hadn't gotten to yet.
[0] https://nitter.poast.org/henryquantum/status/208362369543662...
And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.
Truly jagged beyond belief.
The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
not sure how many will get this reference but "AI" for science and math is like super-shoes for runners
at first we are blown away by the impossible improvements including sub-2-hour realworld marathon and every other PR/CR/WR is dialed down
but then the improvements slow and reach a stall point because of the limit of technology and the source of the achievement
ie. sub-2-hour marathon yes, sub-1-hour never happening (rollerblade inline-skate record is 1-hour marathon)
I heard that Gary Kasparov was impacted by AI chess, but at least he still seems to have a job, so don't give up.
But it's not clear if LLMs will produce abstractions that humans would find elegant.
https://garymarcus.substack.com/p/two-critical-updates-re-as...
As always, PR hype. Goalposts have not moved.
Guys, please use critical thinking. The haters don't hate by default, we hate because we're gaslit about this stuff every day and it's annoying. Extraordinary claims require proof, and they're not giving us information that would be essential to knowing if this is actually significant or not.
Maybe you’re the one who needs breaking out of your cached beliefs.
That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.
so far there is no end to this progress in sight so it's full steam ahead on this singular domain. once it plateaus you should expect to see the greatest disruptions in human endeavors ever as all the training flops will start flowing to other domains to disrupt and dominate.
but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)
If I had to create a tagline to describe my opinions about AI in a single sentence, that’d be it.
It’s possible to both hate AI and be impressed by it at the same time. Lying to ourselves about its capabilities does us no good. It’s emotionally difficult to do, but people need to come to grips with what’s happening and shake themselves out of a state of denial.
The AI results are clearly impressive. But these sorts of things are also in the ballpark of what human effort could solve given enough attention and time. Though it is hard to say.
You seem to overlook a simpler barrier. To make these advances, they have to be possible. A 15% improvement in GPU kernels doesn't evidence that significantly more improvement has been left on the table.