11 comments

[ 2.7 ms ] story [ 13.0 ms ] thread
Scott Aaronson remarks in this article that the labs are sitting on a whole pile of incredible results that they cannot figure out how to release given the reception to the Navier-Stokes proof.
Satisfying to see Aaronson, who is usually downplaying some claimed advance, correctly panicking. It makes me feel better about my own.

For there is always a time to panic, no matter how exhausting the requirement might seem. And that time is now.

I have a friend who earnestly told me I sound like an evangelical about AI. He worried for my sanity.

I told him that if I saw that Jesus was alive, walking amongst us, performing actual miracles - I would definitely try to convert him. I'm not one to look away from hard data just because it's inconvenient.

That's exactly what's going on right now. People's reactions to it are so weird.

> "Do you publish a paper that lists “GPT-6 Astra” or “Claude Fable” as the author—but then let the AI profusely thank you in the acknowledgments for suggesting such a wonderful problem to it?"

"I thank Human Mathematician for their words of encouragement and for reminding me to believe in myself".

On the contrary, I'd urge people who actually adopted LLM agents for work to skip the second one. Reading it risks giving an aneurysm at worst, and wasting one's time and effort at best.

The first link however will be surprisingly delightful.

> On the contrary, I'd urge people who e.g. adopted LLMs and agentic workflows for work to skip the second one. Reading it risks giving an aneurysm at worst, and wasting one's time and effort at best.

I have “adopted LLMs and agentic workflows for work” and I neither had an aneurysm nor felt that my time was wasted after reading it, and I think it showed the very tempered expectations that you reference.

Care to actually engage with the points in the article?

There's a reason I didn't: the article itself did not bother. They just put things out there, "thought-leader to thought-leader just like that", to quote a recent HN comment I found hysterical in an unrelated thread. So to elaborate felt like an isolated demand for rigor on my part.

Nevertheless, point by point:

1. price competitiveness and autonomy:

They falsely assert that "current frontier models need laborious oversight and guardrails on even the simplest tasks" - this is simply untrue. We have countless workloads at work that are easily doable for an agent. I experience this time and time again, and have been for half a year now. There are lots of colleagues of mine who only do such tasks. They do a worse job at them (sometimes way more), and take longer to do so. Using guardrails is basically unneeded, and the oversight burden is also far from laborious.

I also work with those "crappy engineers" they speak of, and I wish it was only a bottom quartile. Haiku outperforms them; they need laborious handholding. Handholding that they do not comprehend, because their grasp of the English language matches their general technical expertise (i.e. you could hire a random guy off the street and they'd be better). Mind you, I do also keep running the numbers, and they're already not worth it. This entire section was just blatantly wrong, at least in my sector (cloud operations and devops).

2. jagged intelligence / the models are idiot savants

Yes, they are. Their intelligence rises with the amount of parameters, and remains mostly local to their "area of expertise", so they clearly scale that way. This has been discussed to hell and back, and should be more than familiar enough to anyone reading this forum; it's trite.

3. specification burden

It's the exact same as with crappy people. Literally the exact same. If you find yourself specifying things so hard, the model is simply not good enough for that workload. You can indeed absolutely dig yourself into a hole and make things not worth it, this is also extremely trite. It's been an adage with automating anything since forever. "You spend 5 minutes doing something that annoys you, or spend 30 minutes automating it." Unless you never heard this adage before somehow, this won't be new.

4. same as the previous one

5. navier-stokes and proof hacking

All of these caveats are well understood by the relevant community, and are well accessible to those still within tech but outside of math, too: https://news.ycombinator.com/item?id=49672339

6. human review

It is not an alternative, it is inescapable. See my previous link. We're already a bottleneck, which is again obvious to anyone who uses these things on the daily.

7. (they start from 1. again) cheap iteration is key

Yes, just like with people. And again, this is just readily apparent if you ever tried to guardrail an agent. It's an uphill battle. More cheaper failures basically always beat fewer expensive failures, this is nothing new, and nothing necessarily specific or unintuitive.

8. repetitive work is more easily performed

Yes... just like with people. That's how the whole manufacturing line and the various specialized stations came to be. Businesses are built around continuously ossifying their processes into more standard, more textbook ones. This is exactly why agents are eating the bottom layers, and why people here keep saying "the bar is rising".

9. doesn't check out cost wise, what does is already code-automated

No, there is absolutely a space between the unautomated and the code-automated where these agents slot in fine.

That's all their points, the rest is just a conclusion. These are very basic experiences anyone can identify, littered with some tropes everyone sees here every day, delivered in a quirky format that people are hyping the sn...

“Yes, there’s still enormous uncertainty about what the rest of our lives will look like, but as far as I can tell, there’s no longer any real uncertainty that it’ll all mostly revolve around AI, and the extent to which we succeed or fail at directing its power toward human flourishing.

Good to know that in the throes of coping with apparent existential unknowns, even smart people can still wiggle in some hubris. Imagine, using his own analogy, upon the resurrection you try to steer the son of the Hebrew god toward the perceived benefits of western moral philosophy.

> Accepting the reality of the coming machine god after it’s solved Navier-Stokes

One thing that stands out to me about this whole ordeal is how little appreciation there is for the fact that instead of there having been one machine god involved, it was the work of ten thousand or more.

I think it's all too easy to forget, maybe even mistakenly assume otherwise, that the prompt you put into the chat textbox is not going to receive the same kind of machine attention the NS problem did.

The models are really impressive, but I don't think it's cope to remark that this was in essence a 1:10000 chance outcome. This is assuredly closer than it ever has been, or was even imagined to be, but is still far from the image these announcement paints in one's head.

Astra's successor is not going to solve millenium problems for you. Even if they double in capability each generation as claimed, you're still to wait until GPT-16 till you will have a millenium problem capable model in your service; and even if they release twice a year, that's still almost a decade away.