Terry Tao using coding agents to build apps means we're one step away from a Fields Medalist asking an LLM why his Docker container won't start, just like the rest of us.
"as such [LLM-coded interactive] supplements are not mission-critical to the core of the paper, I again feel that the downside risk of using guided interaction with LLM agents to generate such visualizations is acceptable."
It's a tool. Good for some things but not others and generally not to be trusted.
There are many AI bulls who adamantly disagree and cite Tao’s statements about LLMs for mathematical proofs as an example of how advanced and autonomous these systems already are
> It’s a tool. Good for some things but not for others and generally not to be trusted.
I agree completely you always need to check the work of LLM agents, but it does strike me as a tiny bit funny to anthropomorphize AI by using ‘trust’ while warning against anthropomorphizing the AI by using unchecked output. ;) Generally speaking, “trust” in AI has been going up very quickly as the models & harnesses improve, and as people figure out effective workflows.
I trust my hammer with nails but not screws… does that mean the hammer should generally not be trusted? The problem with AI is we don’t know the difference between nails and screws. (This may be where my analogy breaks down. :P) But I feel like saying don’t trust it isn’t as helpful as saying something like you should expect to spend more time planning and iterating than before, and you should expect tot spend more time reviewing and checking output than before, and learn how to use skills and context and subagents, and learn to use AI on some non-production low-consequence projects first. Saying ‘generally not to be trusted’ implicitly suggests not using AI, and doesn’t leave the reader with how to use AI. The goal is to build trust by building good workflows and by understanding what works well and what doesn’t, right?
I don't see how "trust" is anthropomorphizing. Do I trust this bridge to hold my weight? Do I trust that this tool in my hand will perform as expected?
I don't understand what trust means in this context. Even if I were able to hire Donald Knuth to write all my code, I wouldn't "trust" it to be bug-free, let alone to be the right fit for my needs.
> Donald Knuth's code would be much more likely to meet my standards during a code review than some LLM output.
Just to nitpick - will it actually? I don't know what your relationship with Donald Knuth is, but to the extent I'm familiar with him and his approach to work, I would expect that even if I could afford him, I would not be able to get him to agree to accept my coding standards, whereas I found that LLMs (although flawed in many ways) can be cajoled relatively effectively to adopt my particular standards.
Indeed. LLMs produce truly atrocious code, unmaintainable and unreliable. If you're vibecoding a toy to amuse yourself or something similar low-stakes, that's perfectly fine! For higher-stakes code, it's definitely not.
The article's awkward opening statement proves it wasn't written by AI.
I have been interested in machine-assisted ways to do and teach mathematics from as far back as 1999, when I started coding several applets in Java 1.0, both for my complex analysis and linear algebra courses, to visualize various mathematical objects I was interested in (such as honeycombs or Besicovitch sets).
It’s very much Terrence Tao style. His style is having long sentences that could have been broken down into shorter sentences but he chose not to. It doesn’t really affect reading comprehension.
I am far from a mathematician but I am excited by the possibilities of using AI for generating more math. Math in my mind exists purely in the world of forms, and cannot be appropriated for profit, but is downstream to everything else. I am keen to see what this enables.
I always enjoy these "domain expert has fun using AI to do something in their domain" articles. But it's always a hobby project, never something serious.
His website using mathematical knowledge is refreshing. There's a small UI bug, but personally, I wish more educational materials were this rich in audiovisual content.
Many visualizations that I have always wanted but just didn't have the time to build, I now have.
To give an example, I wanted a simplified 8-bit computer to complement the 16-bit teaching computer I use and designed this in a few days with the help of claude:
There is infinite latent demand for software, most especially outside the traditionally software-focused spaces. If LLMs stopped improving today it would take us 10 years to catch up to the new software-writing abilities that have become available. This is a great illustration of that fact.
Running legacy educational Java applets, especially around math and physics, has been a longstanding popular use case of our CheerpJ Applet Runner extension, running Java bytecode in the browser via WebAssembly.
I am not sure how to feel about agents solving the problem via proper modernization. It's certainly positive that students will be able to interact with this content in a modern and more accessible way, but the educational use case for our product, although not commercially important, has always been a source of pride.
even though there's still a lot of work to push things over the finish line, i have enjoyed how much it has reduced the activation energy for starting and finishing "one of these days..." projects!
It's probably a matter of short time until it's possible to disassemble any sophisticated software, rewrite it entirely with better features and usability, generate all needed artifacts, port to any platform. The only moat left is probably remote massive data storage. So if you want to replicate YouTube or TikTok, it's not impossible, but requires a lot more hardware assets than say anything that runs entirely locally (like operating systems or most video games).
I wonder if LLMs modify their output when they realize they are interacting with a famous person.
By famous I mean someone whose biography is in the training data. All models know a lot more about Terrance Tao than they know about me, when he's working on his projects do the models know they don't need to explain "Besicovitch sets".
Since the system prompt likely includes something about not insulting the user, does the LLM modify it's responses if it realizes it's talking to famous politician, like "dont mention the time $politician was cancelled".
47 comments
[ 6.1 ms ] story [ 88.6 ms ] thread[0]: https://en.wikipedia.org/wiki/Martin_Hairer
[1]: https://www.hairersoft.com/
"as such [LLM-coded interactive] supplements are not mission-critical to the core of the paper, I again feel that the downside risk of using guided interaction with LLM agents to generate such visualizations is acceptable."
It's a tool. Good for some things but not others and generally not to be trusted.
There are many AI bulls who adamantly disagree and cite Tao’s statements about LLMs for mathematical proofs as an example of how advanced and autonomous these systems already are
I agree completely you always need to check the work of LLM agents, but it does strike me as a tiny bit funny to anthropomorphize AI by using ‘trust’ while warning against anthropomorphizing the AI by using unchecked output. ;) Generally speaking, “trust” in AI has been going up very quickly as the models & harnesses improve, and as people figure out effective workflows.
I trust my hammer with nails but not screws… does that mean the hammer should generally not be trusted? The problem with AI is we don’t know the difference between nails and screws. (This may be where my analogy breaks down. :P) But I feel like saying don’t trust it isn’t as helpful as saying something like you should expect to spend more time planning and iterating than before, and you should expect tot spend more time reviewing and checking output than before, and learn how to use skills and context and subagents, and learn to use AI on some non-production low-consequence projects first. Saying ‘generally not to be trusted’ implicitly suggests not using AI, and doesn’t leave the reader with how to use AI. The goal is to build trust by building good workflows and by understanding what works well and what doesn’t, right?
Donald Knuth's code would be much more likely to meet my standards during a code review than some LLM output.
Just to nitpick - will it actually? I don't know what your relationship with Donald Knuth is, but to the extent I'm familiar with him and his approach to work, I would expect that even if I could afford him, I would not be able to get him to agree to accept my coding standards, whereas I found that LLMs (although flawed in many ways) can be cajoled relatively effectively to adopt my particular standards.
https://www.reddit.com/r/mathematics/comments/1tryyw7/terenc...
I have been interested in machine-assisted ways to do and teach mathematics from as far back as 1999, when I started coding several applets in Java 1.0, both for my complex analysis and linear algebra courses, to visualize various mathematical objects I was interested in (such as honeycombs or Besicovitch sets).
https://htmx.org/essays/universities-and-ai/#demos-visualiza...
Many visualizations that I have always wanted but just didn't have the time to build, I now have.
To give an example, I wanted a simplified 8-bit computer to complement the 16-bit teaching computer I use and designed this in a few days with the help of claude:
https://bdp.cs.montana.edu/
I am not sure how to feel about agents solving the problem via proper modernization. It's certainly positive that students will be able to interact with this content in a modern and more accessible way, but the educational use case for our product, although not commercially important, has always been a source of pride.
https://chromewebstore.google.com/detail/cheerpj-applet-runn...
By famous I mean someone whose biography is in the training data. All models know a lot more about Terrance Tao than they know about me, when he's working on his projects do the models know they don't need to explain "Besicovitch sets".
Since the system prompt likely includes something about not insulting the user, does the LLM modify it's responses if it realizes it's talking to famous politician, like "dont mention the time $politician was cancelled".