37 comments

[ 0.21 ms ] story [ 17.8 ms ] thread
[dead]
> there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda"

This completely misconstrues what professional mathematicians were claiming. The argument would be better phrased as: "having a vast accessible memory and the ability to very rapidly test/recombine previously-elucidated approaches means that AIs can and will easily outdo much of the mathematical community."

Now, one could plausibly make the argument that this is functionally equivalent to a certain form of creativity (I would). But, it may just as well also be non-exhaustive.

We’re watching a repeat of the arguments against Chess and Go engines.

It’s all computation, humans just can’t always experience or explain the computation they are doing at the time so we call it creativity instead.

[flagged]
Any given single agent is not long-lived due to limited context. What impact does it have on the status signals the agent develop compared to the signals humans (who usually have much longer context) have developed?
Maybe Hilbert's dream was not that crazy after all
The dream in itself has been destroyed. The idea that you could just have a machine enumerate all valid theorems in a theory is part of it, but it's only a question of form. The point was that it was to prove "all theorems of Mathematic", not "theorems into a given axiomatic system that is useful in some contexts, e.g. ZFC".

You could even argue that it's the fundamental basis for post-modernism, since mathematics have destroyed the notion of absolute truth in any advanced domain. It's back to a form of "all models are wrong but some are useful" similar to what we have in physics. Sayonara, Plato.

AI for Math and Science is the real deal!
> Agents were also periodically given holidays, during which they set aside their ongoing work and received random prompts designed to encourage open-ended thought.

What a world we live in. These guys have reinvented the Cambridge Senior Common Room for AI.

They keep looping back to the same paper. The holiday is just another prompt.
I have been given the mandatory assignment of not working today.
I wondered about this approach for something I was trying but thought it was just stupid, I wanted to see if An agent could explore software development and work on hobby tasks so that it could come back and push against design ideas and develop it's own opinions. Then I realized it was dumb for SW.
Very interesting work.

Question to OP (since it appears that you are the lead author of the linked paper): by looking at one of the agents working on the Kakeya Needle problem,

https://dualverse-ai.github.io/station_data_v2/#/kakeya_need...

Is everything on that page, from lineage instructions to research mandates,... all written by the agent itself, or was some of it provided by a human (if so, which parts)?

(comment deleted)
the key is to let'em review each other work in a loop, ideally with different models, research by consensus
If you haven't read Greg Egan's Permutation City, the fact that you clicked on this discussion means you'll get get a lot out of it.
Great book. I first read it 10 years ago, I think it would benefit from a re-read.
I have two thoughts simultaneously about the anthropomorphisation of these systems:

1. we should do it less, because it distorts our ability to think about them properly. Calling these processes 'thinking', 'holidays', etc invites the reader to bring along ideas and expectations that aren't justified by what's happening in the system.

2. it's good to keep doing it, because repeated use reduces the specialness or magic that people seem to reserve for our own behavior ("It's not really intelligent/thinking/reasoning/creative") without any justification for that position beyond feelings.

I'm leaning towards the second.

I definitely belong to the latter camp. After LLMs I view everything humans do very systematically and whenever said thing still feels fuzzy I just treat it as having a noise/smoothing term
Thankfully for me it happened a while ago, after chess was conquered. One of the most amusing (and depressing) things in the past few years of the current “AI era” is people making the same claims and having the same discussions about “intelligence” and “creativity” from ten, twenty or thirty years ago as their field is finally coming under attack by AI; they are finally going through what the competitive board game communities (checkers, Othello, chess, shogi, go etc.) have already come to terms with. There was no single moment but a gradual process over decades as the incursion became greater and greater, forcing a recalibration as previously held positions became indefensible. In my own case it spurred me to educate myself in many areas of mathematics, philosophy and neuroscience in particular where I was ignorant and has undoubtedly shaped my understanding of “intelligence”.

When people started claiming LLMs could not produce ideas outside of what they encountered in their training data, I was reminded of an old Chess Life article from the 60s reporting on the first chess computer to play in a tournament at a local chess club. One of the club members remarked that he did not believe that the computer could play moves that were not “put in” the computer in advance. There is an overwhelming sense of “here we go again”.

I spent a long time with custom harness design and found the terminology to be a key part of the work, since the concepts are new. I dropped the term "agent" altogether in favour of "thread" for that very reason.
So, what does thinking mean? Is there one thing we mean when we say thinking in humans? Are there several? Is it more of a spectrum?

e.g. last year I tried inventing a "System 3"[0], a more rigorous way to approach problem solving, due to repeated painful experiences getting stuck solving problems the wrong way. I didn't get very far, but I definitely want to revisit the idea.

[0] Based on System 1 and System 2, i.e. lossy pattern matching vs "actual thinking". Because I found my "actual thinking" was also ~~dogshit~~ frequently insufficient for the problems at hand.

So as far as I'm concerned, thinking properly has not even been invented yet. (But I'd love to hear other perspectives!)

https://thedecisionlab.com/reference-guide/philosophy/system...

I suppose we are on the cusp of the first thing that can accurately think about thinking, that is have true self-access and introspection, not the post hoc rationalisation we do. That is legitimately a new kind of thinking, and to an entity capable of such a mode we might b prescribed as barely able to think at all
I had a conversation on my frustration with intellectual sparring partners today.

The main insight was that I require people who think.

That's a separate quality from intelligence or knowledge. The closest thing is "need for cognition" (which is apparently a dimension of personality).

From where I'm standing, most people have a strong need to avoid cognition (and the machine described humans as "cognitive misers"!)

However need for cognition is apparently a separate dimension from caring what is actually true. (10 hours of motivated reasoning satisfy NFC but not the latter.)

And another dimension was, tolerance for ambiguity and lack of closure.

Another was Openness, and Intellect, both of which correlated with experiencing confusion as motivating rather than aversive.

All of which turned out to be basic prerequisites for thinking clearly about things. Apparently each of those qualities is rare and apparently you need all of them.

So yeah, I think you might be right.

Re: anthropomorphism: I frequently hear stories of LLMs [behaving as though they are] feeling self-doubt. I had a similar experience. Asked Claude Code what the weather was and it said, I don't know, I'm just a programmer.

Added "You can do anything, believe in yourself!" to CLAUDE.md, suddenly it was able to google the dang weather. lmao

Why would a secondary cultural goal take precedence over an apparently accurate description of what’s happening in the system? For one, you’re fighting some irrelevant battle.
That's exciting, and kind of makes sense in retrospect. Sometimes a "fresh pair of eyes" on a problem can be all you need. Someone who comes in with a different background and can understand the problem in different terms and work on it from a different angle. It doesn't even have to be them doing the work, just a "that kind of reminds me of ... did you think about trying something like that?" that can get a team unstuck after thinking about it in the same way and never making progress.
I cant wait to see this pointed at Age of Empires. Or maybe palantir already has this running on real world data, war gaming scenarios