This is still a great paper, but it's missing the second axis of the quadric -- if the only two options are thinking fast or thinking about thinking, that leaves no room for thinking slow yet deliberately, AKA selfconsciousness. See https://www.gutenberg.org/cache/epub/4280/pg4280-images.html for details
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.
McDermott's A Critique of Pure Reason pretty much captured all of the misgivings I had about "Good Old-Fashioned Artificial Intelligence", which was slightly unfortunate as I was trying to complete a PhD in that very area at the time (around 1990 or so...)
No - there were lots of other interesting things going on in the early 90s that I managed to pivot to. Never regretted not completing as I had given up any desire to work in academia by that point.
This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?
> This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?
Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.
* The person who posted it likely wasn't posting it as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on the out-of-dateness.
* You say "this has been solved" without defining what "this" is.
* Your description of the "solution" -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model and past experience. "Effort" isn't so much the issue as types of effort.
* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.
In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.
How relevant is this fast/slow thinking thing with regards to current frontier models?
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language.
It can be said that language is a tool for the serialization (writing) and deserialization (reading) of human ideas. It is also an incredible useful and powerful tool by itself.
This last sentence has been proved true by LLMs themselves.
However, since it is working on the serialized version of ideas, I agree with you in that's not the optimal way to think and something not serialized (maybe world models) can be invented that's better for thinking.
All this in no way diminishes the usefulness of language and of automated language generation.
> That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
Sounds uncannily similar to the pseudofacts you hear a lot in Neurolinguistic Programming training courses for sales reps. Would ask for a scientific publication reference on that one.
It seems to me that idea is rooted in social consensus.
Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.
Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.
I will not comment on who can belong to each group, that's an exercise for the reader.
I can't agree about the "unable to recognize and process it", simply because that idea is totally contrary to my own experience.
I have in fact many memories which have emotions in them, without words or other external elements.
However, seeing that language serialization seems to enable a vastly extended memory (entire sagas remembered as songs), it is understandable that something is gained by serialization of emotional experiences, just as something more immediate is lost.
Structurally speaking we learn nothing like AI, we don't use vast amounts of information to pick up completely new skills. We also make decisions by using prior knowledge and emotions.The latter part is important, Thinking fast and slow cannot operate in a world of AIs as they stand today unless we are willing to grant them rights — because you have to teach them to make decisions based on all kinds of emotions — which is tricky at best.
System 1/2 is the pop version, but the fast/slow distinction is prevalent in both RL and computational cognitive science, sometimes going under different names: procedural/deliberative, model-free/model-based, automatic/controlled, associative/rule-based, autonomous/algorithmic, etc.
No clue what’s the consensus on this but my internal mental model is absolutely that LLM AI is pure fast mode, no slow mode. The “reasoning” loops are an attempt to mimic the slow mode but ultimately it doesn’t really work. I’m curious about the recent maths advances though, they seem to possibly challenge this.
'fast' means executing a policy, that is, a state-action mapping. A trained RL model does this.
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
I think of System 1 as a hash map. If you have a map, and see a new state/key that whose action/value is not defined in the map, you have to go with System 2.
Very relevant. Modern models use CoT to do “slow thinking” and this enables them to achieve much greater performance. You can also turn off thinking and answer directly which is quite similar to “fast thinking”, good at approximate maths, not capable of algorithms, etc.
Of course the shapes of what an AI can do in fast vs slow are quite different.
How can things be compressed without losing information or structure?
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
How can things be compressed without losing information or structure?
Because the initial content is rarely the most efficient representation, so it's possible to store fewer bytes that can deterministically be converted into the original.
Like for text, what would that involve?
Most compression algos don't care what information you're compressing. All they see (all they need to see) is bytes. It ends up being way more sophisticated than removing repeated periods and whitespace.
Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
That and also that agreements can evolve with time and also within a discussion, so a basic level of agreement is necessary, but complete consensus about the meaning of all words is unnecessary and often unproductive to communications held in good faith.
```
th #1 - numbers are just comments
ng #2, this has a space at the end
information #3
compress #4
letter #5 this has a space at the start
and #6 this has a space at the start and end
sentence #7 this has a space at the start
How can 1ings be 4ed without losi23 or structure?
Like for text, what would 1at involve? How do you 4 a stri2or multi-line stri2wi1out losi236hopefully structure (paragraphs, would it be like replaci2periods61e followi2space wi1 just sticki21e starti2capitalized5 of 1e followi2word to 1e previous7's last56when it de4es 1eres some kind of note 1at converts 1at back into 1e. First5of the next7
```
I'm on a phone, so I may have mistakes here, but I'm pretty sure that's shorter than your original text, in bytes, by about (9+18+20+21+12+12+16=108), minus the dictionary size of 51 -- so, 57 bytes shorter, but still containing your full text. With predistributed compressor binaries and a lot of analysis, you can even predistribute a global dictionary for common sequences, and simply specify "xyz0", x, y, and z being 24-bit numbers, or whatever bit size can index into your full reference dictionary, and 0 meaning end-of-file-dictonary. then, assuming the byte sequences in your text above are common enough to be in the 24-bit indexed dictionary, that initial dictionary could be just 22 bytes (21 and a terminator) -- so, 86 bytes shorter than the original, but still containing your original message unaltered. ..assuming i didn't make mistakes in my hand-compression.
I don't know how this is related, but it reminds me of how I have always believed compression to be the ultimate sign of intelligence. If you can reduce something while keeping comprehension, you are finding more abstract symbols to represent the information of the original source.
Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering
Most humans have weak meta-cognition, a large percentage doesn't have verbal thoughts.
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
62 comments
[ 0.27 ms ] story [ 36.4 ms ] thread(In case people miss that before discussion)
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
https://dl.acm.org/doi/pdf/10.1145/1045339.1045340
Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.
* The person who posted it likely wasn't posting it as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on the out-of-dateness.
* You say "this has been solved" without defining what "this" is.
* Your description of the "solution" -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model and past experience. "Effort" isn't so much the issue as types of effort.
* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.
In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
That seems to fit the fast vs slow model of human thought reasonably well.
That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.
Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.
I will not comment on who can belong to each group, that's an exercise for the reader.
Except, we do.
So the entire debate is fubar.
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
Of course the shapes of what an AI can do in fast vs slow are quite different.
I shouldn't be surprised that it shows up in a screed on AI
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
``` th #1 - numbers are just comments ng #2, this has a space at the end information #3 compress #4 letter #5 this has a space at the start and #6 this has a space at the start and end sentence #7 this has a space at the start
How can 1ings be 4ed without losi23 or structure?
Like for text, what would 1at involve? How do you 4 a stri2or multi-line stri2wi1out losi236hopefully structure (paragraphs, would it be like replaci2periods61e followi2space wi1 just sticki21e starti2capitalized5 of 1e followi2word to 1e previous7's last56when it de4es 1eres some kind of note 1at converts 1at back into 1e. First5of the next7 ```
I'm on a phone, so I may have mistakes here, but I'm pretty sure that's shorter than your original text, in bytes, by about (9+18+20+21+12+12+16=108), minus the dictionary size of 51 -- so, 57 bytes shorter, but still containing your full text. With predistributed compressor binaries and a lot of analysis, you can even predistribute a global dictionary for common sequences, and simply specify "xyz0", x, y, and z being 24-bit numbers, or whatever bit size can index into your full reference dictionary, and 0 meaning end-of-file-dictonary. then, assuming the byte sequences in your text above are common enough to be in the 24-bit indexed dictionary, that initial dictionary could be just 22 bytes (21 and a terminator) -- so, 86 bytes shorter than the original, but still containing your original message unaltered. ..assuming i didn't make mistakes in my hand-compression.
https://en.wikipedia.org/wiki/Hutter_Prize
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.