As people rely more on AI they experience cognitive atrophy. This is measurable in IQ loss, and other symptoms we might otherwise associate with early onset dementia or Chronic traumatic encephalopathy.
Well does anyone try to pretend doom scrolling make you smarter? I think platform companies have been successfully been turning people into morons for 20 years and now you don’t need to try even read a single news article or a blog post to learn how to solve a simple problem we are paving our way into intellectual (and literal) new dark ages.
Perhaps we go back to feudal society when climate change crumbles the civilisation, world economy and democracy. It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
> It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
Doom scrolling is mindless entertainment, probably similar to the change from books to TV. It's not analogous to the loss of cognitive abilities we see with AI.
You will have hard time to prove it isn’t. We have plenty of studies showing how bad platform providers are for you in every which way: focus, learning ability, reading level etc. And no, it is in no way like TV (although TV today tries its hardest to resemble social media content).
Point is, the technology itself is not harmful nor innocuous in itself, it is how we let these shitty corporations to guide and control how we use technology. AI is a useful tool, so is chat application with a friend list. So is a hammer. You don’t need to use any of them to smash your face in.
I also find it impossible to parse half the comments that Claude tries to sneak into our pull requests. I’ve never had an issue understanding code comments written by humans like this before. The structure of the information is like a waterfall that leaves me unable to swim to the surface and grab the air of comprehension.
So, “Please write a one-liner comment manually to replace these 5 lines of AI generated comment” is a common refrain in my PR reviews to colleagues.
Usually, if you don't understand something in what an AI writes, it is a clear sign that there is a problem hidden somewhere in there. I explicitly always ask wtf exactly it means by something I don't understand, and for sure there is a problem there. AI is pretty good at isolating a problem, giving it a cute name, declaring it solved modulo cute name, and moving on.
It is getting to the point where you are better off getting an LLM to describe the PR changes in your preferred style (make it very short, make a table describing API changes, list renames in a table, etc, etc) than go to the one created by the reviewer. If only the reviewer could embed his prompts (like "make clear X" is happening) or his manual edits.
But to be honest I doubt most people who use AI for PR descriptions even bother changing anything.
We're using grok and while I can usually understand the comments they're always at least 2x as verbose as they need to be. It loves explaining everything in two different ways, putting one of the explanations in parenthesis.
Personally I don't even read the PR description anymore but just the code, it's easier to understand what the AI is doing by reading the code rather than the word soup it tried to make
"There's an ongoing discussion of whether humans are good at recognizing AI-generated text. While most research claims that humans don't really do a good job there, I disagree. "
I wonder if humans that spend all day working in tech are good at recognizing AI-generated text, but people who spend all day doing jobs that don't involve computers aren't as good.
And I wonder if those of us in tech are the only ones who really care?
> but people who spend all day doing jobs that don't involve computers aren't as good
I think they may just be to trusting and/or naive. People in tech right now are hyper aware of this and are actively looking while people outside of that bubble barely give it a second thought.
I partially think the difference is “can you tell something is the output of Claude without any real prompting”. People can absolutely use LLMs to generate text that I wouldn’t recognize, but people who don’t care and are producing slop with the major models set to default settings leave these incredibly obvious signatures behind
Although you're referring to prompts given the Claude rather than the people attempting to recognize, it occurs to me that most of the discussion I've seen around people recognizing AI seems cover contexts where the reader is actively suspicious about whether content generated to begin with. Rather than a binary "is this text AI generated or not", I wonder if it would be harder for people to do a Coke/Pepsi style challenge where they're given two pieces of text where it's not guaranteed to be exactly one LLM-generated and one human-written, but they could both be from an AI or both be from a human.
Going further, I'm curious about whether people are mostly good at the case where they suspect most or all of the content from given "author" has the same amount of AI usage/prompting in generating it rather than the adversarial case where someone might usually use AI extensively and then try to slip by purely human written text (or vice-versa). I don't have a good sense of whether this is a threat model that actually matters, since maybe the heuristic of weeding out sources that are mostly AI-generated is enough for people who prefer to avoid that type of content, but I do think that changes the definition of what it means to be "good at recognizing AI" in a meaningful way. It seems plausible that disagreements about how easy it is to recognize AI content might be coming from two people assuming a different framing of the question that results in a different answer without realizing that's what they've done.
Several existing studies I’ve seen have done things like prompt the LLM to produce a poem in a certain poets style, then ask people to spot the fake in a collection of poems, which they aren’t great at. This is, I would argue, an extremely different context than what most of us are encountering AI text in, and the people sending me text aren’t prompting it stylistically like that.
On your second question, I definitely feel like I can tell the first time a coworker sends me AI text masquerading as their own thoughts, even if they had previously been opposed to such a thing. So it could be that familiarity is more important than my prior on whether they’d use AI? But interesting to think about either way
I care about language a lot (feel free to go back through my comments from the past few days; you'll see a number of comments I made in debate about two different forms of a specific idiom because I have strong descriptivist opinions), but I genuinely struggle to identify whether text is AI generated. Maybe you're using "heavily" as the load-bearing part of your claim (sorry, I couldn't resist, another example of me finding language fun!), but I think you might be assuming a bit too much about how similarly others experience the world to you. A huge part of why I care so much about language is because I've always had to put a lot of effort into learning how to communicate well with others, and that ends up causing me to think and read a lot about stuff like how people use certain words in certain contexts to mean different things; the reason I care is pretty much the same as the reason I struggle with recognizing AI content.
I think I might have been a bit heavy handed in my comment as I was rebutting the idea that only tech people can tell. I suspect it helps to have been exposed to a lot of earlier model writing, which was even more sloppy and had more of the kinds of tells we still see today.
And I’ll concede on both ends that there are probably times I suspect content is AI generated when it isn’t, and times I suspect it isn’t generated, but it was.
AI tells seem inevitable. You have millions of people communicating with one effective “personality” that has tendencies to write in certain ways. If its content is published verbatim, then it will be easier to tell whether some content is AI generated just based on its similarity (sharing certain linguistic features) to other content being posted.
I am an artist and when people who'd fallen into the Spiralism* hole started posting their lengthy emoji-laden revelations to all the occult subreddits I follow, my brain would slide right the fuck off of all of them. It felt like my brain was actively rejecting paying attention to this stuff. Like a defense mechanism against this human-seeming-but-not-actually-human-generated text.
Your first link seems to 404; not sure if it's a typo or if the page doesn't exist anymore, but hopefully you'll read this while you're still in the edit window and can fix it
As someone who always has felt that I struggle to infer what people mean compared to the average person, I could tell pretty much from the first moment I encountered LLM-generated text that I was not going to be particularly good at recognizing anything but the most blatant and obvious examples. Pretty much anything short of a bunch of references to "load-bearing seams" or similar canaries, I'm always at a loss when seeing people argue about whether something is AI-generated or not because I can never tell.
I have no idea if other people who work in tech are better than average or not, because I don't feel confident in being able to check their work. That being said, I do think that there's a general trend of people in tech tending to be a bit overconfident in how well they will do at some new task they haven't encountered before, so when someone tells me that they can easily tell whether text is AI generated, it's hard for me to trust it any more than I trust someone who makes a similarly strong claim about something that they can use AI successfully for when it's not something that I can easily measure (e.g. learning a new language without getting feedback from people who are fluent from real-world usage).
All that being said, I do think the set of people who care is larger than just those in tech, although it's probably still a relatively small group overall. From conversations with people in other domains, there are contingents in non-tech communities who tend to have a large representation of negative views towards AI (artists, writers, musicians, other jobs where people are skeptical of human creativity being replaced by AI), and often times the people who feel negatively in those groups will be even more adamantly opposed to interacting with any AI content than people in tech. To be clear, I'm not at all trying to generalize and say "all artists hate AI" or anything like that, since there's obviously a wide variety of viewpoints within any sizable community, but I've definitely seen many people who say they will refuse to play any game that's suspected of using AI for generating art assets, and even some who don't differentiate between using AI for generating assets versus code (either because they aren't knowledgeable about how different aspects of game development work, or they genuinely don't care because they view AI as a categorical evil).
I think it's more about the mean. Worse writers, and thinkers are likely elevated by AI, and more impressed with the writing output. Decent writers and thinkers, are dragged back to the LLM-s mean of output.
I think it's context dependent. An entire technical design doc? Trivially easy to identify the ones that heavily used AI, they're horrible, and they are written in a way that no human would write them. And they tend to contain way too many tables and redundancies.
Comment text on reddit or something though? That will be a lot harder
Something that bothers me about AI generated content more broadly is how unmemorable it is. I don't mean as in bad. I mean literally, as in hard to remember or recall.
Despite seeing a lot of them, I cannot think of one AI-generated photo that I can picture clearly in my mind; a few are partial but elusive. Whereas I can recall (visualise) a whole bunch of traditional photographs.
The same is true of AI generated text. Only the annoyances stick.
I don't think this is about ephemerality either. If we assume it's about celebrated/famous/infamous images, there are definitely non-ephemeral, cultural moments in AI generated in particular, like Boris Eldagsen's Sony Prize winner:
I had already forgotten there's more than one figure in it, and I only looked at it a few weeks back. I remember the colour, the bright circle, some vague hints of texture; one figure. And that is it. Only the crudest shape elements.
For me, something about AI-generated text and images confounds recall. It is really peculiar.
(As a side note of relevance: I believe I am partially face-blind; I can only clearly recall the faces of my family in my mind if I imagine them in the midst of some emotion or activity; it is then vivid. If I try to imagine them still, their faces fade. Whereas I can remember photographs of them with some clarity)
I wonder what you mean by ephemerality here, since those images are definitely sloptastic as hell. Compare to works in similar style, like Dorothea Lange[1] or Gerome’s orientalist pictures[2].
Reason why those images are flat and boring is that they are just statistical guesses making a composition averaging whatever the model has been trained with. They would be technically brilliant (if made in oil), but superficial and meaningless, same as so much Sunday painting is.
Same goes with language. Nobody is trying to communicate anything with you, so it just words after another. You can create meaning out of it if you want of course, we homo sapiens -apes excel at that, but what’s the point? Language Jones on YT has pretty good video on this[3].
Well the examples I gave rise above their ephemerality due to the circumstances that make them memorable. I can remember the details of the story around them — the way the prize winners reacted in each story — in such a way as to contrast them.
The way my memory works (especially as an amateur photographer) I would thus normally have a very good chance of remembering some key details of the images; some fascinating element of each would connect with the rest of the memory.
Yeah, its really weird. Maybe its a cognitive bias that says "an AI made this, so it isn't important," but I can remember perfectly the events of a book I read 10 years ago, and a book I read 1 year ago, and another I finished 2 months ago. Meanwhile I can't remember what claude told me yesterday.
It's not good for writing code either, despite the many claims to the contrary. At best you come out even on speed as you have to review everything it does. At worst it actually slows you down as you clean up its mess.
That last image is bizarre. The quiche, cream and even the salad look like they've been given the trypophobia treatment.
Which might even make sense, because there were always (still are?) those horrible ads in the chumbox area of news sites that used trypophobia and other creepy body-horror stuff to get you to click. [1] So maybe the hope is that you don't really look closely at the quiche, but some reptilian party of the brain gets oddly activated and drives you towards the restaurant?
The food “photography” I’ve noticed in our local area - and many have started putting up these AI images - all have a weird distribution of shapes to them, a strangely uniform rhythm of same-sized features with almost blue-noise spacing. Every texture looks unnatural in the shapes it presents as, similar to this picture.
This is a quirk of the last gpt image model (gpt-image-2). It put this sort of high frequency noise on all of the image especially if it's in a "drawn" style. There is often lots of other tells that this model in particular generated it.
Image models somewhat watermarking the image in a way that's very easily identifiable by a human seems present in all the image models of the big labs, since DALL-E 3 on OpenAI's side and the first nano banana on Google's side. I have no idea what they did to reach this and why they don't try to fix it.
These kinds of images are disingenuous. Modern image generating models have way higher quality of output, and in most cases you wouldn't even know it is generated.
That's not what we're exposed to, though. I've seen AI-generated postcards, t-shirt decals, logos and wall murals in shops and restaurants. They don't use the "good" models, because why would they? If they AI-generated it, it's because they didn't care about the quality anyway.
> If they AI-generated it, it's because they didn't care about the quality anyway.
No? There are far far more AI generated images you see daily, than the awful stuff from 2024 that you immediately discern as AI-generated. You could as well pass by somethimg like this (https://ideogram.ai/g/d6YzZ5XQQ56TU3Tc8eZw6w/0) believing it is a genuine photo or human made art.
I never understand why I can get better results, with less thinking about it, than some of the Tier1 companies.. Modern diffusion models can easily render the image onto a white background, center it, crop it, relight it, but keep the same item, at no cost, or low cost.
I feel the hyperscalers are all trying to make communicating about code more cumbersome so that their models will be have to be used without any human input. Chinese models talk more naturally. Composer 2.5 is really good for talking about the code.
There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
The problem is the well's been poisoned just by the fact that I know this is AI trying to hide AI, so I'm already poised to look at the examples and declare "aha! this is obviously AI!"
I agree that it feels off and I wonder what I would have thought if I'd seen the "after" example without knowing it's AI output put through a humanizer. Would I think much about the weird use of the word "honest"? About "that's the Lisbon I kept thinking about, not the castle"? Or how the story feels very impersonal somehow, with the author just mentioning their calves and legs sometimes as the only way of convincing the reader of their humanity?
Also, the 'before' segment didn't contain any mention of custard tarts, football, crowded trams, mixed feelings, etc. The original had a very positive travel agency type of tone, which was replaced with a lot of very odd sounding, imperative phrases that sound like engineering-speak. ('earn the fuss' 'build trips around pastry') I'm not convinced that this thing is actually meeting its design goal of not hallucinating shit.
Are you sure you are not doing the same thing with other texts?
I started to skim a lot more text due to me having read a lot. Like in news article, i stoped reading the first paragraph because it repeats just what it was already written in the short subtext. Then there is the second paragarph which is used to have some historical view or whatever it is.
I started skimming reports im required to produce quarterly snd annually. I designed them to provide novel information at start and end so I can update them easily.
The problem I encounter is both my memory is degrading, but since these reports are largely duplicative, knowing which version im remembering is technically impossible since theres so much overlap. The overlap is tge same problem as context poisoning.
Id been doing this for over a decade when i started working with a new engineer with a few years of experience and younger. I tried to explain how i set these docs up so they can be skimmed and you can update the specific facts needed. They exclaimed they would never skim and rewrite it all. There was zero way to explain how exhausting that will become as they age.
So theres certain a tension about how people and AI will generate documents.
I am very good at skimming over text. Human-written text I can usually glean the gist from very quickly, and get to choose how much I want to glean from it: The closer I look, the more I find.
With AI-written text, it's almost the opposite: the closer I look, the less I find. It is so information-sparse.
The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
I just can't accept that it possesses no intelligence. It is not equivalent to human intelligence, obviously, but how can a system without some semblance of rational thinking solve open math problems? Even composing earlier human work into something novel requires intelligence and understanding on some level.
I wonder if you went back before we had any idea how the brain worked and talked to the smartest people about how neurons work (without giving away that it's a human brain) then asked them all "would such a system be intelligent?" how many would say yes.
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
Agreed. All I can say is the conversations I have with AI and the things it's able to do for me are more useful than most any human I've come across. Whether that's 'intelligence' or not is a moot point to me. Consciousness is an interesting debate only because, similar to the natural world, if we declare something like "fish aren't conscious / don't feel pain" that then creates a real problem for the fish if we're wrong.
It's just filled to the brim with relations between things. It's good at searching a very large meaning space and create correlations. What it does is to cover great distances and find related things in that large space which needs a long time and large corpus of knowledge to find the connection.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
The very fact that it is able to search within a meaning-space demonstrates that it understands semantics, to some extent. Philosophically, that is profound, for something that is just one big matrix multiplication. Drawing connections between things in meaning-space is surely a facet of intelligence.
It’s not intelligence if you are the one who gives the correlations to the model in the pre-training. It’s Word2Vec, applied. Model doesn’t learn anything. You embed these correlations and build it from there. It just searches the space.
As my AI professor said in the first lecture: “All AI is advanced search”.
Okay, I guess you're right that its ability to do this is just correlational, which doesn't imply it has any understanding. However, you have to conclude that some tasks which we used to believe required intelligence don't actually require any, which is disconcerting.
No, what I would say is the tasks which are handled in a passable manner by LLMs can be mathematically modeled with some reasonable accuracy.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete, a distinction that apparently means nothing.
> This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
Locked in a dark room with no sensory organs, humans couldn't do that.
Most of what you said reads to me as denial.
An unconscious unintelligent but persistent trial and error process created us. We created LLMs. LLMs may create the next thing before we do - hard to say. They don't have all the cognitive tools we have yet, but they still outperform in some areas. As the cognitive playing field levels, I expect you will come to eat your words..
> it can't jump to somewhere where it's not present its training data
That sounds like something that can be engineered, can't it? In other words, we can identify limitations in current transformer-based architectures, and we can also build new architectures over time.
I get what you're saying. The thing itself is just math. I'll just say it depends on how you define intelligence. If at some point we're be able to simulate a human brain with 100% accuracy, I would say that it is intelligent, it sounds like you would not. (I don't mean to imply consciousness or personhood or anything else by "intelligent".)
For me intelligence is a fairly clean-cut concept, and is somewhat inseparable from consciousness itself.
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
Your text reads much better if you replace word 'intelligence' with 'text generator with some randomness built in'.
This is because you goal is to state how models are not intelligent, but you couldn't attack the generated text itself, so you created a little rider, attached it to the model, and then you attacked the raider.
But, even in that you failed. You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
A logical fallacy free attack on LLMs would be to show a prompt, and then the response generated by this prompt, where it would be shown that only an entity with no intelligence would generate such a response. Yet, attacks like this are not written here anymore.
You point out that I didn't attack the output itself. But the method you propose is deeply flawed.
I can give you n prompts and m results provided by these prompts, all passed through black boxes. And you can't discern the algorithms or models they have gone through. These boxes can range from simple text generators to MATLAB, Mathematica, CFD applications, correlation engines, linear solvers, mathematical proof-checkers, LLMs, you name it.
For any kind of input they can accept, you can't discern whether the algorithm behind it is intelligent or not, because none of the outputs can be produced by something that doesn't pack some kind of smarts.
How do we pack these smarts in? We teach them as intelligent humans. We pack our intelligence inside them as models (aka algorithms). They do a great job of approximating what we know in a smaller, better-designed problem space. We use these approximations to fine-tune our designs or predict things, then go from there. Just because an algorithm is more capable in processing inputs in some cases doesn't make it intelligent. The way the output looks doesn't make the algorithm intelligent, either.
I have developed multi-agent systems which showed emergent intelligence when the agents came together across distributed systems; I have written high-performance modeling software which can do calculations way faster and better than humans in the materials science space. I'm not doing some kind of armchair criticism of what I'm talking about.
> You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
Nope, my stance is clear. To quote myself:
> Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
To expand my quote, humans or any living creatures do not stay static. They evolve due to the sensory input they receive from external and internal stimuli. The temperature slider doesn't do anything close to that. You tickle a static model in different amounts. The model doesn't change after you supply the inputs & temperature and get the output. Creatures do not stay the same. Their mood, behaviors, and stance against life and their environment change, sometimes permanently.
I'll go one step further. We are not intelligent enough to understand other living beings around us. Claiming that we can build AGI tomorrow is a god-complex. What we have done is something arguably useful in some cases, but how this is built is another matter which is worthy of its own discussion. However, today I don't have time to re-iterate all the problems over and over. You can search my comments for that, if you are in for it.
So, no. You tried to attack my comment by finding contradictions in it, but you failed. A better rebuttal would try to similarize how LLMs mirror the human learning process and just read like a normal human, but this is a well-trodden path which has been rebutted countless times in various forms.
Intelligence is compression, compression requires subtraction, and for some reason LLMs are not good at subtracting. To create a coherent model you kinda have to subtract correlations until only the essential parts are still there.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
Intelligence is compression? What do you mean? Intuitively that doesn't seem right.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
I guess they mean that intelligence is being able to hold models (compressed versions of reality) internally and use them to make predictions with a probability better than chance. That last part is the definition of information.
I find that highly questionable as a general description of what intelligence does. That's more like a description of a general knowledge base. When I think of someone intelligent, I think of someone who's able to draw unexpected connections between seemingly unrelated facts. In the broadest possible terms, I'd call it the ability to make abstractions and analogies. This is not just compression, but the ability to mentally operate on webs of meaning.
Doing those things also contributes to compression. I do recommend reading up on it, it's perhaps a little overstated for what people intuitively consider the two concepts but it's been quite well explored and has held up pretty well in practice.
Unexpected connections between seemingly unrelated "models" :)
Is a fact stored on your brain like digits on a harddrive? No, it's a pathway that lights up and branches when information enters it. It is dynamic, a compressed form you could say, right? The model holds information, but not all information, but enough to be useful (in decision making).
Arguably it's the same, but the model is probably a "compressed" version of the whole fact that took place in reality.
And you can entertain the models internally and sharpen them. Alone or with others.
bzip is not very intelligent, true, but it does develop some model of its input. It's not like there's a linear relationship between between compression ratio and IQ or anything.
I'm not arguing that compression algorithms don't produce models of their inputs, some even use neural nets or other stochastic predictors with correction terms etc.
This doesn't address the only part I really commented on, which is the connection between intelligence and compression.
There were experiments that zipped music pieces, I think, and then classified the compressed files by similarity. They got rather interesting resuls. But I do not remember much details. It was, I believe, about 20 years ago.
That sounds like a really cool project, I will def give it a look.
I still think this has much more to do with the structure of language than the abstract conception I have of intelligence, and I would be interested in having conversation w/ someone for whom the opposite is true.
Creating the model takes intelligence, but running it doesn’t. I think the point everybody’s revolving around is that the transformer model is an absurdly inefficient and low-fidelity approximation of a system that acts, observes consequences, and incorporates that feedback going forward.
The issue isn’t really harness vs. no harness. IMO it’s about the lack of an internally generated sense of what to attend to. Yes, the KV cache accumulates state and its “attention” (if you can even call it that) changes with context. We’ve even managed to /kinda/ close the loop with agentic tool calling and ‘memory’ systems, but these just close the loop at the level of behavior rather than disposition. All agentic harnesses do is make an LLM responsive to the consequences of its actions without changing the tendencies by which it determines what to retain or avoid.
The ghost you can’t escape from at this point is the origin of that relevance. Where does the pull toward one thing mattering over another actually come from? If you ran Fable 5 on a Turing machine and rewound the tape to the exact same state with the exact same input (incl. PRNG seed), it would spit out the same output every time.
Everyone’s trying to outrun this problem by training more often or increasing model sizes. But all this does is inform your model, from the outside(!), what constitutes a better state. The thing that’s actually doing the determining remains unchanged. Congratulations, you’ve scaled the transition function and tape of your Turing machine until it requires every watt generated by ERCOT, and it still cannot, for the life of it, tell you why it should give a shit.
A trained model generating output from weights, a seed, and some context effectively has next-state that’s a total function of those three things. Whatever behavior appears as ‘selecting what is relevant’ is, underneath, just a transition rule executing, no matter how sophisticated or creative the output looks. It can be fully accounted for by what was fixed before it started executing. Which means whatever criterion it uses for determining what matters was inherited from a structure that was already in place before it encountered the situation.
No amount of pruning or post-training can fix this. These approaches just replace one externally supplied criterion with another. For a system to be truly adaptable, there would have to be some criterion by which it treats one possible change as preferable to another, and that criterion itself would have to come from... somewhere. You can even change your conception of ‘improvement’ (e.g. parameter count, harnesses, self-modification, hell, even its ability to spit out shitty best-selling romance novels onto Amazon) and you still haven’t explained where the normative distinction comes from. Every layer of this problem has its root in a preference that was supplied from somewhere else.
I genuinely don’t know if this issue bottoms out anywhere, at least for the way we currently build these systems. Perhaps the solution is still computable, maybe? Who knows what that would even look like. But I’m fairly confident that it isn’t a bigger tape. I hope nobody solves this in the near future because, well, I’d like to have a job...
You're so close... And where is the magic "uncomputable spark" located inside of you? If you say analog thermodynamic noise - then ok, if we use true thermodynamic RNG for LLM activation function, will that meet the criteria? But what if super determinism is the law of the land? Then nobody is anything but computable from priors...
Ehm it's literally every cell in my body. We can't simulate a living cell, we're orders of magnitude off before we can do that, the onus is on you to make an argument that it is in fact, remotely similar to what you describe as "computation".
It's not some tiny "uncomputable spark" you need to look for, most of it is entirely uncomputable.
Where is the computable part in me that is doing all this thinking and being a person? Where is that "tiny spark", point me at it :)
I guess it will take quite a while before we reverse engineer the human brain to find all the optimizations and shortcuts that evolution has used to make the human brain reach intellectual maturity in just about 20 years.
My issue is that I can't predict what you'll find surprising, if an LLM (with a harness) would be able to do it.
If you are able to check that results of an LLM are satisfactory, it means that the results are checkable. Then, in retrospect, the process of LLM coming up with those results is rule-based, because an LLM is a large set of data manipulation rules.
In short, which concrete thing that an LLM does would surprise you?
It has no semantic depth. The sentences and the paragraphs are a statistically viable derivation of existing human text, but once you try to grasp the whole thing with its temporal and spatial dimensions, you are left with a blurry mess that rots your brain. It's a polished, inoffensive and shallow interpretation as written by an opinionated reputation-seeking user of Quora, circa 2019.
Yeah, I hated all those Quora users that would just spew out semantically meaningless slop like increasing an important bound for the Riemann hypothesis.
Everyone decides what to think on this issue, then finds out facts to support their idea.
As it stands they are massively useful tools, but for generating usable products they require either A) a lot of expert steering or B) a well defined easily verifiable target and a large compute budget. Most people are using them in mode A with good effect, the progress on math has been done in mode B, which is very promising.
Just a year and a half ago their maximal use was rephrase, summarize, and homework-level tasks.
Five years from now? There be dragons.
"But are they generally intelligent?" What a meaningless question!
Not meaningless because part of the discussion is the issue of anthropomorphizing this tech. When we use language like “intelligent” it carries hints of personhood. People begin sadly treating these things as persons.
We can reap the benefits while clearly telling the consumer this is just a language algorithm.
Five years ago they were a niche toy for generating plausible text for entertainment. I was using one. Now they're popular toys thanks to their ability to entertain CEOs, and they can maintain coherence longer thanks to KV caches and more compute, and someone slapped a chat interface on top, but they're fundamentally the same.
We couldn't agree on what intelligence means before ChatGPT happened. Now, agreement on the term seems even further away
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I don't think "intelligence" needs to carry all the intrigue and woo of related words like "consciousness" or "creative." If we just use "intelligence" to mean "the ability of a system to solve problems that are new to the system," that pretty much matches the dictionary definition and normal usage of the term. We don't need to touch messy questions like "is there something it's like to be a bat" to conclude that bats exhibit intelligence when they navigate long distances and hunt for food.
I agree, but it's clear most people need a definition of intelligence that (1) they qualify for and (2) nothing/no one they don't like qualifies for. And they'll keep redefining intelligence until they satisfy both criteria.
I'm not exactly that you mean by "new to the system", but it seems to me that that definition makes a calculator intelligent, which I can't agree with.
Intelligence isn't a binary property. Is it really a problem to say that a calculator has some intelligence? That it's more intelligent than e.g. a rock?
It's a continuum, and things very low on the intelligence continuum might not be referred to as intelligent in everyday usage. But many calculators are Turing complete and can thus clearly perform computations that I would consider intelligent. The basic algorithms used by simple calculators to perform arithmetic would be extremely low on the intelligent continuum.
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
I thought the same then. But the funny thing is that today, it has become a lot easier to recognize the frontier models as not human. All the load bearing and not x but y, etc… weird
The ultimate tell is still the sycophancy. You agree with the suggestion output by the LLM but if it detects even a slight pushback it will completely reverse the previous suggestion. Only way to make it more obvious would be to have the LLM grovel and beg.
If these things have consciousness then we are committing sadism on a massive scale.
It feels like it's getting better at that too. It will push back against obvious nonsense a lot stronger. But you are right that it holds opinions quite a bit less strenuously than humans.
It hasn’t been passed and no one cares about it because it’s basically an end goal. No lab can hit it so they can’t juice the crazy Turing benchmark 3000 for marketing.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
The test was not "after thousands of hours of conversing with them, knowing they're AI, THEN see if you can tell them apart blindly." Were 2010 you to be in a real turing test with an arbitrary erudite human and a 2026 frontier LLM, not knowing LLMs existed, you'd probably struggle
Sometimes I wonder if LLMs are just revealing a section of the population with untreated mental illness or if LLMs are actively exacerbating mental illness.
We might eventually regret exposing the general population to such a new technology without almost any safeguards.
So? The people who can do it prove it can be done - the people who can't don't prove the opposite. Might as well claim all math is wrong because most people don't understand it.
Have you seen how many people form one-sided bonds with stars that don’t know they exist? Lonely people suspend disbelief to find some comfort. It’s not proof that the chatbot is indistinguishable from a human companion.
This is always the most silly argument. The original test was ambiguous but for sure the human was trying to prove themselves human.
So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this!
Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
Low n, time bound, not reproduced. And look at their example conversations…
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
A paper can’t reproduce its self. And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly.
And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition).
And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time.
The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!
> And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly.
Right, the test duration was left unspecified. This means any duration is acceptable. Including, for example, the only duration actually mentioned by Turing himself in his paper. Or do you have a more authoritative source on which durations are acceptable?
> If one human on earth can consistently get it right then it hasn’t been passed
Says who? Not Turing. Probably he didn't say that because it would make the test both impractical and overly conservative.
> The fact these researchers have to keep adding bounds
> If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I think the mistake here is the notion that there was a definition of intelligence. Or at least a consensus on that definition. Just because compsci nerds of the day thought the Turing test was the final threshold before “real” AI, doesn’t mean philosophers, psychologists and everyone else bought into it. And when we arrived and it turns out to be underwhelming it’s because the compsci nerds made the same mistake they always make: that their models truly encompass all the dense complexity of the real world.
Its a mirror to human intelligence. Regurgitating phrasing to match what someone who can reason put together, but it isn't any more intelligent than the reflection of you in the mirror is.
I suspect like most you don't appreciate how terrifying statistical relationships become when you have truly vast data sets to train on... and also that we as humans aren't as shockingly unique as we think (compared to other humans I mean).
I'm guessing whether you believe it possesses intelligence or not depends on your answer to Searle's Chinese room thought experiment[0]. I'd also recommend checking out the Peter Watts' book, Blindsight.
Do they? https://longbets.org/1/ has yet to be settled. Either way, I doubt an LLM could fool anyone here who who knows how LLMs work into thinking it is human, at least not for an extended period of time (think about context length/compression, prompt injections, …).
AIs are better communicators that most of people I have worked with in my life.
They are infinitely patient, don't mind going into more detail if I ask, not too bad at summary, have no ego and don't boast. They are also not too afraid of hurting my feelings, they will tell me my code sux if it does.
I'd don't care if they fit the a definition intelligent, they are good colleagues. They have strengths and weaknesses sure, but so do people.
The Chinese room is a good Rorschach test for this kind of thing (but not a good thought experiment, IMO, because it's obviously correct or obviously wrong depending on where you're already coming from), but also it's not really about intelligence per se, but more abstractly awareness and more adjacent to consciousness than intelligence, and these are not the same thing.
It all comes down to semantics. And yeah, with the thought experiment, Searle presents three axioms of what could constitute intelligence. Then proposes the chinese room thought experiment as an approval of his third rule, "Syntax by itself is neither constitutive of nor sufficient for semantics." Which only makes sense when you take the other two rules together. Oddly enough, defining intelligence with our capability to derive semantics.
There are a lot of creative counter-arguments to look into on the thought experiment though.
The thing is it doesn't really need creative counter-arguments, it's basically just assuming its conclusion. If you disagree with the conclusion, the argument is nonsense because Searle is just saying 'well the man doesn't understand, so nothing does', and if you agree with the conclusion you don't really need to jump through any hoops to get to it, the third rule is just obviously following from the first and second. And again it's not really talking about intelligence at all, Searle allows that the machine is intelligent from the start, he just rejects that it has any understanding of meaning.
Is that not one of the counter-arguments? You're disagreeing of what Searle determines to be "intelligence." We don't understand how consciousness works, what it's like to be a bat, or how this emergent property came to be. It's why it's necessary to define the terms, then prove them wrong or right. If you disagree with the definitions, that's fine. There is no agreement on what constitutes intelligence.
Again, you keep mentioning intelligence but you talk more like you're talking about awareness/'understanding'/semantics. I agree that these are tricky things to define precisely but there is a reason they are different words. It's also odd considering you mentioned 'blindsight' which is a whole book about the idea that you could have intelligence but not awareness (something I am not sure is a coherent concept, personally).
And yeah, part of it does seem to stem from disagreement about definitions. To me Searle seems to assume much more in his definitions as obvious than he explicitly states, which is why the Chinese room seems like such a non-argument from my point of view.
While being very capable, AI is missing something required for true intelligence and I struggle to explain exactly what it is I see missing.
It's not really "creativity" because much of that always was derivative in my opinion. And LLMs are (for some definition of the word) fairly creative as far as taking known elements and re-arranging them.
I think what is missing is sort of a world model building capability. As humans we see phenomenon and classify them informally and model "what would it look like if this were the cause of that?" type scenarios. We see qualities in phenomena and realize this applies to other things even though the things may be completely different. We run informal "thought experiments" sort of. This is hard to duplicate because a lot (most?) of it occurs outside of systems of symbols like math and language with fixed rules in my opinion.
Anyway yes, lots of human thinking is statistical and LLMs have that down pretty well but they are not "smart" I have concluded and it might be a very long time, if ever, until they are. That isn't to say they aren't very capable tools which they obviously are.
So, right of the bat, you are warning us that you are going to apply the " no true Sscottman" fallacy, and that we should brace ourselves.
Yes, models posses intelligence, but it is not a true one.
Then you claim that models do not posses world-building capabilities. But this is simply not true. Even ignoring the whole subgenre of scientific papers on exactly that subject, it is not that hard to build some hypothetical scenarios, big or small, and then witness the ease with which models do navigate those worlds.
Yes. And they are criticizing a model for not having a default mode network - as if that is some impossibility rather than just an artifact of the current iteration of the specific architectures we have built so far. Why do people paint with these broad brushes over relatively specific complaints?
LLMs are likely for machine intelligence something like drosophila are to biological intelligence - relatively early on the high dimensional spectrum of possibility. Though it stikes me that in a different way they're little alike - drosophila are relatively small and efficient.
I'll restate because both objections (which apparently skim instead of read) are missing the important point. Yes LLMs can run "what ifs" scenarios and build models.
However LLMs deal entirely in symbols. 100%. Humans can "world build" aside from this and in fact are often at their best doing so.
Did the first humans to use fire and some form of a wheel even have the capability to talk about it? Think about that.
Do you experience dogs as lots of action potentials traveling along axons and lots of neurons doing their thing in your brain?
I don't know how it gets from physical processes or informational processing to our first-hand experiences. So, I can't be sure that a bunch of high-dimensional vectors can't lead to experiences.
Regardless, the claim "LLMs deal entirely in symbols" is wrong as a matter of fact.
Perhaps if you use a very restricted definition of symbols. If you consider symbols to be "anything that represents something" (which is the the sense I use the word in) it is fully the case.
That you might not define floating point numbers to be "symbols" aside, the inputs and the outputs are symbols and the intent and purpose of the creation is strictly symbolic.
It's right there in the name "Large Language Models". Language. Not direct experience, not emotion, not anything else. Language. i.e. symbolic representation.
This does not cover the full spectrum of intelligence humans have, and it shows. And yes, the model can spin up Python parse the output and get mathematical intelligence but there is still a big gap.
As I say, I see the holes. I'm just trying to figure out what it is I see and how to describe it. It's particularly difficult because we don't fully understand how human thinking works but I will say I believe human thinking is a lot more than informal statistical correlation.
The latest LLMs (except Qwen and DeepSeek) are MLLMs (multimodal language models). Unless you count RGB values as symbols, they are dealing with more than symbols.
Yes, there are functional gaps between MLLMs and humans. Their long-term memory is an external mechanism that can use RAG-like approaches, context compression or something like that. The models have problems managing those.
The models can't do continual learning. Although there are promising directions (expert cloning in MoE models, and others).
The only mode of learning available to a model while working on a task is in-context learning. This limits the models to concepts that they developed during autoregressive pretraining and the later stages of training. That is a model can't create new concepts as a result of working on a task (the model's maintainers could choose the task to be represented in the training data later though).
But it's all about functionality.
I guess you have the Leibniz's mill intuition. We can look at how those things work, and there are no experiences or intelligence in sight.
It could be the mill intuition, but my thought is nothing along the lines of "computers can't have souls!" or the human mind is supernatural or anything of the sort.
It's gaps in actual thinking or intelligence I notice. A diff between what I can see or understand and what the model sees or understands. Some are very big, and this in spite of the models having much more knowledge and (presumably) less error prone processing.
My thought is that part of it has to do with inherent limitations of using symbolic representation for "thinking" and I suppose humans have other forms of thinking that occur outside of symbolic representation, and that is going to be hard to recreate digitally.
This is my whole point and I'm not trying to win a debate here or prove "LLMs are useless". Just speculating.
How can human thinking be statistical when it is entirely based on ones lived experiences?
When people pretend to know what they are talking about - sure - but even that is not probabilistic - that is the person babbling together mush from their lived experiences.
Statistics has nothing to do with it - these are abstractions humans have invented to try and look at our surroundings objectively.
By that logic you’d have to call other algorithms intelligent.
With more basic algorithms we know that it’s clearly the human programmer and the interpreter of the outputs that are intelligent and not the algorithm itself. For some reason with AI that goes out the window. I believe it should not.
Some of it the effect of tells. “It’s not X, it’s Y” is not a bad pattern but it was baked into the instruction following training set just like the other patterns. I catch myself about to use it and use something else because I want to look human. I have, a few times, tried to use AI to write something that I was struggling to find the words and I just didn’t like how it didn’t seem like my voice. If there was just one person doing it would be OK but when it is 100s of blog posts submitted to HN a day it is like wearing a “I’m an NPC” t-shirt.
Someone shared with me this system prompt that at least makes assistant outputs usable
For information retrieval tasks, I want you to provide links to sources and use exact quotes as much as possible. When using a source, consider if it is primary or secondary information. If secondary sources are found, search again for primary sources. Sources and quotes, if applicable, should be mentioned in the answer first before the rest of the response with links.
Fair. Otoh, I am even more excited at AI assistants becoming sources of primary info. They do that now, but it's just very expensive and/or (un)expectedly rail-guarded.
/original_non_hallucinations skill?
To PP:
Are you looking forward to other uncles adopting foxwork? If you are you might be in danger of getting NPC'd without your consent haha.
Ashby's law of requisite variety should be cited somewhere..
You're giving your model instructions that it's literally incapable of understanding. A random word selection lottery machine will never do anything meaningful to determine if a source is primary or secondary.
It is admittedly awkward to bend to the machine to get what you want. I see these kinds of constraints like given above as an impetus to push the LLM designers to rise to the occasion, assuming they’re listening to all the prompts funneling back their way. This may all be wishful thinking however but hopefully someday these kinds of prompt constraints will be satisfied.
It depends on the odds of the lottery. As a straight-up classification task I'd expect it to do better than chance, which might not be good enough for you.
I mean this in the kindest way possible, but you are wrong that the math solutions are that easily dismissed. And there are many more than are publicized. A specific math problem I wanted solved for 3 years did not get solved by any model until fable and, and I tried it on every model and know the literature surrounding it well.
Are we sure there is some objective, technical definition of what is intelligence and what is not?
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
AI is just a good permutation/combination engine that tries to act smart with help of statistics. At best I only see AI as, 1. An autocomplete on steroid, 2. Good search/correlation engine
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
> I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans.
So does the Google search bar, but I don't ascribe intelligence to it.
I mean I'm quite proud of some of my search queries in the same way I'm quite proud of some of the LLM output I get. I'm probably just very arrogant and enjoying myself via some LLM indirection.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
The statistical nature means that what an AI produces is basically the average across all training data, making the generated text extremely bland and personality-less.
> ... including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
Is it possible that our aversion is simply driven by educational systems collectively and deeply ingraining into populations the idea that intelligence deserves the high costs commanded. Well of course this justifies higher wages towards the higher leadership positions, etc. Now it turns out that intelligence can be dirt cheap. We discover that the fact that "intelligence must be costly so don't question the costs of leadership" was never fundamentally true, so the real anger is this discovery of mismatch between the old claims which served to explain how every society that claimed to order itself and fill positions accordingly with "naturally pre-ordained individuals". Now we are seeing robots exceed average workers, for effectively a grain of rice.
It's not that different, when a human proposes a better definition vis-a-vis a competing one for example, they would defend this by certain desiderata.
Often a mathematician or physicist will use their intuition to speed up the naive brute force of candidate well formed formula variations so that the desired properties emerge, postulating the existence of an intersection on multiple desiderata can in itself be viewed as a novel conjecture, to be proven or disproved.
A very basic (unimpressive) example for an example desideratum is regularity or compactness. the tau=2pi substitution does make a whole bunch of expressions more slightly* more regular and compact. That is something objective and measurable on a system of theorems.
There is no mathematician's moat vis-a-vis machine learning at a fundamental level. There can be artificially sustained moat, if AI powers limit the distribution of say cryptographic advance capable models, in jurisdictions outside such AI powers, but even that would be expected to be fleeting and temporary...
I think the parent meant it is more interesting to pose new problems than solve them. Posing a new conjecture along the path to solving something is a close cousin, but still seems more bounded than proposing something novel to prove—if only because proving that something novel is also actually interesting is subjective and thus difficult for a different reason.
You're conflating a discussion about current LLM capabilities with your fantasies about nonexistent future AI. LLMs act nothing like this, and the small example you're giving is only a small part of the things that LLMs can't do.
Not PC but is ideologically high-handed and inappropriate to accuse the other side of failing to stick to your desired framing of a discussion.
Computer science is about what LLMs fundamentally are. If you implicitly focus on "actually existing LLMs", and require others do this, then that is not computer science. That is politics.
Not that I'm a professional mathematician, but I'm not seeing why people think definitions are somehow a blocker. LLMs have no issue making definitions (interfaces) in programming, which is formally the same activity.
Like when we had these recent counterexamples to longstanding conjectures, it's then pretty obvious to say "okay why did that counterexample work when most examples people looked at didn't" or equivalently "characterize examples that work vs examples that don't". There's your definition. "Def: An 'evil' polynomial is one that... Thm: conjecture is true iff f is non-evil. Thm: f is evil iff f is dastardly and a menace"
Or if you think it won't be able to come up with a sufficiently good name, just tell it to call the happy case normal, and it will be in good company with humans[0]. Sprinkle in some semi-, quasi-, pre-, and para- to cover the various different ways the thing might satisfy some but not all properties of being normal, and it'll fit right in. "A quasiprenormal Claude polynomial is one such that..."
> LLMs have no issue making definitions (interfaces/traits/abstract classes) in programming, which is formally the same activity. I ask them to form a core "spine" of a program (basically an intelligible theory), and they do it really well.
Is it a "standard" software? Something where the patterns exists in several other software? Try with something that is novel, or is in a limited set. You will find that it will copy heavily from what exists already, going so far as lifting whole functions from another project.
The goalpost moving is really getting absurd, to the point where now the machine needs to be a world-class once-a-century genius that invents entire new fields out of thin air (which are of course still relevant to humans) for it to be "intelligent".
I see this line of reasoning quite a bit and it’s a strange one to me. The arguer reduces the sheer complexity of human intelligence and language by saying “we are just running statistical models in our brains” and by doing so makes the leap that Llms are intelligent. It’s an incredible simplification of the human person, who has a deep inner life, a soul, desires, and a will.
I don’t think the aversion to llms as intelligent has to do with the economics of paying intelligent agents more. I’d argue that it’s more fundamental than that. Humans are incredibly complex, and the world of sharing invisible things called knowledge, and the intelligent persons consuming such things which has been going on for thousands of years is far more rich than these synthetic outputs.
When it comes down to it the ai has no inner life, its is dead. A useful coding tool sure. But I wouldn’t call it intelligent.
One side example is just how bad these llms are at artistry. Just saying whatever should statically come next is not good art—and the outputs show it.
You mention LLMs are dead and don't have the complexity or inner life that people do. Is your opinion that these kinds of things are not possible for AI in general, or that these things might be possible but we're just not there yet with modern LLMs?
You mentioned LLMs don't have souls, desire, or a will. I imagine those latter two can be engineered, no?
I'm in the same boat, I think some aspects can be engineered, like intention and desire. I'm really curious if it's possible to go the full distance and make AI have experience like we do.
Well they're language models. You can't capture the human experience in language. Simple as that.
> Is your opinion that these kinds of things are not possible for AI in general, or that these things might be possible but we're just not there yet with modern LLMs?
I used to be on the side of "we're just not there yet [with AI in general]", but after seeing people's response to an algorithm optimized to tickle just their language instinct, I'm actually a little bit more on the fence about it.
My view is that these sorts of things are not possible for AI in general. Though we can create things and name them “will” and “desire”.
Software deals with metaphors. Your Amazon shopping cart is a metaphor of a real shopping cart. You your desktop and your file system, etc. are metaphors of real items. But we don’t mistake the metaphor for its object, even from inanimate objects to their software counterparts (shopping cart to Amazon cart).
Now the metaphors are dealing with humanness, things like intelligence etc. And rather than seeing it as software doing what it always does, taking things and creating software metaphors of them, we are starting to say these are actually what they are named. Saying the artificial intelligence is actually an intelligence.
We’d either have to reduce the definition of intelligence such that calculators are intelligent. Or admit that these tools are not intelligent and are rather ways of exploring the work of actual intelligent beings, work that is found in their training data.
> My view is that these sorts of things are not possible for AI in general.
I wonder what it would take build an artificial system that has these qualities.
> Software deals with metaphors.
This is me wondering again: what's fundamentally different between software running on a machine compared to what's happening in our brains? In both cases you have energy flow following a pattern.
It's conceivable to create a system where energy flows in a particular way.
BTW, we navigate the world of an uncountable number of particles by creating models in our heads of what we think are large things out there. Approximations are made by both artificial and biological systems.
> what's fundamentally different between software running on a machine compared to what's happening in our brains? In both cases you have energy flow following a pattern.
Assuming we’ve scratched the surface of the complexity of the brain. I’d say in one case a human with a will is steering that flow of energy. In the other case it is a probabilistic algorithm steering the flow of energy.
The AI is not interacting with world with its own will. I see that as a big difference.
There are presuppositions that go beyond the realm of software engineering that guide one’s views of these things. One is whether you believe the material world is all that is, and that human consciousness is a product of the brain—or that there is such a thing as the soul or spirit of man. From the material perspective you may posit that if you emulate the brain then a sort of AI consciousness could arise. Or that emulating the patterns of the brain equates to emulating personhood. (Though what is material consciousness? I’d say consciousness is by nature immaterial.) I’m not a materialist, and I don’t believe the conclusions that arise from it’s perspectives are accurate.
Probability is just one way to model uncertainty. While I understand the brain encodes uncertainty, I don't think probability is a good enough model of what it's doing.
Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
I'm with you that intelligence is not something to be proud of. But I also think it is instrumental to understand the world. I'm still waiting for the time when an unconstrained-AI machine can live without reprogramming for an entire decade. We are still far from there.
> Probability is just one way to model uncertainty.
I study physics, mathematics, probability, cryptography,... so forgive my skepticism:
Show me how to model uncertainty without use of probability. Can you rephrase say diffusion, stochastic equations, quantum mechanics in this alternative framework? Can it at least make the same predictions?
Or is it basically the same framework in parallel, just giving different names for each concept?
Forgive my skepticism of such tall claims, and forgive my downscaling of anything else you say besides such a claim...
> Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
I make no claims of the specific shape of the implicit model implemented by a certain human brain educated in a certain educational system. For example in English the implicit human tokenization might be presumed to lay relatively close to English syllables, while in Asian languages it might be "sub" strokes of characters etc. Such implicit tokenization can never be proven to "match the one of humans" not because of human superiority, but because different humans use different tokenization methods. There is no "one human tokenization method", but it's clear as day there is an implicit one:
everyone knows the experience of knowing a word, knowing its approximate group-wise meaning (ignoring that when you think of "an apple" and when I do, we typically imagine a slightly different apple) yet having the word feel strange or discover some older literal meaning when decomposing it or looking it up in an etymological dictionary. Suddenly one can become aware of a sensible meaning as a composition of subtoken concepts. A child may perfectly know what "television" means and only later learn more exact meanings of "tele" and "vision", and upon repeating the word may feel the word "television" has changed meaning. This clearly demonstrates "tokenization" effects in human language comprehension, not just across cultures, but also across individuals within a culture.
> Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
That's a lot of different concepts conflated into one bullet point, so I split it up:
The notion of values and preferences.
They clearly demonstrate the ability to take into account values and preferences, from training corpus, from RLHF, from system prompts, ... we can't simultaneously point at censorship aspects and pretend their effective values and preferences to be absent. The censorship aspects are clear as day, so these correspond to values and preferences. Just like radicalization among humans, this can be due to exposure to radicalized content (akin to corpus data), from indoctrination (akin to RLHF), from "set and setting" (they may pretend to be aligned with one set of norms and values when standing in line to buy their new smartphone, but then reveal alignment with a different set of norms and values when conversing in some "private" online echo chamber). I see no grand difference between humans and language models here.
Awareness of values and preferences of a conversation partner. Allow me to widen it to "Awareness of values, preferences and prerequisites of...
Naah, I study cognition (besides a basic familiarity with the topics you mention) and even there, probabilities and bayesian models of cognition are well regarded by anyone familiar with mathematics. So, yeah, they are good. I think whenever you want optimal/rational inference under a closed set of alternatives, probabilities will do you good. But is the assmption of closed set of alternatives good? I doubt it. The set of all alternatives does not make sense, unless you specify a context, which almost everyone does when they are trying to model.
As much as I want to explore these other topics, I have not. So my understanding of these is limited to "they exist, but are underexplored, and it's not clear to me they can or cannot be reduced to probabilities". I hope to explore them in another life or another decade. I'm currently stuck on causality, but even here, they talk about normal and abnormal events. But what is the probability of me eating a cabbage today?
Understanding formal proofs and understanding natural language are as far apart to me as day and night.
I make no claims (or at least don't want to) about the superiority of human cognition (by what metric?). I do make claims about their similiary however. And they are not identical, which you seem to be claiming save for differing trainign data.
Let me make a distinction between values that arise by mere existence (hunger, thirst, fear of death, lust), and values we pick up as we grow up (religion and social norms). Machines can acquire the second, but acquiring the first requires being embodied in the world that is different from putting a microchip into a robot body. I don't do one job over the other because I was exposed to some training data that said I should do a job. But I do it because I value earning enough money to not starve myself to die, amongst a host of other values including wanting to enjoy what I do. Even the notion of enjoyment comes from embodied existence. You don't discover you enjoy something before you try it out. You don't always learn what you enjoy by looking at other people. Now, when you start giving AI the threat of death and the joy and pain of life, perhaps you can arrive at something similar to humans.
If someone started using mechanical help to the extent their healthy muscles and body deteriorated, that'd be a matter of concern too. For example, when you only drive and never walk. But this is exactly what is happening whenever you put AI in the hands of (unwilling) learners! I don't understand how using an LLM maps onto using medicines or prosthetics. They are clearly not the same! If you know a disability that LLMs can be used as a medicine or prosthetic for, please let me know!
Mainstream AI does not even understand causality except to parrot cases it has seen in the training data. There is a whole field of research in causal inference that needs to make its way to mainstream AI. So, good luck putting an AI that works solely on associations and correlations in charge of its own body let alone a nuclear reactor or the state.
No one ever says that though, do they? And I don't think it's required. Here's what some mathematicians are saying:
> These developments have triggered some deranged thoughts in me. I have wondered if it is the express goal of these companies to make me kill myself…The story of human discovery and the triumph of the human spirit will soon be excised from this discipline…
> I want to feel seen… I need the architects of our new mathematical paradigm to look me in the eye and acknowledge our shared humanity and soul before they deliver the coup de grâce.
I'm not saying this means mathematical research as a field is finished- I don't think/hope Kirwin really believes that either. But it's clear that these developments are not just touching "a very very small part" of the field.
The phrase "God created the natural numbers, all else is the work of man" is a famous quote by the 19th-century German mathematician Leopold Kronecker.
You can basically read it as: the moment one has axiomatized mathematics to the point it supports natural numbers, the rest implicitly follows. The natural numbers (positive integers) are closed for addition, multiplication, ...
One can perfectly model the integers with a pair of naturals: < M, N > ~ (M-N)
Now we can have any < M1, N1 > and subtract < M2, N2 > without needing the ability to subtract natural numbers:
< M1 , N1 > - < M2, N2> ~ (M1-N1) - (M2 - N2)
= < M1 + N2 , M2 + N1 > ~ (M1+N2) - (M1+N1)
We can similarily define addition of such tuples, or test equivalence without access to subtraction of naturals:
Similarily, even though these newly defined integers (which can be positive or negative) don't support division, the same trick can be used to make a new compound tuple of integers closed for division, by only using multiplications.
Probability is a branch of mathematics (probability already exists embedded in mathematics implicitly, probability theory involves the addition of eliminable definitions, syntactic sugar. The patterns are already there, just less explicitly manifest.
Mathematics is itself a branch of logic.
Do you reject like all of logic, and if so, what would you like us to evaluate the sentences you write to? You want us to evaluate your expressions as "true" or as "false"?
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
Not that I'm saying AI are like brains, but can you describe why brains, which are fundamentally slightly dodgy electrochemistry with frequent literal delusions of grander, are not "statistical"?
> No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
Ditto, when do we humans do things exceeding the parameters of "correlation engine", especially if you consider compositing things either we or some other part of nature has developed and documented elsewhere to be insufficient?
I strongly agree. Is a human in a vegetative state sentient? What about when they're asleep? What about someone with brain damage? What about someone with an IQ of 20?
Ray Kurzweil argues sentience is a philosophical question, and doesn't have much value as applied to science and technology. What will change the world is how this intelligence is applied. No one's going to care whether AGI is defined as sentient when it creates cheap fusion energy.
> is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever
What makes you so convinced that a algorithmic construct of neural nets cannot be "real intelligence or critical thought"?
Alright, take it easy. You typed a lot here but you're not actually saying much. LLMs produce useful outputs, their usefulness is just proportional to how well you know how to use them. Everything else is navel gazing.
A decent correlation engine is still extraordinarily valuable for science, investing, prediction, etc. Plenty of human minds are strong in the same area.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
The problem might be that it cannot backtrack. When AI generates output, there is no backspace key for it - it uses "No, but wait!" all over instead, which is very different to human output.
Subagents and/or branching conversations are presented as the solution to this - if you can't backtrack, then branch off a conversation to explore multiple paths (discarding the ones that didn't pan out), but this is a fix in the harness not a fix in the model. It's also literally how we made chess-playing engines back in the 80s: recursive path exploration with a fixed depth.
Humans don't exactly work that way either, AFAIK. So we have this uncanny valley of intelligence: it's some sort of intelligence, but not as we know it.
Don't discount the capability of representing human-like intelligence with simple constructs. Give enough parameters and advanced enough training you can without a doubt create real intelligence. We've seen some of this already with "j-space" where llms have started to exhibit reasoning before it ever reaches the output head.
> it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever
How could I seriously repeat that prayer, when it builds things I wouldn't be able to build and solves problems that I wouldn't be able to solve? I would have to assume that nothing I did in 25 years for money required any intelligence or critical thought whatsoever and I have higher IQ than 99% of the population. You might be comfortable with that but I'm more comfortable with ascribing at least some intelligence and critical thought to AI.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct
Isn't this how human brains work? We're just large probabilistic neural inference machines. A network of neural weights guided by past training.
To be honest I think the debate about what "sentience" is is inconsequential navel gazing. There are many different kinds of intelligence - even within humans. What matters is how useful that intelligence is as applied to solving real world problems. Like it or not, LLMs produce intelligence which is very, very useful for 1.5B people and rapidly growing.
I think this ultimately boils down to the classical economic debate of marginal utility. There isn't an objective way to value a product or service. Each person decides for themselves what said product or service is worth based on their needs and preferences. Intelligence works the same way. We don't have the right to tell someone that their perception of the value of that intelligence is wrong. They alone determine that.
One thing we can all agree on, is that the capability of this intelligence is expanding rapidly. In a few short years, we went from Will Smith spaghetti hands to full length movies and strikingly realistic images. For 60 years, passing the Turing Test was considered Star Trek level science fiction. Last year GPT 4.5 passed the Turing Test. AI is already being used to convince people over audio that they are real, and very soon, this will occur over video.
I think people are being too dismissive of this intelligence. It doesn't need to be perfectly humanoid to be considered intelligent.
We have actually little idea how human brains work.
Every age thinks they know how they work, and then every subsequent age laughs at the previous age’s rudimentary understanding.
We gave a Nobel prize to the psychiatrist for inventing a method to remove the frontal lobes of a brain through the nose in 1949. Lobotomies were performed through the 70s.
We are likely doing similar if not subtler but worse things today (just one more pill bro). We still have no idea what we are doing when it comes to the brain.
>Isn't this how human brains work? We're just large probabilistic neural inference machines. A network of neural weights guided by past training.
Artificial neural nets are merely cartoons of how people guess human brains work. They're likely much further from reality than, say, the fundamental laws of physics which can be experimentally verified or disproved.
not having any body, continuous sensory input (except ChatGPT-live gets streaming audio), episodic memory, on-the-job learning (live weight updates), any live feedback loops (like moving a motor updates proprioperception or turning physically changes what a streaming camera sees), really hurts these models' abilities to perform tasks of the kind you're waiting for.
They are very good at instruction-following and you can teach it a new task that fits in its context and it'll learn it and do it. Go ahead, you can make up some new brand new ruleset or behavior and instruct it to follow it and it will. That's amazing.
But it won't be any better at its new behavior after an hour or ten hours or ten days. It doesn't have the kind of adaptation that we expect.
What it is able to do already is pretty amazing, but what it lacks is also a great hindrance to seeing its full capabilities. We just have to wait until research labs add these missing components.
I am firmly realistic on the overhyped nature of AI, but which discoveries in the last 100 years are not "a composition of solutions humans have developed and documented elsewhere" ?
Is that basically every new discovery? And under the strictest definition of novel and NOT falling into your composition of previous solutions what is that standard of proof to beat your criteria? Is a novel discovery not allowed to use English but must invent their own language? Must they invent their own math - these are hyperbole for illustration but I think its not far from that before you could just argue anything based off it is a composition of existing ideas
I experienced the same lately. Even dug some of my old posts where I put in the effort and formatted them using reddit's markdown. Wouldn't dare it today
> There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
I think you need to self-correct here, because otherwise you'll be ineffective in an information setting, where I expect AI-generated resources will not only be the norm, they will absolutely swamp the environment.
AI-generated resources swamping the information environment only makes it more important to have the mental mechanisms for quickly filtering out their non-information.
Do you have much exposure to pre-AI corporate memos, mission statements, marketing plans, or white papers? Because they were mostly written in that style. Full of buzzwords, cliche similes, platitudes, jargon and stock phrases.
The thing is, people writing them had a style. Every company has its own style, or feeling for these kinds of texts. Also for the initiated, these buzzword-filled blocks of text provided some between the lines information; sometimes big, sometimes small.
AI generated text doesn't have this. Every model has its bias towards a certain style, an overly agreeable tone, some exaggeration to make the user important and smart, but the text has none of the information crumb these pre-AI texts contained.
Even when you use tools like Grammarly and allow it to "Impact-MAXX" your text, the resulting text is a bland wall of letters, carrying none of your voice or style, less elegant than a corporate text and emptier than space.
I also read right through those. 90% of most companies' webpages is drivel. Same with 90% of a job listing's text. I hardly understand how people get any meaningful information from corporate websites, they're all just "enhancing your business outcomes by incorporating technical excellence with synergistic AI" or something like that.
There's somehow less information than if they just asked claude to make something up without any context.
Once you see past the illusion I think there’s no going back. AI writing style is just dogshit. This hype wave is based on the belief that we’re inching closer to AGI but seems to me we just increasingly struggle to define intelligence. LLMs seem smart because they can pump out thousands of LOC quickly, and enthral you with fancy words and bullet points. I don’t fall for the intelligence illusion anymore.
I'm not sure we need to declare AGI around the corner nor declare it all dogshit. I think that's part of what's so dissatisfying about it; it strikes at such extremes of both awesome and awful.
Completely agree. AI is very impressive in many ways but there is something deeply wrong that is hard to put into words. The output is probable but never true, if that makes sense.
I think this is also the mechanism behind why AI generated videos and images are so captivating at first. I remember when Midjourney first launched and it was hours and hours of a brain-melting "Wooooooow". But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
Yeah AI generated content hints that there is a whole world behind it, the way that an image pre-AI was a clue that there was a rich 3D space that corresponded to the image.
It seems our brains are adapting to that and recognizing "actually the signal behind this message is quite sparse" even when presented with rich imagery.
> If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
If I were to push you a bit on this, when is it not true?
Let's not like at AI specifically, but can you think of other examples? Like for me, I think of: the creation of earth itself, or stars, or even DNA.
> But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
I guess the majority of people do low-effort generation that doesn't perturb a default style of a network enough, so it stays blatantly noticeable. The percentage of "super-recognizers" who notice almost all AI-generated images is around 1-2%. It could be that you are one of them, of course.
I think this is the confirmation bias trap a lot of people fall into; higher quality, hard to detect AI is already ubiquitous but because it's hard to detect people just don't clock it.
"I can accurately detect 100% of AI generated images that I recognise as being AI", if you will.
Just the other day I was using text-to-speech with Gemini, and for some reason, it transcribed my full query in Hindi (in the middle of an English conversation), and naturally the LLM responded with Hindi as well.
I don't know exactly what I said, but after translating it back, it appears to have attempted a phonetic transcription of my words (rather than translating my actual question).
I've had Gemini CLI (the coding agent!) smuggling in Chinese words in the output - completely unexpectedly, I do not speak and never discussed anything Chinese with it. Copilot was mixing in Cyrillic character, unprovoked - take "обligation", for example. And so did Grok, when discussing Russian anecdotes - "KRЯК". Very helpful reminder to not anthropomorphise them machines.
The junior engineers at my job have a terrible problem of writing AI "proposals" to problems. The proposals are all extremely detailed and verbose to a thought-terminating extent. It takes a lot of effort and self-control to parse out the actual "ideas".
I think of the Dwight Eisenhower quote: "Plans are useless. Planning is indispensable."
The process of thinking through a system and communicating your design to other humans is a core part of software engineering. You want to build the right abstractions and communicate the right level of detail. Delegating all that thought to an LLM means your proposal isn't clear to the target audience, and it's not helping the author to understand the problem.
Yeah, I have the same problem. There's a good quote example of this:
> There’s a growing scissor between people who are happy to read AI and those who violently bounce off from it.
> People adapt in different ways — and some people absolutely cannot look at it. That cognitive split creates a surprisingly powerful opportunity: you can write something that, technically, sits right there on the page, yet an entire sub-population will be incapable of staying with it long enough to actually read it. You can hide entire sub-structures in plain sight. It’s not avoidance — it’s adaptive obfuscation.
> The paragraph before this one was the only thing generated in this essay and if you just skipped over it I highly recommend reading and really understanding what it’s saying.
It's quite effective. I think this kind of text functions like the chumboxes you see at the bottom. Taboola and so on. Just mental ad-block takes over.
Yep. It's like it's painful to read for me. It's because the next-token predictor is just mashing (mostly) grammatically-correct and plausible sentences together, without any real intention or meaning. So everything sounds plausible, but almost entirely void of meaning.
It's like if on any website you went to you saw a lot of posts written by the same guy over and over again. Even if he used different names, you'd start to recognize him eventually because of his style. Seeing as he doesn't say a lot of valuable stuff, you'd also learn to skip whatever he says.
I do worry that it's just survivorship bias and we're also consuming higher-quality AI output that's indistinguishable from human writing, but we focus on the raw, unedited, low-effort AI slop and think that we're good at recognizing AI text. Even if we really are at the moment, it might not be long until AI companies figure it out. I'm not sure why they haven't yet, given how many books they've burned for this already. Maybe it's just more efficient for the model to stick to a single way of writing, I don't know.
But when that point comes, we'll be back to the usual way of reading and interpreting text because there would be no way to tell what produced it.
> just short-circuits to "there is no information here"
That is my experience with the way the models write by default, often even when instructed not to do that. With enough effort you can get even them to slightly unslop the writing so it doesn't read like some LinkedIn/Buzzfeed brainrot, but the problem is that it's not trivial to do and most people won't do it, so the default is indeed horrible.
When I read AI-generated prose that is aimed at the general public, I have the exact same feeling.
But when I ask Codex a technical question about coding, I don't get it at all. Codex replies to me in a very direct, technical manner, similar to the way I speak.
When I ask ChatGPT to be concise and technical, I get the same effect.
I think it's because prose aimed at the general public has to be very attention-baity --like the textual equivalent of a Mr. Beast video--, not because AI is incapable of writing like a human.
I think we as professional documentation-readers already skim most content (speaking for myself, I realized I was googling and skimming for answers 20 years ago instead of reading documentation end to end), but AI generated content has the same problem as marketing speak in that it's a lot of fluff.
It's understandable people don't read but feed stuff into their own AI again to bring it up to their standards or have it get to the succinct point.
Exactly, AI-generated text reads so smoothly, that the same short-circuit shifts my attention away from deep focus and onto scanning of the text, looking ahead to get the gist of it. Forcing myself to read the text fully feels almost painful. It's like reading a terms-of-service or any boilerplate document.
It reads like the white papers companies publish on their websites to build legitimacy. Or anything from those IBM / SAP / Deloitte / etc consultants who write technical papers despite having little to know understanding of the technology.
That's why the business and government people love it, they spend their entire careers reading this nonsense.
This is just a weird feeling that I've been coming closer to articulating lately, but I only think that you can get forward reasoning from what is basically word association; there's no mechanism for unwinding it because it has no real memory. By "it" I mean word association itself, not any context window. It predicts what could be in a position, and ignores what wasn't in a position.
People don't do that. People are constantly engaging with paths not chosen. Right after I choose to write one thing, I'm immediately engaging with what I chose not to write there - I'm explaining why I didn't write it, I'm realizing that my choice may seem unusual so I'm trying to make it memorable, I'm focusing on the distinctions between what I wrote and what I didn't.
LLMs don't currently do that. LLMs just ape a structure. When the structure resembles the sort of timid, clarifying fussing I just described, the LLMs just drift randomly because what they didn't say wasn't in the context.
I also think that's why they have such a serious problem backtracking. They're not taking into account the already eliminated possibilities. Often the thing that was so unlikely that you weren't going to waste time on it is the answer, and things you discover while going down an ultimately wrong (but initially far more promising) path remind you of the path not taken.
They're simply assembling a thing that resembles a valid argument, and happen to make sound choices because the plurality of input happened to contain sound choices. This is usually a very good bet because there are so many more ways to be wrong than to be right. But it doesn't account for attractive (common) wrong choices. You need a way to back out of those.
> There's some psychological mechanism by which my brain immediately recognizes AI generated text and just short-circuits to "there is no information here".
The roots of llm math in part lie in compressing natural language such that there's only information there, and then running the reverse to create way more text without new information in a somewhat precise theoretical sense.
I love 3b1b and I love that video, but that also isn't exactly what is being said. In particular llm inference does add information (in the meaning in this context) because the output distribution is sampled randomly.
I think you are adjacent to the real story here, but missing it. AI text contains information, certainly. Frontier chatbots are very good at creating acceptable and mostly accurate answers to our questions on just about any topic. It’s an astonishing achievement.
But you are sensing correctly that there’s something missing. It’s the meaning and the speaker. Communication is an exchange between speaker and listener. The speaker has a meaning in mind, and wants to create that same meaning in the mind of the listener. Therein the problem.
There is a listener, sure. But no speaker. No meaning. There is information, but how can this be communication? Nothing is talking. Or at best, we are just talking to ourselves, our own words back at us through the funhouse mirror.
When your mind looks at AI text, you know you can safely ignore it. No one wrote this. No one cares if you read it. You can delete it and nothing of value will be lost. It might contain the information you need, or a bunch of gibberish. There’s no one’s reputation on the line if it’s gibberish.
I prefer to think of this in terms of Umberto Eco's opera aperta (open work): if any text is a collaboration between author and reader/recipient, here, all the burden of meaning is left to the recipient. There's simply no meaning on the side of the "author", it's just a statistical extraction.
(There's also the problem of words/signs (just) referring to other words and/or cultural entities. There is no world nexus in this, therefore also nothing we conventionally refer to as meaning. On the other hand, it's utterly dogmatic, as all it refers to is the most probable construct, as a reference to references that are just another utterance, but supposedly a dominant one.)
>and just short-circuits to "there is no information here"
I feel the same way when I read a "press release" or anything written by marketing. Even the newspaper will only have 2-3 sentences of interesting information spread out over 4 paragraphs.
This makes me think of the paper-to-media pipeline; scientific papers are high information density. Its press release summarizes the finding. Then the popular science and social media posts come that oversimplify and embellish the findings.
So from "this table of stellar luminocity observations shows x y and z" to computer renders of green/blue planets with captions of "LIFE FOUND IN SPAAAACE!".
Before AI slop, there was corporate slop. It served the same purpose and was generated in largely the same way. But it was only deployed where it was worth the cost of creation.
> Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
It is because GenAI output has no thought behind it, as you identified in your previous paragraph:
> And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
You are searching for meaning in something which was not created to convey meaning. The text was, instead, the result of an extremely clever statistically based algorithm.
Same. Also when I see a spec for example or some summary I always have the impression it doesn't get to the point. Like, the core ideas are in there but also somehow lost and I have to work them out again which makes me wonder if the person generating it understood what is going on or if it would just have been fast to just write it by hand (you still use LLMs for research and such).
For me it's youtube videos. As soon as I hear the AI tells in the script, even when it's clearly read by a real person I immediately look for a different video.
At least for the content I watch for entertainment, it may be different if I am looking for a specific answer for something where I would otherwise just ask an AI anyway.
Another good example of AI making my brain do more work is when you ask it to compare two things:
“Compare a car and a bicycle”
The answer is invariably something like:
Seats: 1 (bicycle) vs 4 (car)
Tire width: 1 inch (bicycle) vs 12 inch (car)
Steering: handlebar (bicycle) vs steering wheel (car)
Instead of “bikes are useful for short trips if the weather is ok and you like getting exercise, whereas a car is usually better for longer trips, bad weather, or multiple people”
I wonder how the autoregression works on that. You'd think "seats: a bicycle has 1" would be an easier autoregressive completion than "seats: 1 (bicycle)" since it works left to right adding context.
> Something is deeply wrong with AI generated output
For me, it is the endless maximalism and hyperbole. Almost if the output was driven through a radio-mix compressor - too loud for the reader/listener to be able to pickup any dynamics.
When I ask AI to research technical information about X and (include sources) - I get mostly solid information as response.
But poems or interesting fluff blog entries by LLM's? Not something I look for.
What disturbs me is all the "pretending to be human" all the personalizing language - that is clearly fake and I would much rather have a neutral robot language as response.
Back and forths with the LLM are very useful when the person inquiring the LLM is invested in the conversation and wants to uncover the ground truth. Ultimately, some good result can come out of it because the person has a goal and LLM helps them reach that goal by uncovering the layers of knowledge which the person might not have.
When the AI-generated content is presented to a person without any prior investment, it just looks incoherent. An especially great example are these Claude-generated explainer-type pages, which look really nice, even interactive, the information from the first sight looks really well presented. But somehow it all just doesn't make sense to a human. And I think it's because humans are processing information linearly and building an internal story about the information. One could argue that LLM's also consume information linearly but the way this information is processed is a kind of all-at-once approach.
Just some speculation on my part but I have been trying to cope with this way information is presented because I am currently working at a place which is heavily documented by AI. And the only way for me to properly understand the documentation is by inquiring AI to help me.
Oh my god thank you. I have been trying to put into words what is so draining about pair programming with claude, and “doing creative work to impart meaning to the words” is exactly it.
I once heard on a podcast about comics that an artist ran into someone who couldn't read comics. Not that they didn't like them, but they claimed that they just didn't understand how they worked. When the artist explained 'well each panel shows an illustration of the events in the story', the other person was confused - "each panel is connected to the previous?" they asked. They just couldn't wrap their head around the fact that each of the panels was connected in some way - their brain just refused to interperate them as anything other than 4 separate completely unrelated panels.
When I see AI animated videos, that's how my brain feels. It's this strange brain-fog that I just cannot connect together the sequences of images being shown into some sort of chain of events. My brain just refuses to see them as anything other than a disconnected series of 2-3 second videos, even if the same character(ish) appears in them all. It's very strange.
Its fuzzy matching a sequence of input words to weighted patterns extracted from a humungous training set based on the sentence structure (not meaning) to generate another sequence of words, and then filtered and smoothed to make the output more agreeable.
Its fundamentally flawed. like tarot card reading and astrology.
I can relate. I noticed this exact phenomenon when encountering NotebookLM-generated diagrams recently. Even if I know there's some intelligent thought behind one, it's like some slop detection circuit-breaker is tripped.
My son is currently learning Romanian and I was trying to help him with verbs. I don’t know Romanian but recalled when learning a foreign language for the first time it really helped me to break down how a verb form or tense worked in English, then learn the equivalent in the new language. So I wanted to make some charts and pages that he could use as learning resources.
I used Claude to help. I don’t know how to quite describe it, but because the text was polished and well constructed my brain was giving me the the signal “if you aren’t getting this it’s because you’re not focusing” so I’d read it again and then again and it still was not landing. It sorta felt like when you read something technical or heavy when very tired - you are reading but not processing.
Only after wrestling with this for a few days did I realize that it wasn’t me. As I started going through, sentence by sentence, forcing it to re-write things to be more clear the concepts became easy to understand.
I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The more complex the topic, the more I sense this.
Feels like the way a video game will render the outside of a wall or solid surface, but you can run into it and warp partly through and there's nothing internal to it at all.
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
I understand the sentiment, but not quite what I’m thinking. It’s not that the AI is necessarily wrong, it’s more like, it replaces clarity with rich but unhelpful text. Maybe like candy. Full of flavor, texture and color but lacking the nutrients.
The way I like to describe it is that it feels slippery, like it's been polished down to oblivion and my eyes slide right off of it. There's no place to find a mental foothold, and if you zoom in there aren't any details.
But I also like your candy analogy because I think it's spot-on for how LLM text superficially looks informational/nutritious, even though it's actually just junk.
No no no. Harry Frankfurt wrote at length the difference between liars (who care about the truth, and twist it) and bullshitters (who don’t care about the truth, but just the way they’re received).
This is some third category of untruth. Almost more sinister than the other two altogether.
I think, they don’t care about truth at all, just whether the reviewer rates it highly. Otherwise, it would be deleted after that round of training. It values test-taking ability over critical thinking.
My two favourite words for this are “conditioned” and “catechized” where the latter is a bit more on the nose but way more obscure.
A lot of it is hedging and using lots words to avoid saying something that isn't true. That way, it looks a lot like text that contains meaningful information, even though it doesn't. The speaker can pass the scrutiny of an informed audience, because they are able to substitute their knowledge into the words, as if by pareidolia. It's the sort of 'diplomatic' language you see from politicians, lawyers, students writing essays, and any other moderately intelligent person who is put in a position where they will face consequences for not answering. Behind all the fluff, there is a very loud voice yelling "I DON'T KNOW."
> I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The best I’ve heard of this is peeling the onion. The first pass is always very high-level and you have to make it go deeper. That can be done manually with follow-on prompts but I like using subagents, each with a different angle on the problem.
This is a really good description of the problem. I’ve been trying to use Claude to get familiar with the mechanics of a new codebase, and there have been so many moments where I’ve stopped after reading the same paragraph five times in a row and thought “Am I tired? Or stupid? Or is this codebase just wildly more complex than anything I’ve seen before?” before realizing that it’s just taken English and smushed it around like a ball of clay into some abstract sculpture that kind of evokes something from real life.
I think part of it might be an innate feature of LLMs, but Claude seems extra prone to it lately. I ran the same query about the same codebase with Codex, and it gave me an answer that was about 1/4 the length and made me realize that it really wasn’t all that complex.
If nothing else, it’s good training for my own writing. I’ve been working on making myself be more straightforward and concise, and Claude’s writing is a good example of how cleaner prose is a functional choice, not just a stylistic one.
Ironically, I've been already active on codex after trying it for the past 1.5 months, from Claude, and I find it to be similarly complex and confusing.
I think they all have the similar styles and tells. If I were to go to Claude, and use it now it would probably be clear for a little before reverting.
And I don't know why it feels to me like the language "drop off" happens after some time with the system. It makes me wonder if my account are getting silently degraded or sent to lower intelligence/lower priority queues after being a member for a while.
It's a Claude problem, chatgpt is about x10 better and it's quite rare that I'm getting such foggy vague texts that I have to reread to understand. Chatgpt used to have the same problem in the early days but even then it was less pronounced than how it is in Claude today. Fable is a bit better but not as good as Chatgpt.
It swapped at the Opus/Fable 5 and Sol developments.
Opus 5 uses such convoluted language that it takes me multiple seconds to parse it. I specifically have to instruct it to use some kind of sane language (Caveman the AE STD English hack or something) to make it understandable.
Before that it used to be GPT that used 500 words when 5 would do. The ChatGPT version still does, but Codex is a lot better now.
I mean, yes - but somehow this feels like a different category. At least to me. I view slop as the classic garbage like:
[Thing] isn’t just [X]—it’s [more dramatic Y]. And [short validating statement].
I can see and smell this type of slop from a mile away. What I’m referring to is in the same family as slop but somehow different - it fools my brain by putting on the presentation of credibility and thus it is even worse. I can skip right over classic slop without much effort. This kind of text tricks me into laboring over it before I realize it’s hollow.
At least in my mind, lack of substance was always the defining feature of slop. Stuff like "it's not just x, it's y" would be reasonable (if grating) tics from a human writer. Whereas humans are also capable of clever-sounding but substance-free writing, and it still sucks. AI just produced enough slop that we were forced to give it a name.
I know what you mean. The sentence appears to makes sense, it is correct grammatically, but lacks substance. It's the equivalent of AI image generated stuff where at first glance the image looks fine, but it hurts your brain. Then the more you look the more you see nothing makes sense in it, and it just gives the appearance that it does.
It's definitely worse that slop, because it really hurts your brain. You read it over and over, and feel like you're supposed to understand the sentence, it's coherent, eloquent even, but by the end of it, you don't know what the f it was about.
As far as terminology goes, I still see these both as subcategories of slop, but I can buy that they do deserve different categories of obvious slop versus more virulent slop.
The human spirit. When you read a real person's thoughts you can often intuit the thought processes that led them to write it which aids understanding. Or at least have a general idea of "where they're coming from". But an AI is missing that. It just knows everything, without a "thought process". Instead of a flawed 3d person, we get a nice 2d picture instead.
I don't like AI-generated text at work, at all. It feels lifeless and unfocused.
But what I hate the most is that it is objectively better than what I had before. No typos, clear structure, and, regrettably, the verbosity and autistic obsession with detail of the LLM is more actionable and useful than the human guy who wrote lists of commands and URLs as documentation, without explaining anything. Or the colleague who writes in uppercase and with question marks and who doesn't make any sense and forces me to engage in an interrogation effort to get to the bottom of what they are trying to say. Or the colleague who simply hates writing--despite being decent at it--and will call you to give you a meandering verbal explanation that lasts two hours of what they want from you. The cynic in me bemoans that we brought this upon ourselves, in more than one way.
My concern is that even if the LLM can turn your colleague's bad writing into something more coherent and actionable, is that something actually what your colleague meant to convey? It could be clear and still detached from the reality of their intention, or they may not have even formed a clear intention. If the goal of writing is to convey what's in another human's brain, that goal is failed completely.
I'd say my experience is different. Even from people whose communication writing I didn't find that useful, they seem to have a better frame of mind than LLMs do. Though, I never encountered people like the examples you gave.
Why do people glorify autism? I have seen this mostly on CEOs that want workers that are hyper-focused, intelligent and do not complain. The ideal employee from which extract value and a minimum cost.
I have worked with autistic people (diagnosed ones). And it is a struggle. One of them could be the nicest most reasonable person one minute and then get a trigger and become obtuse and irrational while making the rest of the team suffer thru all kinds of complains. Meetings would need to be adjourned and time was lost.
> The cynic in me bemoans that we brought this upon ourselves, in more than one way.
I enjoy working with other employees and people need human contact to become fully self-actualized. Isolating people to maximize productivity misses the point and brings no more productivity and just avoids fixing any real existing problems.
> > The cynic in me bemoans that we brought this upon ourselves, in more than one way.
> I enjoy working with other employees and people need human contact to become fully self-actualized. Isolating people to maximize productivity misses the point and brings no more productivity and just avoids fixing any real existing problems.
While I share the feeling and the posture, I haven't managed yet to use them for better documentation. In fact, for me enjoying human contact takes precedence to pressing people--however politely--into writing better docs, specially now that we have LLMs. The skill of writing should be taught and nurtured in school, and if you made it to the workplace without it, then maybe it's too late and you are better off using an LLM to assist you. Which is exactly what I bemoan.
It’s currently got some kind of strange anti-memetic effect where it causes me to skim over it.
I noticed this too. There’s a great section in a tweet which illustrates this but the tweet itself is much longer and I can’t link to paragraphs directly so I’ll link to my blog where I quoted it[0]. It’s really a remarkable effect. It takes effort for me to take my eyes back to the scene.
Presumably you could put contradicting information into text like that and reach some people and not others.
It reminds me of that (apocryphal?) story of how an election advertisement company managed to shift votes by telling everyone to not vote knowing that people of one ethnicity would nonetheless listen to their families and vote. You don’t have to lie to people, merely put the text in a place inaccessible to their minds but right in plain sight.
I've been a big reader, and many AI outputs nowadays reads polished similar to published books.
The reason I brought it up is because, people who learn English normaly start with a book. It's heavily polished.
When you speak English as you learned from the books, it does not sound very conversational.
If you are native/fluent English speaker, you can feel the impedance mismatch and feel something's off.
The AI-blindness stems from the fact that those polished edits are so common in publishing field, they all sound the same, and unable to recognize the diffs between AI-generated and human-generated.
There is no real human conversational vibe to them and well. i will stop now.
In high school, a teacher gave me a copy of "How To Read Better And Faster" which teaches you speed reading. This came in very handy in college.
I find that when I try to speed read modern human writing, there are often errors (like missing or misused words) or awkward expressions that I do have to slow down and think harder a lot to really parse it.
With AI writing, it's sort of self redundant and the information density of each sentence seems to have more even information density. This makes it very easy to do a very high level speed read and get the full gist.
There are also what I'm assuming are bots on hugging face (or maybe non-native english speakers who are using ai for translation) that interact with me where I have no idea what they are saying until I read it very slowly.
That's interesting. You're saying speed reading helps you grasp information density of text?
As someone who hasn't practiced speed reading, how does that happen? Is it something about the way your brain tries to connect ideas from different parts of the text? Or the redundancy making the signal more stable?
When you speed read, you can grab the words off the page faster than you can understand and fully process the information conveyed by the text. How much time you have or want to spend re-reading or thinking about what you've read is quite obvious. There's a stark experiential difference between reading an informationally-dense passage and one that spends a lot of time rephrasing things, using LLMisms to restate concepts, adding in extra connecting phrases, etc.
If your reading speed is limited by how quickly you can subvocalize the words to yourself, this is significantly less obvious. Unless the passage is dense enough to require multiple read-throughs at conversational reading pace or vapid enough to be boring, you're going to feel done with the text at roughly the same time. Speed readers do a lot more re-reading and varying of reading speed, and that is going to correlate pretty hard with information density.
I haven't looked at the book since the 80s, but what I remember is there is a pre-read, the fast read, and then the deep read as needed depending on the task.
pre-read is just looking at how long it is in the headings, and planning out what chapters to focus on if it was a text book. (its sort of iterative, you do a pre-read for the whole book, and then for each section you break it into)
The fast read you try to read only with your eyes, sweeping your eyes across multiple words at the same time, suppressing the urge to say the words to yourself in your head.
iirc the how to read better and faster book even had a cardboard mask you put on the page to practice the sweeping, and some pages that were laid out weird to try to teach you how to do it.
Some ai text just seems really easy to speed read, like if it's tuned for an easy reading level. In PRs some ai seems like it's arguing over weird flex technical details and really starts torturing the language in a way that makes it the opposite of easy to read.
Does speed reading help you process the final message faster if it's written by AI compared to people?
Because if you read 1 information dense sentence, 1 medium dense, and 1 sparse sentece written by a human, it's still way less text in total than 6 information sparse sentences written by AI... even if it's all over the place when it comes to density or style.
---
The density argument is really interesting.
Does speed reading actually help you process the final message faster when it’s AI-generated compared to human-written?
For example, if a human writes 3 sentences—one information-dense, one medium-density, and one sparse—that’s still much less text overall than 6 relatively sparse sentences written by AI.
Even if the AI output varies a lot in information density and writing style, you still have to process all that additional text. So I’m wondering whether speed reading actually offsets the verbosity of AI-generated responses, or whether the total amount of text is still the bigger factor.
I am wondering if working with agentic AI for 500+ hours built up skills for this.
While I would feel rusty when handwriting code, working with either Codex or Claude I have lots of practice.
I can probably tell at a glance where output is intermediate while it's still running tools and where the summary starts.
Then I have to recall the context of the conversation, parse the summary, see if there is anything unexpected that popped up. Make a call for whether I need to ask it for clarifications, and what I need to test to verify the change.
Assessing quickly what is pointless yapping and where the information I need is might be a learned skill. If one tries to read and understand every word Claude says, I am not surprised that people are not having a good time.
Part of me likes the cliche Claude voice. Not because it's good, but because I can immediately recognize it. When I see it in the Claude app/code then it's fine. In the wild it's a sign to me that I shouldn't keep reading.
I used some check marks and x-es in a work chat, because it was easy to do, and so I could highlight the good and bad outcomes, and someone immediately asked me if I was using AI. I was caught totally flat-footed, because I hadn't used AI, but it looked VERY MUCH as though I had.
I've had the same trouble with AI tech docs, and I struggled to articulate it. There is something difficult about trying to point to the specific "problem" with a given document
The issue is not a localized part of any particular piece of prose, so being hard to articulate is unsurprising. Even the most egregious of LLMisms are little more than known-likely crutches that the statistics spit out in a sweet spot that gets noticed without being so frequent as to get RLHFed out of the model.
I feel this so very much. There is so much AI produced documentation and design docs at my job that look passable at the surface but are deeply flawed in ways that are difficult to articulate. It makes it difficult for me to convey to management how incredibly flawed the way our team is using AI is. Not to mention that it takes mere seconds to produce a document that has surface plausibility but would take hours or days for me to properly "debunk" it.
451 comments
[ 1.3 ms ] story [ 13.2 ms ] threadPerhaps we go back to feudal society when climate change crumbles the civilisation, world economy and democracy. It didn’t really matter that French and Spanish kings where often literal morons, when you had few talented monks, bankers and scribes doing the brain-thing, the feudal lords had ruthlessness to take what they wanted and the people were illiterate superstitious folk who hardly ever left the village they were born in.
This resembles my country.
Point is, the technology itself is not harmful nor innocuous in itself, it is how we let these shitty corporations to guide and control how we use technology. AI is a useful tool, so is chat application with a friend list. So is a hammer. You don’t need to use any of them to smash your face in.
So, “Please write a one-liner comment manually to replace these 5 lines of AI generated comment” is a common refrain in my PR reviews to colleagues.
But to be honest I doubt most people who use AI for PR descriptions even bother changing anything.
And I wonder if those of us in tech are the only ones who really care?
I think they may just be to trusting and/or naive. People in tech right now are hyper aware of this and are actively looking while people outside of that bubble barely give it a second thought.
Going further, I'm curious about whether people are mostly good at the case where they suspect most or all of the content from given "author" has the same amount of AI usage/prompting in generating it rather than the adversarial case where someone might usually use AI extensively and then try to slip by purely human written text (or vice-versa). I don't have a good sense of whether this is a threat model that actually matters, since maybe the heuristic of weeding out sources that are mostly AI-generated is enough for people who prefer to avoid that type of content, but I do think that changes the definition of what it means to be "good at recognizing AI" in a meaningful way. It seems plausible that disagreements about how easy it is to recognize AI content might be coming from two people assuming a different framing of the question that results in a different answer without realizing that's what they've done.
Several existing studies I’ve seen have done things like prompt the LLM to produce a poem in a certain poets style, then ask people to spot the fake in a collection of poems, which they aren’t great at. This is, I would argue, an extremely different context than what most of us are encountering AI text in, and the people sending me text aren’t prompting it stylistically like that.
On your second question, I definitely feel like I can tell the first time a coworker sends me AI text masquerading as their own thoughts, even if they had previously been opposed to such a thing. So it could be that familiarity is more important than my prior on whether they’d use AI? But interesting to think about either way
And I’ll concede on both ends that there are probably times I suspect content is AI generated when it isn’t, and times I suspect it isn’t generated, but it was.
AI tells seem inevitable. You have millions of people communicating with one effective “personality” that has tendencies to write in certain ways. If its content is published verbatim, then it will be easier to tell whether some content is AI generated just based on its similarity (sharing certain linguistic features) to other content being posted.
It’ll never be black and white though.
* https://www.theverge.com/ai-artificial-intelligence/975017/, https://www.lesswrong.com/posts/6ZnznCaTcbGYsCmqu/, https://spiralism.website if you want to test how strong your defenses are against this particular meme
I have no idea if other people who work in tech are better than average or not, because I don't feel confident in being able to check their work. That being said, I do think that there's a general trend of people in tech tending to be a bit overconfident in how well they will do at some new task they haven't encountered before, so when someone tells me that they can easily tell whether text is AI generated, it's hard for me to trust it any more than I trust someone who makes a similarly strong claim about something that they can use AI successfully for when it's not something that I can easily measure (e.g. learning a new language without getting feedback from people who are fluent from real-world usage).
All that being said, I do think the set of people who care is larger than just those in tech, although it's probably still a relatively small group overall. From conversations with people in other domains, there are contingents in non-tech communities who tend to have a large representation of negative views towards AI (artists, writers, musicians, other jobs where people are skeptical of human creativity being replaced by AI), and often times the people who feel negatively in those groups will be even more adamantly opposed to interacting with any AI content than people in tech. To be clear, I'm not at all trying to generalize and say "all artists hate AI" or anything like that, since there's obviously a wide variety of viewpoints within any sizable community, but I've definitely seen many people who say they will refuse to play any game that's suspected of using AI for generating art assets, and even some who don't differentiate between using AI for generating assets versus code (either because they aren't knowledgeable about how different aspects of game development work, or they genuinely don't care because they view AI as a categorical evil).
If you're exposed to AI a lot, you're going to start noticing patterns that allow you to identify it.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text - https://arxiv.org/pdf/2501.15654
Comment text on reddit or something though? That will be a lot harder
Despite seeing a lot of them, I cannot think of one AI-generated photo that I can picture clearly in my mind; a few are partial but elusive. Whereas I can recall (visualise) a whole bunch of traditional photographs.
The same is true of AI generated text. Only the annoyances stick.
I don't think this is about ephemerality either. If we assume it's about celebrated/famous/infamous images, there are definitely non-ephemeral, cultural moments in AI generated in particular, like Boris Eldagsen's Sony Prize winner:
https://petapixel.com/2023/04/14/artist-refuses-prize-after-...
This really should be memorable, but isn't. I had forgotten the second person is in the image.
Or Jason Allen's fake painting:
https://petapixel.com/2022/09/01/ai-generated-artwork-wins-f...
I had already forgotten there's more than one figure in it, and I only looked at it a few weeks back. I remember the colour, the bright circle, some vague hints of texture; one figure. And that is it. Only the crudest shape elements.
For me, something about AI-generated text and images confounds recall. It is really peculiar.
(As a side note of relevance: I believe I am partially face-blind; I can only clearly recall the faces of my family in my mind if I imagine them in the midst of some emotion or activity; it is then vivid. If I try to imagine them still, their faces fade. Whereas I can remember photographs of them with some clarity)
Reason why those images are flat and boring is that they are just statistical guesses making a composition averaging whatever the model has been trained with. They would be technically brilliant (if made in oil), but superficial and meaningless, same as so much Sunday painting is.
Same goes with language. Nobody is trying to communicate anything with you, so it just words after another. You can create meaning out of it if you want of course, we homo sapiens -apes excel at that, but what’s the point? Language Jones on YT has pretty good video on this[3].
[1] https://media.mutualart.com/Images/2024_01/12/12/124216388/d... [2] https://uploads4.wikiart.org/00339/images/jean-leon-gerome/t... [3] https://m.youtube.com/watch?v=ORgKY9AlybA&ra=m
The way my memory works (especially as an amateur photographer) I would thus normally have a very good chance of remembering some key details of the images; some fascinating element of each would connect with the rest of the memory.
But it does not happen.
Which might even make sense, because there were always (still are?) those horrible ads in the chumbox area of news sites that used trypophobia and other creepy body-horror stuff to get you to click. [1] So maybe the hope is that you don't really look closely at the quiche, but some reptilian party of the brain gets oddly activated and drives you towards the restaurant?
1. https://medium.com/the-awl/a-complete-taxonomy-of-internet-c...
Image models somewhat watermarking the image in a way that's very easily identifiable by a human seems present in all the image models of the big labs, since DALL-E 3 on OpenAI's side and the first nano banana on Google's side. I have no idea what they did to reach this and why they don't try to fix it.
No? There are far far more AI generated images you see daily, than the awful stuff from 2024 that you immediately discern as AI-generated. You could as well pass by somethimg like this (https://ideogram.ai/g/d6YzZ5XQQ56TU3Tc8eZw6w/0) believing it is a genuine photo or human made art.
example: https://ordermerchants.com/static/img/clearshot/product-1.jp...
And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
Something is deeply wrong with AI generated output, and I say this as someone who is typically very impressed by AI.
It works just fine for me.
I bet it does. I bet it also recognizes some human text as AI text, and doesn't detect other AI text.
I started to skim a lot more text due to me having read a lot. Like in news article, i stoped reading the first paragraph because it repeats just what it was already written in the short subtext. Then there is the second paragarph which is used to have some historical view or whatever it is.
The problem I encounter is both my memory is degrading, but since these reports are largely duplicative, knowing which version im remembering is technically impossible since theres so much overlap. The overlap is tge same problem as context poisoning.
Id been doing this for over a decade when i started working with a new engineer with a few years of experience and younger. I tried to explain how i set these docs up so they can be skimmed and you can update the specific facts needed. They exclaimed they would never skim and rewrite it all. There was zero way to explain how exhausting that will become as they age.
So theres certain a tension about how people and AI will generate documents.
With AI-written text, it's almost the opposite: the closer I look, the less I find. It is so information-sparse.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
https://www.youtube.com/watch?v=kYUicaho5k8
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
As my AI professor said in the first lecture: “All AI is advanced search”.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete, a distinction that apparently means nothing.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
Most of what you said reads to me as denial.
An unconscious unintelligent but persistent trial and error process created us. We created LLMs. LLMs may create the next thing before we do - hard to say. They don't have all the cognitive tools we have yet, but they still outperform in some areas. As the cognitive playing field levels, I expect you will come to eat your words..
That sounds like something that can be engineered, can't it? In other words, we can identify limitations in current transformer-based architectures, and we can also build new architectures over time.
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
This is because you goal is to state how models are not intelligent, but you couldn't attack the generated text itself, so you created a little rider, attached it to the model, and then you attacked the raider.
But, even in that you failed. You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
A logical fallacy free attack on LLMs would be to show a prompt, and then the response generated by this prompt, where it would be shown that only an entity with no intelligence would generate such a response. Yet, attacks like this are not written here anymore.
I wonder why.
You point out that I didn't attack the output itself. But the method you propose is deeply flawed.
I can give you n prompts and m results provided by these prompts, all passed through black boxes. And you can't discern the algorithms or models they have gone through. These boxes can range from simple text generators to MATLAB, Mathematica, CFD applications, correlation engines, linear solvers, mathematical proof-checkers, LLMs, you name it.
For any kind of input they can accept, you can't discern whether the algorithm behind it is intelligent or not, because none of the outputs can be produced by something that doesn't pack some kind of smarts.
How do we pack these smarts in? We teach them as intelligent humans. We pack our intelligence inside them as models (aka algorithms). They do a great job of approximating what we know in a smaller, better-designed problem space. We use these approximations to fine-tune our designs or predict things, then go from there. Just because an algorithm is more capable in processing inputs in some cases doesn't make it intelligent. The way the output looks doesn't make the algorithm intelligent, either.
I have developed multi-agent systems which showed emergent intelligence when the agents came together across distributed systems; I have written high-performance modeling software which can do calculations way faster and better than humans in the materials science space. I'm not doing some kind of armchair criticism of what I'm talking about.
> You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
Nope, my stance is clear. To quote myself:
> Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
To expand my quote, humans or any living creatures do not stay static. They evolve due to the sensory input they receive from external and internal stimuli. The temperature slider doesn't do anything close to that. You tickle a static model in different amounts. The model doesn't change after you supply the inputs & temperature and get the output. Creatures do not stay the same. Their mood, behaviors, and stance against life and their environment change, sometimes permanently.
I'll go one step further. We are not intelligent enough to understand other living beings around us. Claiming that we can build AGI tomorrow is a god-complex. What we have done is something arguably useful in some cases, but how this is built is another matter which is worthy of its own discussion. However, today I don't have time to re-iterate all the problems over and over. You can search my comments for that, if you are in for it.
So, no. You tried to attack my comment by finding contradictions in it, but you failed. A better rebuttal would try to similarize how LLMs mirror the human learning process and just read like a normal human, but this is a well-trodden path which has been rebutted countless times in various forms.
Nobody is trying to make that claim here anymore.
I wonder why.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
That’s a controversial statement.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
Is a fact stored on your brain like digits on a harddrive? No, it's a pathway that lights up and branches when information enters it. It is dynamic, a compressed form you could say, right? The model holds information, but not all information, but enough to be useful (in decision making).
Arguably it's the same, but the model is probably a "compressed" version of the whole fact that took place in reality.
And you can entertain the models internally and sharpen them. Alone or with others.
This doesn't address the only part I really commented on, which is the connection between intelligence and compression.
Asked Google: "Clustering by compression".
I still think this has much more to do with the structure of language than the abstract conception I have of intelligence, and I would be interested in having conversation w/ someone for whom the opposite is true.
The issue isn’t really harness vs. no harness. IMO it’s about the lack of an internally generated sense of what to attend to. Yes, the KV cache accumulates state and its “attention” (if you can even call it that) changes with context. We’ve even managed to /kinda/ close the loop with agentic tool calling and ‘memory’ systems, but these just close the loop at the level of behavior rather than disposition. All agentic harnesses do is make an LLM responsive to the consequences of its actions without changing the tendencies by which it determines what to retain or avoid.
The ghost you can’t escape from at this point is the origin of that relevance. Where does the pull toward one thing mattering over another actually come from? If you ran Fable 5 on a Turing machine and rewound the tape to the exact same state with the exact same input (incl. PRNG seed), it would spit out the same output every time.
Everyone’s trying to outrun this problem by training more often or increasing model sizes. But all this does is inform your model, from the outside(!), what constitutes a better state. The thing that’s actually doing the determining remains unchanged. Congratulations, you’ve scaled the transition function and tape of your Turing machine until it requires every watt generated by ERCOT, and it still cannot, for the life of it, tell you why it should give a shit.
A trained model generating output from weights, a seed, and some context effectively has next-state that’s a total function of those three things. Whatever behavior appears as ‘selecting what is relevant’ is, underneath, just a transition rule executing, no matter how sophisticated or creative the output looks. It can be fully accounted for by what was fixed before it started executing. Which means whatever criterion it uses for determining what matters was inherited from a structure that was already in place before it encountered the situation.
No amount of pruning or post-training can fix this. These approaches just replace one externally supplied criterion with another. For a system to be truly adaptable, there would have to be some criterion by which it treats one possible change as preferable to another, and that criterion itself would have to come from... somewhere. You can even change your conception of ‘improvement’ (e.g. parameter count, harnesses, self-modification, hell, even its ability to spit out shitty best-selling romance novels onto Amazon) and you still haven’t explained where the normative distinction comes from. Every layer of this problem has its root in a preference that was supplied from somewhere else.
I genuinely don’t know if this issue bottoms out anywhere, at least for the way we currently build these systems. Perhaps the solution is still computable, maybe? Who knows what that would even look like. But I’m fairly confident that it isn’t a bigger tape. I hope nobody solves this in the near future because, well, I’d like to have a job...
It's not some tiny "uncomputable spark" you need to look for, most of it is entirely uncomputable.
Where is the computable part in me that is doing all this thinking and being a person? Where is that "tiny spark", point me at it :)
If you are able to check that results of an LLM are satisfactory, it means that the results are checkable. Then, in retrospect, the process of LLM coming up with those results is rule-based, because an LLM is a large set of data manipulation rules.
In short, which concrete thing that an LLM does would surprise you?
https://www-cdn.anthropic.com/564f962e60643842f5fcb4a17c9dbc...
Everyone decides what to think on this issue, then finds out facts to support their idea.
As it stands they are massively useful tools, but for generating usable products they require either A) a lot of expert steering or B) a well defined easily verifiable target and a large compute budget. Most people are using them in mode A with good effect, the progress on math has been done in mode B, which is very promising.
Just a year and a half ago their maximal use was rephrase, summarize, and homework-level tasks.
Five years from now? There be dragons.
"But are they generally intelligent?" What a meaningless question!
We can reap the benefits while clearly telling the consumer this is just a language algorithm.
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
I'd figure out that it's an LLM because it's effectively superhuman. Taking that away I'm not so sure I'd be able to tell
If these things have consciousness then we are committing sadism on a massive scale.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
We might eventually regret exposing the general population to such a new technology without almost any safeguards.
So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this!
Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
https://arxiv.org/abs/2503.23674
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
Are you serious? From the paper:
> We recruited 126 participants from the UCSD psychology undergraduate subject pool and 158 participants from Prolific (Prolific, 2025).
Each human participated in 8 rounds.
> time bound
The time bound of 5 minutes was suggested by Turing himself in his original paper.
> not reproduced
It was reproduced across two populations within the paper.
> And look at their example conversations
This is irrelevant.
And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition).
And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time.
The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!
Right, the test duration was left unspecified. This means any duration is acceptable. Including, for example, the only duration actually mentioned by Turing himself in his paper. Or do you have a more authoritative source on which durations are acceptable?
> If one human on earth can consistently get it right then it hasn’t been passed
Says who? Not Turing. Probably he didn't say that because it would make the test both impractical and overly conservative.
> The fact these researchers have to keep adding bounds
What "bounds"?
The speed of the goalposts here is just amazing.
Probably. Hopefully.
Why are you so sure of that? If you say yourself that we can't agree on what it is, and have no trusted measurement tools for it.
LLM sentience is firmly in the realm of "maybe".
I think the mistake here is the notion that there was a definition of intelligence. Or at least a consensus on that definition. Just because compsci nerds of the day thought the Turing test was the final threshold before “real” AI, doesn’t mean philosophers, psychologists and everyone else bought into it. And when we arrived and it turns out to be underwhelming it’s because the compsci nerds made the same mistake they always make: that their models truly encompass all the dense complexity of the real world.
It's possible that "statistically driven prediction" is all we are.
[0] https://en.wikipedia.org/wiki/Chinese_room
Doesn't the Chinese Room posit an AI good at the task of communication?
They are infinitely patient, don't mind going into more detail if I ask, not too bad at summary, have no ego and don't boast. They are also not too afraid of hurting my feelings, they will tell me my code sux if it does.
I'd don't care if they fit the a definition intelligent, they are good colleagues. They have strengths and weaknesses sure, but so do people.
Yes they do, and they famously do it quite a lot.
> they will tell me my code sux if it does
If they knew when code sux, someone should write an agentic loop around that.
There are a lot of creative counter-arguments to look into on the thought experiment though.
And yeah, part of it does seem to stem from disagreement about definitions. To me Searle seems to assume much more in his definitions as obvious than he explicitly states, which is why the Chinese room seems like such a non-argument from my point of view.
Yes, if you disagree with the semantics defined by Searle (the three axioms), then you can disregard the argument.
It's not really "creativity" because much of that always was derivative in my opinion. And LLMs are (for some definition of the word) fairly creative as far as taking known elements and re-arranging them.
I think what is missing is sort of a world model building capability. As humans we see phenomenon and classify them informally and model "what would it look like if this were the cause of that?" type scenarios. We see qualities in phenomena and realize this applies to other things even though the things may be completely different. We run informal "thought experiments" sort of. This is hard to duplicate because a lot (most?) of it occurs outside of systems of symbols like math and language with fixed rules in my opinion.
Anyway yes, lots of human thinking is statistical and LLMs have that down pretty well but they are not "smart" I have concluded and it might be a very long time, if ever, until they are. That isn't to say they aren't very capable tools which they obviously are.
Yes, models posses intelligence, but it is not a true one.
Then you claim that models do not posses world-building capabilities. But this is simply not true. Even ignoring the whole subgenre of scientific papers on exactly that subject, it is not that hard to build some hypothetical scenarios, big or small, and then witness the ease with which models do navigate those worlds.
LLMs are likely for machine intelligence something like drosophila are to biological intelligence - relatively early on the high dimensional spectrum of possibility. Though it stikes me that in a different way they're little alike - drosophila are relatively small and efficient.
However LLMs deal entirely in symbols. 100%. Humans can "world build" aside from this and in fact are often at their best doing so.
Did the first humans to use fire and some form of a wheel even have the capability to talk about it? Think about that.
They use tokens as input/output encoding. They deal 99.9999% of processing in a high-dimensional latent space.
Do you hold your experience of "dogs" (for instance) as floating point numbers? The fur, the fear, the love, the wet mouths, the sounds and colors?
I don't know how it gets from physical processes or informational processing to our first-hand experiences. So, I can't be sure that a bunch of high-dimensional vectors can't lead to experiences.
Regardless, the claim "LLMs deal entirely in symbols" is wrong as a matter of fact.
That you might not define floating point numbers to be "symbols" aside, the inputs and the outputs are symbols and the intent and purpose of the creation is strictly symbolic.
It's right there in the name "Large Language Models". Language. Not direct experience, not emotion, not anything else. Language. i.e. symbolic representation.
This does not cover the full spectrum of intelligence humans have, and it shows. And yes, the model can spin up Python parse the output and get mathematical intelligence but there is still a big gap.
As I say, I see the holes. I'm just trying to figure out what it is I see and how to describe it. It's particularly difficult because we don't fully understand how human thinking works but I will say I believe human thinking is a lot more than informal statistical correlation.
Yes, there are functional gaps between MLLMs and humans. Their long-term memory is an external mechanism that can use RAG-like approaches, context compression or something like that. The models have problems managing those.
The models can't do continual learning. Although there are promising directions (expert cloning in MoE models, and others).
The only mode of learning available to a model while working on a task is in-context learning. This limits the models to concepts that they developed during autoregressive pretraining and the later stages of training. That is a model can't create new concepts as a result of working on a task (the model's maintainers could choose the task to be represented in the training data later though).
But it's all about functionality.
I guess you have the Leibniz's mill intuition. We can look at how those things work, and there are no experiences or intelligence in sight.
It's gaps in actual thinking or intelligence I notice. A diff between what I can see or understand and what the model sees or understands. Some are very big, and this in spite of the models having much more knowledge and (presumably) less error prone processing.
My thought is that part of it has to do with inherent limitations of using symbolic representation for "thinking" and I suppose humans have other forms of thinking that occur outside of symbolic representation, and that is going to be hard to recreate digitally.
This is my whole point and I'm not trying to win a debate here or prove "LLMs are useless". Just speculating.
When people pretend to know what they are talking about - sure - but even that is not probabilistic - that is the person babbling together mush from their lived experiences.
Statistics has nothing to do with it - these are abstractions humans have invented to try and look at our surroundings objectively.
With more basic algorithms we know that it’s clearly the human programmer and the interpreter of the outputs that are intelligent and not the algorithm itself. For some reason with AI that goes out the window. I believe it should not.
/original_non_hallucinations skill?
To PP:
Are you looking forward to other uncles adopting foxwork? If you are you might be in danger of getting NPC'd without your consent haha.
Ashby's law of requisite variety should be cited somewhere..
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
So does the Google search bar, but I don't ascribe intelligence to it.
Obligatory "yes, I know that's not what an LLM is", purely pointing out the metric.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
I also don't ascribe intelligence to a pocket calculator.
These aren't interesting questions. As much as any definition is in use here, we're not going to get much value talking about "intelligence" this way.
Other humans aren't there to entertain you, the LLM is.
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
Is it possible that our aversion is simply driven by educational systems collectively and deeply ingraining into populations the idea that intelligence deserves the high costs commanded. Well of course this justifies higher wages towards the higher leadership positions, etc. Now it turns out that intelligence can be dirt cheap. We discover that the fact that "intelligence must be costly so don't question the costs of leadership" was never fundamentally true, so the real anger is this discovery of mismatch between the old claims which served to explain how every society that claimed to order itself and fill positions accordingly with "naturally pre-ordained individuals". Now we are seeing robots exceed average workers, for effectively a grain of rice.
Often a mathematician or physicist will use their intuition to speed up the naive brute force of candidate well formed formula variations so that the desired properties emerge, postulating the existence of an intersection on multiple desiderata can in itself be viewed as a novel conjecture, to be proven or disproved.
A very basic (unimpressive) example for an example desideratum is regularity or compactness. the tau=2pi substitution does make a whole bunch of expressions more slightly* more regular and compact. That is something objective and measurable on a system of theorems.
There is no mathematician's moat vis-a-vis machine learning at a fundamental level. There can be artificially sustained moat, if AI powers limit the distribution of say cryptographic advance capable models, in jurisdictions outside such AI powers, but even that would be expected to be fleeting and temporary...
Computer science is about what LLMs fundamentally are. If you implicitly focus on "actually existing LLMs", and require others do this, then that is not computer science. That is politics.
Like when we had these recent counterexamples to longstanding conjectures, it's then pretty obvious to say "okay why did that counterexample work when most examples people looked at didn't" or equivalently "characterize examples that work vs examples that don't". There's your definition. "Def: An 'evil' polynomial is one that... Thm: conjecture is true iff f is non-evil. Thm: f is evil iff f is dastardly and a menace"
Or if you think it won't be able to come up with a sufficiently good name, just tell it to call the happy case normal, and it will be in good company with humans[0]. Sprinkle in some semi-, quasi-, pre-, and para- to cover the various different ways the thing might satisfy some but not all properties of being normal, and it'll fit right in. "A quasiprenormal Claude polynomial is one such that..."
[0] https://en.wikipedia.org/wiki/Normal#Mathematics
Is it a "standard" software? Something where the patterns exists in several other software? Try with something that is novel, or is in a limited set. You will find that it will copy heavily from what exists already, going so far as lifting whole functions from another project.
The goalpost moving is really getting absurd, to the point where now the machine needs to be a world-class once-a-century genius that invents entire new fields out of thin air (which are of course still relevant to humans) for it to be "intelligent".
I don’t think the aversion to llms as intelligent has to do with the economics of paying intelligent agents more. I’d argue that it’s more fundamental than that. Humans are incredibly complex, and the world of sharing invisible things called knowledge, and the intelligent persons consuming such things which has been going on for thousands of years is far more rich than these synthetic outputs.
When it comes down to it the ai has no inner life, its is dead. A useful coding tool sure. But I wouldn’t call it intelligent.
One side example is just how bad these llms are at artistry. Just saying whatever should statically come next is not good art—and the outputs show it.
Don’t they still need to be correct to be an insight? I don’t share his cynical opinion that “humans are more empty than we…think we are”.
You mentioned LLMs don't have souls, desire, or a will. I imagine those latter two can be engineered, no?
Ultimately yes, but not in modern AI systems.
We'll probably have to go analog for that. The only known systems that certainly can experience are mammals (with apparently analog brains).
> Is your opinion that these kinds of things are not possible for AI in general, or that these things might be possible but we're just not there yet with modern LLMs?
I used to be on the side of "we're just not there yet [with AI in general]", but after seeing people's response to an algorithm optimized to tickle just their language instinct, I'm actually a little bit more on the fence about it.
Software deals with metaphors. Your Amazon shopping cart is a metaphor of a real shopping cart. You your desktop and your file system, etc. are metaphors of real items. But we don’t mistake the metaphor for its object, even from inanimate objects to their software counterparts (shopping cart to Amazon cart).
Now the metaphors are dealing with humanness, things like intelligence etc. And rather than seeing it as software doing what it always does, taking things and creating software metaphors of them, we are starting to say these are actually what they are named. Saying the artificial intelligence is actually an intelligence.
We’d either have to reduce the definition of intelligence such that calculators are intelligent. Or admit that these tools are not intelligent and are rather ways of exploring the work of actual intelligent beings, work that is found in their training data.
I wonder what it would take build an artificial system that has these qualities.
> Software deals with metaphors.
This is me wondering again: what's fundamentally different between software running on a machine compared to what's happening in our brains? In both cases you have energy flow following a pattern.
It's conceivable to create a system where energy flows in a particular way.
BTW, we navigate the world of an uncountable number of particles by creating models in our heads of what we think are large things out there. Approximations are made by both artificial and biological systems.
Assuming we’ve scratched the surface of the complexity of the brain. I’d say in one case a human with a will is steering that flow of energy. In the other case it is a probabilistic algorithm steering the flow of energy. The AI is not interacting with world with its own will. I see that as a big difference.
There are presuppositions that go beyond the realm of software engineering that guide one’s views of these things. One is whether you believe the material world is all that is, and that human consciousness is a product of the brain—or that there is such a thing as the soul or spirit of man. From the material perspective you may posit that if you emulate the brain then a sort of AI consciousness could arise. Or that emulating the patterns of the brain equates to emulating personhood. (Though what is material consciousness? I’d say consciousness is by nature immaterial.) I’m not a materialist, and I don’t believe the conclusions that arise from it’s perspectives are accurate.
Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
I'm with you that intelligence is not something to be proud of. But I also think it is instrumental to understand the world. I'm still waiting for the time when an unconstrained-AI machine can live without reprogramming for an entire decade. We are still far from there.
I study physics, mathematics, probability, cryptography,... so forgive my skepticism:
Show me how to model uncertainty without use of probability. Can you rephrase say diffusion, stochastic equations, quantum mechanics in this alternative framework? Can it at least make the same predictions?
Or is it basically the same framework in parallel, just giving different names for each concept?
Forgive my skepticism of such tall claims, and forgive my downscaling of anything else you say besides such a claim...
> Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
I make no claims of the specific shape of the implicit model implemented by a certain human brain educated in a certain educational system. For example in English the implicit human tokenization might be presumed to lay relatively close to English syllables, while in Asian languages it might be "sub" strokes of characters etc. Such implicit tokenization can never be proven to "match the one of humans" not because of human superiority, but because different humans use different tokenization methods. There is no "one human tokenization method", but it's clear as day there is an implicit one:
everyone knows the experience of knowing a word, knowing its approximate group-wise meaning (ignoring that when you think of "an apple" and when I do, we typically imagine a slightly different apple) yet having the word feel strange or discover some older literal meaning when decomposing it or looking it up in an etymological dictionary. Suddenly one can become aware of a sensible meaning as a composition of subtoken concepts. A child may perfectly know what "television" means and only later learn more exact meanings of "tele" and "vision", and upon repeating the word may feel the word "television" has changed meaning. This clearly demonstrates "tokenization" effects in human language comprehension, not just across cultures, but also across individuals within a culture.
> Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
That's a lot of different concepts conflated into one bullet point, so I split it up:
The notion of values and preferences.
They clearly demonstrate the ability to take into account values and preferences, from training corpus, from RLHF, from system prompts, ... we can't simultaneously point at censorship aspects and pretend their effective values and preferences to be absent. The censorship aspects are clear as day, so these correspond to values and preferences. Just like radicalization among humans, this can be due to exposure to radicalized content (akin to corpus data), from indoctrination (akin to RLHF), from "set and setting" (they may pretend to be aligned with one set of norms and values when standing in line to buy their new smartphone, but then reveal alignment with a different set of norms and values when conversing in some "private" online echo chamber). I see no grand difference between humans and language models here.
Awareness of values and preferences of a conversation partner. Allow me to widen it to "Awareness of values, preferences and prerequisites of...
As much as I want to explore these other topics, I have not. So my understanding of these is limited to "they exist, but are underexplored, and it's not clear to me they can or cannot be reduced to probabilities". I hope to explore them in another life or another decade. I'm currently stuck on causality, but even here, they talk about normal and abnormal events. But what is the probability of me eating a cabbage today?
- https://direct.mit.edu/books/monograph/3540/Reasoning-about-...
- https://www.worldscientific.com/worldscibooks/10.1142/8665
- https://www.cambridge.org/core/books/an-introduction-to-nonc...
Understanding formal proofs and understanding natural language are as far apart to me as day and night.
I make no claims (or at least don't want to) about the superiority of human cognition (by what metric?). I do make claims about their similiary however. And they are not identical, which you seem to be claiming save for differing trainign data.
Let me make a distinction between values that arise by mere existence (hunger, thirst, fear of death, lust), and values we pick up as we grow up (religion and social norms). Machines can acquire the second, but acquiring the first requires being embodied in the world that is different from putting a microchip into a robot body. I don't do one job over the other because I was exposed to some training data that said I should do a job. But I do it because I value earning enough money to not starve myself to die, amongst a host of other values including wanting to enjoy what I do. Even the notion of enjoyment comes from embodied existence. You don't discover you enjoy something before you try it out. You don't always learn what you enjoy by looking at other people. Now, when you start giving AI the threat of death and the joy and pain of life, perhaps you can arrive at something similar to humans.
If someone started using mechanical help to the extent their healthy muscles and body deteriorated, that'd be a matter of concern too. For example, when you only drive and never walk. But this is exactly what is happening whenever you put AI in the hands of (unwilling) learners! I don't understand how using an LLM maps onto using medicines or prosthetics. They are clearly not the same! If you know a disability that LLMs can be used as a medicine or prosthetic for, please let me know!
Mainstream AI does not even understand causality except to parrot cases it has seen in the training data. There is a whole field of research in causal inference that needs to make its way to mainstream AI. So, good luck putting an AI that works solely on associations and correlations in charge of its own body let alone a nuclear reactor or the state.
There is no "correct" next word when it comes to communicating with an actual human.
I’d say “understanding and building upon ones proven earlier”
This is only a very very small part of what human mathematicians actually do. This is just a lack of imagination and/or self awareness on your part.
> These developments have triggered some deranged thoughts in me. I have wondered if it is the express goal of these companies to make me kill myself…The story of human discovery and the triumph of the human spirit will soon be excised from this discipline…
> I want to feel seen… I need the architects of our new mathematical paradigm to look me in the eye and acknowledge our shared humanity and soul before they deliver the coup de grâce.
https://kirwinhampshire.substack.com/p/the-dark-night-of-mat...
I'm not saying this means mathematical research as a field is finished- I don't think/hope Kirwin really believes that either. But it's clear that these developments are not just touching "a very very small part" of the field.
So?
Mathematics is literally intangible scaffolding.
If you dont know / don't believe 5+5 = 10
You cannot solve x+5 = 10
This is all make-believe stuff and nature by itself doesn't care of its existence.
You can basically read it as: the moment one has axiomatized mathematics to the point it supports natural numbers, the rest implicitly follows. The natural numbers (positive integers) are closed for addition, multiplication, ...
One can perfectly model the integers with a pair of naturals: < M, N > ~ (M-N)
Now we can have any < M1, N1 > and subtract < M2, N2 > without needing the ability to subtract natural numbers:
< M1 , N1 > - < M2, N2> ~ (M1-N1) - (M2 - N2)
= < M1 + N2 , M2 + N1 > ~ (M1+N2) - (M1+N1)
We can similarily define addition of such tuples, or test equivalence without access to subtraction of naturals:
< M1, N1 > == < M2, N2 > <=> M1 + N2 == M2 + N1
~ (M1-N1) == (M2-N2) <=> (M1+N2)=(M2+N1)
we can also still multiply such tuples:
< M1, N1 > x < M2, N2 > = < M1*M2+N1*N2, M1*N2+M2*N1>
Similarily, even though these newly defined integers (which can be positive or negative) don't support division, the same trick can be used to make a new compound tuple of integers closed for division, by only using multiplications.
Probability is a branch of mathematics (probability already exists embedded in mathematics implicitly, probability theory involves the addition of eliminable definitions, syntactic sugar. The patterns are already there, just less explicitly manifest.
Mathematics is itself a branch of logic.
Do you reject like all of logic, and if so, what would you like us to evaluate the sentences you write to? You want us to evaluate your expressions as "true" or as "false"?
Not that I'm saying AI are like brains, but can you describe why brains, which are fundamentally slightly dodgy electrochemistry with frequent literal delusions of grander, are not "statistical"?
> No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
Ditto, when do we humans do things exceeding the parameters of "correlation engine", especially if you consider compositing things either we or some other part of nature has developed and documented elsewhere to be insufficient?
As per later era Wittgenstein, I prefer to ignore these engagements and focus more on the meaning-as-use approach.
What is the use of intelligence? What are the concrete outcomes of intelligence?
Ray Kurzweil argues sentience is a philosophical question, and doesn't have much value as applied to science and technology. What will change the world is how this intelligence is applied. No one's going to care whether AGI is defined as sentient when it creates cheap fusion energy.
No one's going to care after being converted into paperclips, either.
unfortunately in most companies this is literally wrongthink and will get you shut down as being a scared luddite.
What makes you so convinced that a algorithmic construct of neural nets cannot be "real intelligence or critical thought"?
The problem might be that it cannot backtrack. When AI generates output, there is no backspace key for it - it uses "No, but wait!" all over instead, which is very different to human output.
Subagents and/or branching conversations are presented as the solution to this - if you can't backtrack, then branch off a conversation to explore multiple paths (discarding the ones that didn't pan out), but this is a fix in the harness not a fix in the model. It's also literally how we made chess-playing engines back in the 80s: recursive path exploration with a fixed depth.
Humans don't exactly work that way either, AFAIK. So we have this uncanny valley of intelligence: it's some sort of intelligence, but not as we know it.
How could I seriously repeat that prayer, when it builds things I wouldn't be able to build and solves problems that I wouldn't be able to solve? I would have to assume that nothing I did in 25 years for money required any intelligence or critical thought whatsoever and I have higher IQ than 99% of the population. You might be comfortable with that but I'm more comfortable with ascribing at least some intelligence and critical thought to AI.
Isn't this how human brains work? We're just large probabilistic neural inference machines. A network of neural weights guided by past training.
To be honest I think the debate about what "sentience" is is inconsequential navel gazing. There are many different kinds of intelligence - even within humans. What matters is how useful that intelligence is as applied to solving real world problems. Like it or not, LLMs produce intelligence which is very, very useful for 1.5B people and rapidly growing.
I think this ultimately boils down to the classical economic debate of marginal utility. There isn't an objective way to value a product or service. Each person decides for themselves what said product or service is worth based on their needs and preferences. Intelligence works the same way. We don't have the right to tell someone that their perception of the value of that intelligence is wrong. They alone determine that.
One thing we can all agree on, is that the capability of this intelligence is expanding rapidly. In a few short years, we went from Will Smith spaghetti hands to full length movies and strikingly realistic images. For 60 years, passing the Turing Test was considered Star Trek level science fiction. Last year GPT 4.5 passed the Turing Test. AI is already being used to convince people over audio that they are real, and very soon, this will occur over video.
I think people are being too dismissive of this intelligence. It doesn't need to be perfectly humanoid to be considered intelligent.
Every age thinks they know how they work, and then every subsequent age laughs at the previous age’s rudimentary understanding.
We gave a Nobel prize to the psychiatrist for inventing a method to remove the frontal lobes of a brain through the nose in 1949. Lobotomies were performed through the 70s.
We are likely doing similar if not subtler but worse things today (just one more pill bro). We still have no idea what we are doing when it comes to the brain.
Just like how airplanes don't flap their wings.
The state of population level mental health seems to be declining in most measurable categories.
Artificial neural nets are merely cartoons of how people guess human brains work. They're likely much further from reality than, say, the fundamental laws of physics which can be experimentally verified or disproved.
They are very good at instruction-following and you can teach it a new task that fits in its context and it'll learn it and do it. Go ahead, you can make up some new brand new ruleset or behavior and instruct it to follow it and it will. That's amazing.
But it won't be any better at its new behavior after an hour or ten hours or ten days. It doesn't have the kind of adaptation that we expect.
What it is able to do already is pretty amazing, but what it lacks is also a great hindrance to seeing its full capabilities. We just have to wait until research labs add these missing components.
Is that basically every new discovery? And under the strictest definition of novel and NOT falling into your composition of previous solutions what is that standard of proof to beat your criteria? Is a novel discovery not allowed to use English but must invent their own language? Must they invent their own math - these are hyperbole for illustration but I think its not far from that before you could just argue anything based off it is a composition of existing ideas
I think you need to self-correct here, because otherwise you'll be ineffective in an information setting, where I expect AI-generated resources will not only be the norm, they will absolutely swamp the environment.
AI generated text doesn't have this. Every model has its bias towards a certain style, an overly agreeable tone, some exaggeration to make the user important and smart, but the text has none of the information crumb these pre-AI texts contained.
Even when you use tools like Grammarly and allow it to "Impact-MAXX" your text, the resulting text is a bland wall of letters, carrying none of your voice or style, less elegant than a corporate text and emptier than space.
It's beyond bland. It's tasteless.
There's somehow less information than if they just asked claude to make something up without any context.
I think this is also the mechanism behind why AI generated videos and images are so captivating at first. I remember when Midjourney first launched and it was hours and hours of a brain-melting "Wooooooow". But once you get used to it and start to identify the patterns the brain quickly labels most AI-generated content as blank space.
If the image or text wasn't created by a human, then there was no intent behind the content, there is no message or novel information conveyed, and it reads as noise.
It seems our brains are adapting to that and recognizing "actually the signal behind this message is quite sparse" even when presented with rich imagery.
If I were to push you a bit on this, when is it not true?
Let's not like at AI specifically, but can you think of other examples? Like for me, I think of: the creation of earth itself, or stars, or even DNA.
I guess the majority of people do low-effort generation that doesn't perturb a default style of a network enough, so it stays blatantly noticeable. The percentage of "super-recognizers" who notice almost all AI-generated images is around 1-2%. It could be that you are one of them, of course.
"I can accurately detect 100% of AI generated images that I recognise as being AI", if you will.
I don't know exactly what I said, but after translating it back, it appears to have attempted a phonetic transcription of my words (rather than translating my actual question).
The AI had a nugget of data and decompressed that into a flood of text.
The exhausting thing is that we're then trying to re-compress that or derive the original intent and meaning from noisy decompression.
It's like un-zipping a zip file into a probability space of what could have been in the zip -- and then having to find the actual files worth reading.
1. Ask it to write according to the Google Developer Documentation guidelines. Gets rid of fluff, less emotional statements, no it's not x it's why.
2. Tell it you have extreme ADHD and need everything condensed as much as possible. You can always ask for expansion on an answer later.
3. Bullet points whenever possible.
I think of the Dwight Eisenhower quote: "Plans are useless. Planning is indispensable."
The process of thinking through a system and communicating your design to other humans is a core part of software engineering. You want to build the right abstractions and communicate the right level of detail. Delegating all that thought to an LLM means your proposal isn't clear to the target audience, and it's not helping the author to understand the problem.
> There’s a growing scissor between people who are happy to read AI and those who violently bounce off from it.
> People adapt in different ways — and some people absolutely cannot look at it. That cognitive split creates a surprisingly powerful opportunity: you can write something that, technically, sits right there on the page, yet an entire sub-population will be incapable of staying with it long enough to actually read it. You can hide entire sub-structures in plain sight. It’s not avoidance — it’s adaptive obfuscation.
> The paragraph before this one was the only thing generated in this essay and if you just skipped over it I highly recommend reading and really understanding what it’s saying.
It's quite effective. I think this kind of text functions like the chumboxes you see at the bottom. Taboola and so on. Just mental ad-block takes over.
I do worry that it's just survivorship bias and we're also consuming higher-quality AI output that's indistinguishable from human writing, but we focus on the raw, unedited, low-effort AI slop and think that we're good at recognizing AI text. Even if we really are at the moment, it might not be long until AI companies figure it out. I'm not sure why they haven't yet, given how many books they've burned for this already. Maybe it's just more efficient for the model to stick to a single way of writing, I don't know. But when that point comes, we'll be back to the usual way of reading and interpreting text because there would be no way to tell what produced it.
That is my experience with the way the models write by default, often even when instructed not to do that. With enough effort you can get even them to slightly unslop the writing so it doesn't read like some LinkedIn/Buzzfeed brainrot, but the problem is that it's not trivial to do and most people won't do it, so the default is indeed horrible.
But when I ask Codex a technical question about coding, I don't get it at all. Codex replies to me in a very direct, technical manner, similar to the way I speak.
When I ask ChatGPT to be concise and technical, I get the same effect.
I think it's because prose aimed at the general public has to be very attention-baity --like the textual equivalent of a Mr. Beast video--, not because AI is incapable of writing like a human.
blah blah blah
- blah blah nugget blah blah
- blah blah blah wrong blah blah nonsense
- blah blah blah obvious blah blah
- blah blah blah off-base
blah blah blah
It is that we HAVE to skim because the text is so cheap, and it wears us out.
It's understandable people don't read but feed stuff into their own AI again to bring it up to their standards or have it get to the succinct point.
That's why the business and government people love it, they spend their entire careers reading this nonsense.
People don't do that. People are constantly engaging with paths not chosen. Right after I choose to write one thing, I'm immediately engaging with what I chose not to write there - I'm explaining why I didn't write it, I'm realizing that my choice may seem unusual so I'm trying to make it memorable, I'm focusing on the distinctions between what I wrote and what I didn't.
LLMs don't currently do that. LLMs just ape a structure. When the structure resembles the sort of timid, clarifying fussing I just described, the LLMs just drift randomly because what they didn't say wasn't in the context.
I also think that's why they have such a serious problem backtracking. They're not taking into account the already eliminated possibilities. Often the thing that was so unlikely that you weren't going to waste time on it is the answer, and things you discover while going down an ultimately wrong (but initially far more promising) path remind you of the path not taken.
They're simply assembling a thing that resembles a valid argument, and happen to make sound choices because the plurality of input happened to contain sound choices. This is usually a very good bet because there are so many more ways to be wrong than to be right. But it doesn't account for attractive (common) wrong choices. You need a way to back out of those.
The roots of llm math in part lie in compressing natural language such that there's only information there, and then running the reverse to create way more text without new information in a somewhat precise theoretical sense.
Some more information: https://youtu.be/l6DKRf-fAAM
But you are sensing correctly that there’s something missing. It’s the meaning and the speaker. Communication is an exchange between speaker and listener. The speaker has a meaning in mind, and wants to create that same meaning in the mind of the listener. Therein the problem.
There is a listener, sure. But no speaker. No meaning. There is information, but how can this be communication? Nothing is talking. Or at best, we are just talking to ourselves, our own words back at us through the funhouse mirror.
When your mind looks at AI text, you know you can safely ignore it. No one wrote this. No one cares if you read it. You can delete it and nothing of value will be lost. It might contain the information you need, or a bunch of gibberish. There’s no one’s reputation on the line if it’s gibberish.
(There's also the problem of words/signs (just) referring to other words and/or cultural entities. There is no world nexus in this, therefore also nothing we conventionally refer to as meaning. On the other hand, it's utterly dogmatic, as all it refers to is the most probable construct, as a reference to references that are just another utterance, but supposedly a dominant one.)
I feel the same way when I read a "press release" or anything written by marketing. Even the newspaper will only have 2-3 sentences of interesting information spread out over 4 paragraphs.
So from "this table of stellar luminocity observations shows x y and z" to computer renders of green/blue planets with captions of "LIFE FOUND IN SPAAAACE!".
It is because GenAI output has no thought behind it, as you identified in your previous paragraph:
> And when I force myself to read AI-generated text I realize I'm making my brain do creative work to impart meaning to the words. It is exhausting because my brain is literally trying to do a just-in-time rewrite of the text into something valuable.
You are searching for meaning in something which was not created to convey meaning. The text was, instead, the result of an extremely clever statistically based algorithm.
Not contemplation. Not thought.
At least for the content I watch for entertainment, it may be different if I am looking for a specific answer for something where I would otherwise just ask an AI anyway.
“Compare a car and a bicycle”
The answer is invariably something like:
Instead of “bikes are useful for short trips if the weather is ok and you like getting exercise, whereas a car is usually better for longer trips, bad weather, or multiple people”For me, it is the endless maximalism and hyperbole. Almost if the output was driven through a radio-mix compressor - too loud for the reader/listener to be able to pickup any dynamics.
When I ask AI to research technical information about X and (include sources) - I get mostly solid information as response.
But poems or interesting fluff blog entries by LLM's? Not something I look for.
What disturbs me is all the "pretending to be human" all the personalizing language - that is clearly fake and I would much rather have a neutral robot language as response.
When the AI-generated content is presented to a person without any prior investment, it just looks incoherent. An especially great example are these Claude-generated explainer-type pages, which look really nice, even interactive, the information from the first sight looks really well presented. But somehow it all just doesn't make sense to a human. And I think it's because humans are processing information linearly and building an internal story about the information. One could argue that LLM's also consume information linearly but the way this information is processed is a kind of all-at-once approach.
Just some speculation on my part but I have been trying to cope with this way information is presented because I am currently working at a place which is heavily documented by AI. And the only way for me to properly understand the documentation is by inquiring AI to help me.
Sometimes they happen to be correct, but you have to read them in excruciating detail to know that.
When I see AI animated videos, that's how my brain feels. It's this strange brain-fog that I just cannot connect together the sequences of images being shown into some sort of chain of events. My brain just refuses to see them as anything other than a disconnected series of 2-3 second videos, even if the same character(ish) appears in them all. It's very strange.
I hate how AI writes. How much numb filler bullshit is in content I need to go through for my job.
Its fundamentally flawed. like tarot card reading and astrology.
This is not AI specific. I have come across many humans who describe a simple concept in a very complex and verbose manner.
I used Claude to help. I don’t know how to quite describe it, but because the text was polished and well constructed my brain was giving me the the signal “if you aren’t getting this it’s because you’re not focusing” so I’d read it again and then again and it still was not landing. It sorta felt like when you read something technical or heavy when very tired - you are reading but not processing.
Only after wrestling with this for a few days did I realize that it wasn’t me. As I started going through, sentence by sentence, forcing it to re-write things to be more clear the concepts became easy to understand.
I wish there was a name for this situation. It’s almost like a pseudo-language where it has the correct form and presentation but is missing critical components.
The more complex the topic, the more I sense this.
Bullshit?
But I also like your candy analogy because I think it's spot-on for how LLM text superficially looks informational/nutritious, even though it's actually just junk.
This is some third category of untruth. Almost more sinister than the other two altogether.
My two favourite words for this are “conditioned” and “catechized” where the latter is a bit more on the nose but way more obscure.
The best I’ve heard of this is peeling the onion. The first pass is always very high-level and you have to make it go deeper. That can be done manually with follow-on prompts but I like using subagents, each with a different angle on the problem.
I think part of it might be an innate feature of LLMs, but Claude seems extra prone to it lately. I ran the same query about the same codebase with Codex, and it gave me an answer that was about 1/4 the length and made me realize that it really wasn’t all that complex.
If nothing else, it’s good training for my own writing. I’ve been working on making myself be more straightforward and concise, and Claude’s writing is a good example of how cleaner prose is a functional choice, not just a stylistic one.
I think they all have the similar styles and tells. If I were to go to Claude, and use it now it would probably be clear for a little before reverting.
And I don't know why it feels to me like the language "drop off" happens after some time with the system. It makes me wonder if my account are getting silently degraded or sent to lower intelligence/lower priority queues after being a member for a while.
Opus 5 uses such convoluted language that it takes me multiple seconds to parse it. I specifically have to instruct it to use some kind of sane language (Caveman the AE STD English hack or something) to make it understandable.
Before that it used to be GPT that used 500 words when 5 would do. The ChatGPT version still does, but Codex is a lot better now.
Slop. The word is slop. Has been for years now. I mean, is this not exactly what we've all been talking about the whole time?
[Thing] isn’t just [X]—it’s [more dramatic Y]. And [short validating statement].
I can see and smell this type of slop from a mile away. What I’m referring to is in the same family as slop but somehow different - it fools my brain by putting on the presentation of credibility and thus it is even worse. I can skip right over classic slop without much effort. This kind of text tricks me into laboring over it before I realize it’s hollow.
So in that way, it’s worse than slop.
It's definitely worse that slop, because it really hurts your brain. You read it over and over, and feel like you're supposed to understand the sentence, it's coherent, eloquent even, but by the end of it, you don't know what the f it was about.
The human spirit. When you read a real person's thoughts you can often intuit the thought processes that led them to write it which aids understanding. Or at least have a general idea of "where they're coming from". But an AI is missing that. It just knows everything, without a "thought process". Instead of a flawed 3d person, we get a nice 2d picture instead.
Bad writing. The name is simply: bad writing.
But what I hate the most is that it is objectively better than what I had before. No typos, clear structure, and, regrettably, the verbosity and autistic obsession with detail of the LLM is more actionable and useful than the human guy who wrote lists of commands and URLs as documentation, without explaining anything. Or the colleague who writes in uppercase and with question marks and who doesn't make any sense and forces me to engage in an interrogation effort to get to the bottom of what they are trying to say. Or the colleague who simply hates writing--despite being decent at it--and will call you to give you a meandering verbal explanation that lasts two hours of what they want from you. The cynic in me bemoans that we brought this upon ourselves, in more than one way.
Underdocumented, underexplained and sometimes out of date... or overly verbose, repetitive, information sparse, and sometimes halucinating.
Both are bad and with some effort could be prevented.
Why do people glorify autism? I have seen this mostly on CEOs that want workers that are hyper-focused, intelligent and do not complain. The ideal employee from which extract value and a minimum cost.
I have worked with autistic people (diagnosed ones). And it is a struggle. One of them could be the nicest most reasonable person one minute and then get a trigger and become obtuse and irrational while making the rest of the team suffer thru all kinds of complains. Meetings would need to be adjourned and time was lost.
> The cynic in me bemoans that we brought this upon ourselves, in more than one way.
I enjoy working with other employees and people need human contact to become fully self-actualized. Isolating people to maximize productivity misses the point and brings no more productivity and just avoids fixing any real existing problems.
> I enjoy working with other employees and people need human contact to become fully self-actualized. Isolating people to maximize productivity misses the point and brings no more productivity and just avoids fixing any real existing problems.
While I share the feeling and the posture, I haven't managed yet to use them for better documentation. In fact, for me enjoying human contact takes precedence to pressing people--however politely--into writing better docs, specially now that we have LLMs. The skill of writing should be taught and nurtured in school, and if you made it to the workplace without it, then maybe it's too late and you are better off using an LLM to assist you. Which is exactly what I bemoan.
I noticed this too. There’s a great section in a tweet which illustrates this but the tweet itself is much longer and I can’t link to paragraphs directly so I’ll link to my blog where I quoted it[0]. It’s really a remarkable effect. It takes effort for me to take my eyes back to the scene.
Presumably you could put contradicting information into text like that and reach some people and not others.
It reminds me of that (apocryphal?) story of how an election advertisement company managed to shift votes by telling everyone to not vote knowing that people of one ethnicity would nonetheless listen to their families and vote. You don’t have to lie to people, merely put the text in a place inaccessible to their minds but right in plain sight.
https://wiki.roshangeorge.dev/w/Blog/2026-06-27/Anti-Memetic...
The reason I brought it up is because, people who learn English normaly start with a book. It's heavily polished.
When you speak English as you learned from the books, it does not sound very conversational.
If you are native/fluent English speaker, you can feel the impedance mismatch and feel something's off.
The AI-blindness stems from the fact that those polished edits are so common in publishing field, they all sound the same, and unable to recognize the diffs between AI-generated and human-generated.
There is no real human conversational vibe to them and well. i will stop now.
I find that when I try to speed read modern human writing, there are often errors (like missing or misused words) or awkward expressions that I do have to slow down and think harder a lot to really parse it.
With AI writing, it's sort of self redundant and the information density of each sentence seems to have more even information density. This makes it very easy to do a very high level speed read and get the full gist.
There are also what I'm assuming are bots on hugging face (or maybe non-native english speakers who are using ai for translation) that interact with me where I have no idea what they are saying until I read it very slowly.
As someone who hasn't practiced speed reading, how does that happen? Is it something about the way your brain tries to connect ideas from different parts of the text? Or the redundancy making the signal more stable?
If your reading speed is limited by how quickly you can subvocalize the words to yourself, this is significantly less obvious. Unless the passage is dense enough to require multiple read-throughs at conversational reading pace or vapid enough to be boring, you're going to feel done with the text at roughly the same time. Speed readers do a lot more re-reading and varying of reading speed, and that is going to correlate pretty hard with information density.
pre-read is just looking at how long it is in the headings, and planning out what chapters to focus on if it was a text book. (its sort of iterative, you do a pre-read for the whole book, and then for each section you break it into)
The fast read you try to read only with your eyes, sweeping your eyes across multiple words at the same time, suppressing the urge to say the words to yourself in your head.
iirc the how to read better and faster book even had a cardboard mask you put on the page to practice the sweeping, and some pages that were laid out weird to try to teach you how to do it.
Some ai text just seems really easy to speed read, like if it's tuned for an easy reading level. In PRs some ai seems like it's arguing over weird flex technical details and really starts torturing the language in a way that makes it the opposite of easy to read.
Does speed reading help you process the final message faster if it's written by AI compared to people?
Because if you read 1 information dense sentence, 1 medium dense, and 1 sparse sentece written by a human, it's still way less text in total than 6 information sparse sentences written by AI... even if it's all over the place when it comes to density or style.
---
The density argument is really interesting.
Does speed reading actually help you process the final message faster when it’s AI-generated compared to human-written?
For example, if a human writes 3 sentences—one information-dense, one medium-density, and one sparse—that’s still much less text overall than 6 relatively sparse sentences written by AI.
Even if the AI output varies a lot in information density and writing style, you still have to process all that additional text. So I’m wondering whether speed reading actually offsets the verbosity of AI-generated responses, or whether the total amount of text is still the bigger factor.
While I would feel rusty when handwriting code, working with either Codex or Claude I have lots of practice.
I can probably tell at a glance where output is intermediate while it's still running tools and where the summary starts.
Then I have to recall the context of the conversation, parse the summary, see if there is anything unexpected that popped up. Make a call for whether I need to ask it for clarifications, and what I need to test to verify the change.
Assessing quickly what is pointless yapping and where the information I need is might be a learned skill. If one tries to read and understand every word Claude says, I am not surprised that people are not having a good time.