353 comments

[ 0.20 ms ] story [ 41.0 ms ] thread
As I have been saying for years:

Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer. You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

Depends. If the 700 bits were arrived at by the LLM while spending a lot of tokens, and the result is "good", I may want it through you as a middleman because it used up your tokens and won't eat my subscription usage limit to ask the AI to supply those 700. If you spend the tokens and put the result online, plenty of people can spare their tokens because they don't have to ask the AI to derive it. Bonus if that result was run through some kind of testing and verification.

Obviously this doesn't really apply to super simple questions that the LLM can just spit out the answer to right away.

Bits of information depend on the readers prior knowledge, those bits are not absolute numbers. That makes communication not just information exchange but also syncing to a point of larger shared priors. So some amount of redundant information maybe needed/wanted.
Your mental model is out of date and overly reductive. I can't speak for anyone else, but I am well past the point of giving an LLM a bare outline and letting it fill in the blanks. If it's straightforward information transfer, rather than, say, relational/emotional communication or persuasive argument or entertainment, then I'm not going to spend 2 hours compiling and writing everything by hand because you'd prefer to read "why our recent outage was entirely avoidable" in my own handwriting rather than on a typewriter. To wit:

I have megabits of semantic information in my growing knowledge base that I've given the LLM access to. When I want to send you 10000 bits of information, I send 300 bits of 'this and that and this' to the LLM and ask it to pull together the right 10000 bits of information from this KB to send to you. It would not be useful to send you the original 300 bits I blatted into the LLM without the KB beside it, or without several followup prompts in which I argue with myself about some little detail that you'll probably skim over. And I'm definitely not going to give you access to my entire knowledge base, much of which is personal or irrelevant or possibly even illegal for you to access.

For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today. Sometimes such articles interpret or place context around their hidden references, but a lot of the time they just summarize.

Giving someone the text output of a LLM is very similar to publishing a summary without links to the referenced material. When you were querying your LLM, you could have asked specific questions or asked for a custom focus or point of view. Your intended audience might have questions or different concerns, but they're unable to interact with your LLM. What you have delivered is static and unresponsive. It has all the disadvantages of being machine output without the advantage of being interactive, the way your LLM was for you.

It may have to wait until compute is cheap enough that tokens are essentially free, but we need a system to pass "hyperlinks" to LLM's primed with context, ready to be interactively queried on a chosen context. It's being overly generous to assume that people are putting even 300 bits into a LLM for every 1000 bits of regurgitated writing they try to pass off as their own. When people post LLM output as if it were their own, I have no choice but to assume they had zero knowledge of the subject, but this query taught them what they wanted to learn, and now they're sharing that. That's fine, but please pass an interactive LLM link rather than static text.

Once we have "hyperlinks" for LLM sessions, perhaps we can share LLM output a little more usefully and honestly.

I've seen professional journal pieces refer to science journal articles only to go and read the original article and find that it draws a different conclusion than what is implied by the journalist.
My father used to complain about that 40 years ago, though in his case he was reading newspapers rather than professional journals. But he'd point it out to me often enough that I started to see the pattern. Scientist publishes paper saying "We may have found evidence of X, which suggests the possibility that Y may also be occurring". Journalist: "Scientists find X which proves Y".

This has been happening for decades; I still see it happening today*. My cynical suspicion is that words like "maybe" and "suggests the possibility" don't sell enough papers.

* Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote". Summary said "Exposure to X can, on average, cause a 40% higher chance of Y" (where Y was a negative health outcome). I clicked through to the study and read it. Turned out the confidence interval on that chance of Y was so wide, all you could say with 95% confidence was that exposure to X could do anything from reduce your chance of Y by 5 percent, or increase it by 85 percent, or somewhere in between. They had averaged -5 and +85 to get the scarier-sounding 40% number that they published in the summary, but the truth would have been far closer to "this confidence interval is so wide that we really can't conclude anything from this data". But that wouldn't be nearly as likely to get them grants, so they tortured the data in their summary so that it would look better.

There is also incentives to adapt the message to the outlet. If you send that data to peer review and say we found 40% increase, the reviewers would reject it, so they have to moderate themselves. But if you send a summary to the university’s outreach outlet saying that we found something or nothing we don’t know, then they would also reject putting it out. So even for the authors, the incentive is to send a careful conclusion to peer review and an overblown one to popsci.

Btw, I also think a 95% confidence interval is just the wrong statistic to look at given that data, and that they could probably have analyzed it better.

> so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote"

It's never the case that someone misunderstood what scientist wrote. Much like the scientific papers, news articles, including those reporting specifically on the discovery, have their own goals, and the paper being cited is used as evidence or argument for article's own "study". Except for press, the standard is rhetorical, not scientific, it's the conclusions and not the methods that are "pre-registered" at the start, and claims are defended by "hey it's just a point of view", not by statistical significance.

In your own example of worst offender: the scientific study was trying to establish and quantify the connection between X and Y. The summary article was trying to push the angle that "this institution is doing important work". It started with that conclusion, and the paper cited was just the first thing the author found that could be easily massaged into supporting that conclusions by rhetorical standards.

Same paper might get cited by journalist trying to push for "X is bad for you", and they'll do roughly the same as the summary article. And, same paper may be cited by someone claiming they have a miracle cure for Y, and they'll make a honest observation that "absence of X reducing Y is a common bullshit claim based on misunderstanding the paper [citation], that actually shows there's no correlation there, I mean look at the confidence intervals, even the author says that in text nobody bothers to read"... - citation may be honest, but the article itself is still using it to prop up a different flavor of bullshit.

TL;DR: don't believe news. It's bad for your mental and physical health (p<00.05).

> It's never the case that someone misunderstood what scientist wrote.

This is wrong. Reporters frequently don’t understand the science or the nuance in the science.

Reporting and science are two very different disciplines. Reporters rarely have a deep background in science and almost never have a background in the specific area that they’re reporting on.

Hell, even scientists have trouble accurately describing the work of a different scientific discipline.

Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.

> Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.

I'm not inventing them, but maybe conflating two sources:

1. Malice directly intending to hurt or defraud people. Probably not as common as how I make it seem.

2. Not caring. Well, I subscribe to the view that not expending effort to be accurate when talking to other person is as bad as slashing their tires (paraphrasing an old quip), so I very much consider bullshitting and picking a conclusion and then massaging facts to fit it, to be acting in bad faith too.

Nah, even in middle and high school it was plain that people sometimes misunderstood what the teachers meant. It's not like that goes away in adulthood; even when people "care", they still often misunderstand nuance or details, and sometimes even the bigger picture.

Honestly, it's plain weird to say that people never just make mistakes.

PS - worth adding that "I misunderstood" and "I didn't care enough" are not mutually exclusive. You can do both, so saying "they didn't misunderstand, they just didn't care" isn't a reasonable rebuttal. But even setting that aside, there'll be plenty of folks who care but still don't understand.

I'm gonna invoke Hanlon's handgun here: not attributing stupidity to what's adequately explained by systemic incentives promoting malice.

I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.

> I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.

The pattern of personal motives fits misunderstanding. You need to show that there is an organizational pattern of "malice" (your word, not mine), rather than an organizational pattern of "we are trying to publish quickly, and quality accidentally falls to the wayside". I.e., negligence, not malice.

You haven't provided even a shred of evidence suggesting there's malice at the journalist level. Every science journalist I have met genuinely cared about the science (which is why they were writing on it), but they didn't have time to learn enough about the subjects to understand they were oversimplifying things.

Not in science journalism, but I've personally encountered a case in normal journalism that I can only attribute to malice. It was many years ago, but it was so blatant I still remember it.

The 911 call went like this, according to its transcript. Caller: "This guy looks suspicious, like he's on drugs or something. It's raining and he's walking around looking into windows." 911 operator: "Can you describe him? What race is he?" Caller: "He looks black."

How the TV news reported it on the air was: Caller: "This guy looks suspicious ... He looks black."

Omitting excess verbiage is one thing. Omitting words that entirely change the context of the statement, making it look like the caller was racially prejudiced rather than responding to a specific question, is something else entirely. That was the last time I trusted reporting from that particular source (it was NBC, by the way).

My principle is that when someone lies to me, I stop trusting them. By lying I mean not just omitting details, or having an obvious bias, but deliberately telling me A when they clearly know that the truth is not-A. I could not see that report any other way but a deliberate lie, knowing the truth and attempting to make people believe the opposite.

Scientist: "My discoveries are useless when taken out of context"

Media: "Scientists claim their discoveries are useless"

In most cases, they are only useful to other scientists in their field, which does mean that they are useless to the general person until they show up in the form of a new commodity.
> Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote".

It's a similar thing. We live in the "attention economy", and research institutions - particularly after the US President openly went and had his minions cut funding to research purely on ideological reasons, but it's been a problem for decades - are just as susceptible to blow stuff out of proportion to make headlines and thus increase the chance someone might throw some money over the fence.

And media does the same, just to manufacture artificial debate. And so do politicians.

And frankly, I'm fed up with that, we will drive ourselves into a wall.

I have recently reviewed a paper that referenced my own article… but the conclusion was so off, I actually went back and re-read that entire article just to be sure there was no hint to the conclusion that the author derived. There was none. Uncanny experience.
> For a long time...a lot of online news stories would reference websites, papers, polls, etc. without linking to them.

The "$CITY_NAME Business Journal" websites are the absolute worst with this. They'll refer to something specific, for example "$BIGCO's 2025 10-K filing" and it will be a link. That link will go to the 10-K, right? Nope! It goes to another page at the same business journal. Maybe that page is a summary of the 10-K, but probably not. Maybe it's just the general index page for all the articles about $BIGCO at that journal. What it links to, it definitely won't be the specific thing described by the text of that link.

(comment deleted)
Link rot has accelerated in the past decade. Links are great, when they work. If they are a few years old, they often don't. You certainly can't count on it. If you're referencing a paper or an article, it's probably better to cite the title of the article, the name of the journal or magazine or newspaper it was published in, the author, and the date. Then someone might have a chance of finding it again.
Doi url
That works for pieces that have one, although places that index those aren't always free.
Wikipedia has solved this problem by using link + date for references and providingan archive.org link when the original is no longer available.
> For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today.

I agree, this drives me crazy. Ironically, one of my favorite uses for Claude is to ask, "What study is this news article talking about?"

It's pretty good at digging up the source and related sources. And most of the time, if you read the source, the article is nonsense and gets everything wrong.

Not just news stories. This is a huge problem with social media and forums in general too. Lots of people making bold claims about random things, lots of stories they say are based on a third party source, but very few actual links to said sources in question.

I still remember a recent example where one of those trivia accounts on Twitter posted an interesting story about some guy whose life completely changed after an accident, but neither linked to a source or named the person in question.

The only way I was able to verify it was true was through someone in the comments asking the platform's AI chatbot, and the chatbot providing context that I could research and verify...

> There are still news sources doing this today.

It's the opposite, all big news websites do this. Fairly sure it's part of the policy.

I would say only small, niche websites link to sources.

I love this example, partially because it jives with my conviction that LLMs are the ultimate translation machine. Ever since the embedding model days, it is clear that these models are amazing at representing meaning as math. The fact that LLM's most salient use is for coding somewhat agrees with that. After all, what is a programming language but another language? We instruct people with words and machines with code.
I will give you a use case where this is absolutely not the case.

I have a bunch of CLI utils I run for various clients and their peculiar setups. They now have man pages with descriptions and examples in them because the LLM went and read my code and did the needful.

I no longer have to re read my own code, rather I can just use the manual page.

Format and description came from semantics and context that (barely) existed elsewhere and I was not going to retain or transmit, but I have now.

I never liked information theory because information theory as Shannon envisioned it fundamentally did not deal with semantics.

AIT tried solving it? But AFAIK it's a lot of pretty results with not much real application.

A better approximation is something of a "shared model"; then you can actually state things like, the transfer of information sometimes is "trivial" because, well, it's right there in your compressor/decompressor.

My understanding is that to unambiguously quantify information, you need to have a known model within which that information fits. In this context, talking about LLMs adding value (or not) via inserting new semantic information, we're assuming that public knowledge on the internet is not new semantic information, and the quantity of information of contained in text talking about public knowledge is equal to the ~32 bits needed to point to that knowledge.
I was just trying to understand this about a month ago and it is interesting how little there is in terms of semantic information vs Shannon.

An Outline of a Theory of Semantic Information by Carnap was the early attempt.

Fred Dretske wrote Knowledge and the Flow of Information in 1981.

Luciano Floridi has a few recent books.

I couldn't find much else. I don't think AIC really solves the problem of meaning either.

I think the Dretske book was the first time I really understood where Shannon was coming from but I gave up when it got to his actual semantic ideas.

I think I ran across a recent paper that motivated trying to back track what work had been done in this area but I don't recall the name of the paper.

I've have shelved all this for now as over my head.

But then the LLMs aren't adding new information, they're just reading your code and translating that information from e.g. python to english. Rather than sending someone an LLM-generated doc to someone, send the code and your own personal thoughts on the code. If your reader doesn't want to read and understand your code, they can ask an LLM to analyze it, within the context of their specific use case and your personal thoughts if any.
> e.g. python to English.

You're making a big assumption that the code is what is being executed, and not a compiled binary.

Where is the code: My repo? the clients? If it's in mine, the client does not have access and the CLI is a first stop to debugging. They arent in the context of written docs, more likely a production error from a log (thats now spitting out a message to check the CLI).

Less steps, less tools, more context in line and available in an interface your already using.

> they can ask an LLM to analyze it, within the context of their specific use case and your personal thoughts if any.

Or I can skim the man page it generated and make sure it looks good. The "work" (the tokens) dont have get spent over and over again.

I was going to write a post disagreeing with this on the basis of the fact that the reader lacks the background information the LLM has. For example, if I were to prompt "explain the proof of quadratic reciprocity using Gauss sums" most readers would need the entire LLM's answer (and much more, probably) and not just the prompt.

But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text. Most people don't do this as it's extra effort, but it's interesting to imagine a world where this is the default way of engagement with a text, assumed by both writers and readers alike.

> But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text.

That's basically what I've been asking my colleagues (so far a losing battle): Please don't send me AI-generated text. Send me your prompt instead. It is highly likely that I will understand it without needing an LLM, and if not, I can do it myself.

I don’t quite think this tracks. Perhaps you want to communicate 1000 bits that are well known and can be referenced with a 300 bit key. Then the LLM can easily retrieve the remaining information. It’s like sending someone a link to the Wikipedia page instead of explaining something yourself.

No, I don’t want to read LLM writing because it is BAD at it. It doesn’t really understand how humans think (because it thinks differently), and doesn’t seem to understand core principles very well (presumably due to the lack of world model), so it can’t write something humans enjoy yet.

If you have 1000 bits of semantic information that you want to transfer but your default communication combines it with 10,000 bits of noise. Giving it all to an LLM and iterate on reducing that noise while making sure the 1000 bits is still present would enable you to communicate more effectively.

Overall, ideas are ideas. I'm not overly concerned with the fact that it was you who had the idea, as long as the idea is interesting. I don't know most of the people who write the things I read, so it seems to be of no consequence to me at all if they wrote it, as long as it is interesting. LLMs are notorious at creating things that are bland and vacuous, but they by no means have a monopoly on it.

Be the source human, machine, or dolphin, if they write a good article, I'm prepared to read it.

LLMs are wonderful at adding noise and okay at removing noise. My point is that they're not very useful at adding signal. If you wrote 10000 bits of noise and 1000 bits of signal, I would rather receive those 11000 bits from you, and if necessary ask an LLM to remove noise based on what I consider noise. If you can point out to the LLM what it should consider signal vs noise regardless of context, it should have been easy to not write that noise in the first place.
I like this example and it made me think, there is an analogy here to spec-driven development and vibe coding.

Rather than send 300+700 bits, like you said, send 300 (or less!) and let the human intelligence on the other side generate the result. Which supports the even older perspective: “If I had more time, I would have written a shorter letter.”

I’m not sure if this lands on anything very profound, but what about a pattern where, instead of codifying agent output at all, the only artifacts we share are the prompts. And the rewards (respect) accrue to those who generate the most generative among people and AI

The value of writing isn’t always to communicate new information. It’s often to align everyone’s assumptions. For instance, when I say casually to a colleague or an agent “this change will require a db migration” they understand it’s to my teams primary application database. If I submit a design doc to a company wide review which database is changing is critical information. If the 300 bits were truly enough your agent or junior engineer would implement the wrong thing correctly as they often do.
Clarifying which database you're referring to is exactly the sort of thing I'm talking about when I say "new semantic information". It should be communicated, and you can't trust an LLM to choose the correct database, you need to specify that yourself.

What I'm talking about is if you add "database foo" to your prompt, the LLM may then add text describing what that database is, where it is, etc. But that's not new information, it (hopefully) already exists in your team's public docs, slack convos, etc. You should just say "database foo" directly to your reader, and if they want to learn more about that database, they can do that themselves, or you can give pointers to them based on what you consider important.

> say "new semantic information". It should be communicated, and you can't trust an LLM to choose the correct database, you need to specify that yourself.

But I can most of the time, and correct it if it chooses the wrong thing. That's the whole reason LLMs are faster. Its the reason we can give a paragraph prompt and get a kLOC PR back but only need to correct about 5% of it.

What already happens: - People give an LLM a bulleted list of points that they want expanded into a professional sounding document. - The receiver doesn't wanna read all that. They put the full document into an LLM and ask it to summarize it into succinct bullet points.
I first noticed this a year ago when I read a gmail AI summary, then glanced at the main text and saw it was AI generated. I feel like there's good fodder in there for a dystopian sci-fi story about a future where nobody communicates directly with one another, it's all AIs translating, but the AIs slowly start to drift.
Agreed re: the sci-fi story / trope.

I feel like a lot of this is a problem when someone technical is attempting to communicate a complicated technical subject to a less-technical audience.

I can only dumb a thing down so much before the description is useless (when you zoom out too much you lose the details). Even technical people who could understand it but are lazy / "in a hurry" use the summary, without thinking about what detail they are losing.

Even more infuriating is when they then reply to my email, having only read the AI summary, and ask a question that was already answered by my message.

This is the exact same thing that happened pre-AI, with the added step of wasting energy/resources on the AI summary in the middle.

>but the AIs slowly start to drift.

It's essentially the tower of babel. Each person will devolve to speak their own internal language only they understand. Each language will need to be encoded down to its meaning to be reinterpreted. None of us will know if the transformers are accurately decoding, or if the other person is accurately interpreting the decoding (which is arguably already a feature of human language without the computers in-between.)

It was in a gmail promotional video years ago already. Person A would use AI to turn their summary into email and person B would ask the AI to go back to the summary.

They knew it was going to be like that from the beginning.

I’ve been saying the same thing. I’m not saying all, but a significant amount of comms could be bullet points to the benefit of both sender and receiver.
Axios built a hefty business off this simple idea
This is exactly what I want: To communicate with me, have your LLM expand your message to include relevant parts of _your_ context, then I will have my LLM summarize it as briefly as possible with respect to _my_ context.
Just ask them for their prompt, and then put that in your own LLM with your own context. Why do you need their LLM's fluff?
Saying something for years doesn't make it right. Has nobody ever challenged you on that? Our brains prefer to read enjoyable bits, not raw information. If you only got those 300 bits, you aren't going to bother putting them into your LLM to generate 700 bits to make it a more enjoyable read, you're just going to struggle through the 300. Maybe when browsers come with built-in automatic text puff-uppers, then you'll have a case but almost nobody does that now.

Here's a clearer example - would you rather learn a concept from a research paper or a textbook or blog? You say the research paper but they're dense and hard to wade through where-as blogs and textbooks are more wordy but hold your hand, which is something that helps humans learn.

It isn't the transfer of information at all. What's actually happening is you're prompting experiences in a human instead of an AI using text.

Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry, and subtext.

All of that is learned, and writers usually assume they can rely on that learning as the context for the text.

So you don't write to 'transfer information' like a network cable, you write to trigger experiences in the human version of latent space.

Factual information is one kind of experience. But even when that's the goal, there are always layers of implied relationship, social register, role, status, and other implications in everything that's written.

In normal communications the context - business emails, personal messages, mainstream journalism, fiction, and the rest - defines what acceptable language looks like.

The content fits inside that. But it has to fit the context, otherwise it lands in a semantic and psychological uncanny valley - like sending LinkedIn speak to a spouse on a wedding anniversary.

The real problem with LLM writing is that it's good at the technical layer - the grammar and spelling - and has some insights into the rest.

But the default content style is marketing and ad speak. And recently it's developed a weird and unique hybrid style which applies marketing fluff and pretension to technical content like code comments.

So you get one register instead of all of them. It can attempt others, but it's still too limited to generate them fluently. Sometimes the results are outstanding, but often it defaults to mechanical clichés.

So that's why it sucks and sounds so hollow.

Can it be fixed? Yes, but it's very hard work, most people don't have the skills, and it takes time - often too much time to be worth the effort.

That "marketing style" isn't a style, it's the lack of content itself. You just restated their original point despite trying to object to it, because it was correct and there is no way around that.
I don't agree.

Sometimes Claude's problem, such as when I ask it to summarize a long, complex session back to me, is it's too information dense. It uses weird invented terms to gloss over complex parts of the architecture instead of explaining them.

But no matter what - too dense or too sparse - it always sounds like Claude.

It misses the forest for the trees. It feels the need to highlight details not understanding what details are most relevant to a human reader and how to survey the larger problems in a cohesive way that emphasizes the right parts without cliche and undue emphasis.
It's never too dense. There are only more words but not more information.

If I describe all of the individual muscle contractions and joint motions required to walk across a room, it's hundreds of pages of data to say almost nothing. That is the exact opposite of dense.

By contrast a poet can deliver many concepts and many layers and even practically a fractal choose-your-own-adventure in only a dozen words. That is what dense is.

While I agree with some of what you've said and the conclusion you've arrived at, in the end, I think you've missed part of the picture here with regard to "prompting experiences".

> Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry

These are methods of encoding, there's no reason all of these can't be represented in an LLM from a technical point of view.

> and subtext.

This is the other half of the equation to me. Humans communicate by relating shared experiences, an LLM cannot have shared experiences. While it might be able to encode subtext that has been specifically called out and explained, it will never be able to encode the breadth of human subtext, especially that which is reliant on emotion.

I don't believe it is possible to change this until the point mankind truly develops a "wetware interface" to the digital world (and I personally don't want such a thing to exist).

It's worse than that. If it takes others longer for others to consume and understand what you're producing than it does for you to produce it, you'll never be able to communicate with someone efficiently. Communication breaks down at a fundamental level if if you can't keep up with the other side if outputting and they won't slow to allow you do do so.

I see this all the time now with LLM generated output. It's easy to have an LLM generate a chunk of content that can be dropped into a chat or comment, and when it took you 20 seconds to have something written up based on the shared understanding you and an LLM have about the context of the situation, but it takes other people 3-5 minutes to read and understand that content, that fundamentally doesn't scale. It's bad enough when one or two people are doing it, but if the whole team is doing it, the only way to keep up with the stream of information is to also consume it through an LLM. At that point you're likely to be missing much of the nuance, and the amount of errors will explode.

This can be alleviated by people reviewing the output of an LLM and making sure it both includes fundamental information that might be assumed by context and reducing it to the parts that are essential for the new context it's in. This takes time, but is extremely important.

Having an LLM write gobs of text to send to other people instead of doing it yourself is the equivalent of a low yield cognitive zip-bomb. Don't do it.

If only it were 300:700... it is more like 300:100000 and up. But then you use AI to summarize to hopefully retrieve the 300 bits. We should try doing this in a loop with a simple enough sentence to start with and see what pops out after every iteration. Fully automatic Chinese whispers.
"Chinese whispers" - a Chinese colleague teased me about using that phrase in a meeting, which made me question whether I needed to use that phrase in other contexts (e.g. conversation with family, in my home town, or with close friends).

I am not ready to use the Americanism, 'teh telephone game', since nobody knows what that means, yet everyone over a certain age knows what 'Chinese whispers' means.

I’m not disagreeing with your preference for human authored writing, but I don’t know that I agree with the content of your argument.

If you consider that what humans are doing during conversation is a form of compressed encoding / decoding from some latent representation through a quantized signal then if you interpret it that as a compressed sensing problem you absolutely can infer to a very close approximation the original latent representation using far fewer than those 1k bits.

From an information-theoretic perspective, if I can effectively transfer a latent concept to you in 300 bits of quantized communication, then by definition the latent concept itself isn't more than 300 bits. Any bit over 300 in the communication I send to you is fluff.
While I generally agree, the difference is I can also hand an LLM my pile of code and documents which is... a lot more information dense than basically anything I could write. Sure, a one-off prompt in a chat window isn't very helpful.
Summarization—especially of private context—is definitely one of the main exceptions to LLM writing being useless. However, it's still preferable to turn your private context into shared context and then give pointers to that, rather than having an LLM attempt to summarize it.
Oh fully agree. I hate reading LLM generated text, just wanted to provide a counterpoint to GP.
That would be true if there is only one receiver and if i could exactly predict how the receiver would interpret these 300 bits. But i could use LLM to expand these 300 bits to 1000 bits of more explicit information that gives enough redundancy and context that minimizes risk of misunderstanding and increase easiness of understanding, then validate such expanded message (and possibly do iterative corrections if LLM did not understand those 300 bits correctly), and finally release expanded message to (potentially large and heterogenous) group of receivers.
If there's a significant risk of misinterpretation when writing only 300 bits, then the idea you're trying to convey is worth more than 300 bits. In the process of prompting the LLM to add information, you're giving information to the LLM that you could instead just give to your readers directly. If you're just using an LLM to expand your thoughts in the hope that you get a better piece of writing, maybe you should put more effort into your own writing.

(This is all under an information model that assumes the LLM and your readers have equal access to knowledge, which I probably should have made more explicit in my original comment.)

As with everything, it depends how you use it.

I’ll often put a long stream of consciousness on the page, or jot down rough meeting minutes, then ask ChatGPT to “summarise this for an email”. The result is shorter, clearer and easier to read.

AI amplifies the habits of the person using it. If they’re lazy or dim, then it's like giving a monkey a gun.

So... We should only communicate in compressed formats?
Yes, but specifically compress the concepts, not necessarily the text encoding. "100 commas" and ",,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,," may be different amounts of ascii data, but they both represent the same concept that's worth around 20-30 bits of semantic information.
I agree. But there's still a valid use here for LLMs: If I have 1 Million bits of information and I want to communicate a synthesis of 500 bits to you, LLMs can be a viable helper. It is rarely done that way - I agree, but the way I use LLMs: Drop in all the relevant context information (PDFs, HTML, Markdown, Pictures etc. - up to usually 200-300k tokens), then compile a synthesis/summary prompt (usually 1-2 A4 pages of handwritten text. Then copy the output (1-2 A4 Pages), manually edit and send off (often, including an archive of the full original conversation, so the human on the other side can consult the unfiltered original prompt + inferencing, if wanted).

In other words: If I didn't reduce the 500 pages of text for you, you wouldn't know what I mean or what is relevant, or how to filter it yourself.

"Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700,"

You absolutely can if that information is in the code, which it often is.

There should not be that much in the code that needs further elucidation.

Some stuff definitely - but not much.

Usually you need the code and architectural summary + that stuff.

The AI is not very good at it but it will get better.

I think the debate here is about a few different things.

I notice in the attitudes of students towards reading. They will take a 10-page journal article and ask the AI for a summary and then read a summary that's the equivalent of maybe half a page. But if the article could have been half a page, why is it actually 10 pages? It's true that there is some boilerplate, but it's strange to me that people could think that 90% of what they're reading is (to use your phrase) "not true semantic information". It's like if you went to a restaurant and ordered a 10-oz steak and they brought you a little teeny bite of steak and said "Oh, other places will give you a bigger one, but most of that is just filler, we just took out all the superfluous parts." It's a worrying sign for our future if things like this are not tripping people's skeptic sensors and making them wonder if they might possibly be missing something.
A summary is about utility. They are definitely missing something but that doesn't mean what they are missing is useful to them at that time.
I mean I think that points to a related issue of people only focusing on a short-term notion of utility. The point of being a student is largely to learn things that may potentially be of utility to you in some way later, not just to do what meet your immediate needs (in the sense of passing the class).
I get what you're saying and I myself am guilty of surface level learning. But the idea that students should consume all the information is impractical. Learning what is acceptable to discard is part of being a student. I imagine very few students read every college text book cover to cover.

Beyond that people have different motivations and goals and only a limited time to achieve them. Basically I wouldn't be so quick to judge. Plenty of students have dropped out and gone on to do impressive things and that's a bit beyond reading only the abstract for a few assignments.

you think like such a techie. consciousness is not analogous to IO. if writing were only the transference of something from point A to B then what is the technical function of poetry, a question (the open ended kind), a pondering, a wondering, and so on? Furthermore, language is lossy and introduces a large degree of subjectivity, mystery, and uncertainty nomatter which words you choose. Words are by nature lower res/on a lower ontological domain than thought. So your premise is preposterous on both a practical and theoretical level.

Furthermore not all writing is for another to consume; nor even for the author themselves to consume. That is to say it has meaning ipso facto, not dependent on transference, as ritual.

Art is an abstract method of communication. When you choose the words of a poem, you're (hopefully) doing it to convey a feeling within you to the reader. If you write a poem about a beautiful spring day, it's probably because you experienced one, or you're remembering one, or someone was telling you about one, and that evokes feelings within you that you want to put into words, right? Surely you wouldn't write a poem about something you don't care about in any way?

When you talk about something you're wondering about, you're saying that you're missing information. Your ponderings are dancing around the void in your knowledge, defining its boundaries, and maybe imagining what answers might be able to fill that void.

When you put your thoughts into words, they're insufficient. You have so many ideas swirling around in your head, and you can never put them all on a page in the fidelity at which they exist internally. But words are the best we have. Whatever words you write are your best attempt to convey your thoughts to me (barring other media). You're distilling your inner voice that speaks a language only you can understand, into an outer voice that others can understand.

I don't think I'm exactly refuting you here. I think what you've written makes sense, and caused me to think about many things, more so than any other reply to me today. But I also don't think your comment is refuting the point I was trying to make, mainly that LLMs rarely add value in human-to-human communication.

I could probably have pasted my comment and yours into an LLM, and it would have come up with a clearer thread connecting my words to yours. But that thread probably wouldn't have been any of the ones either of us saw, would it?

Thanks for adding a new perspective to the conversation :)

To zero in on your last question, I do think a very good LLM can surface the unconscious (even completely unstated or hitherto nonexistent) premises linking two subjective thoughts and make them explicitly known. And so what if they were not something either of us had in mind?

Surely the ability to do that is worth taking note of.

I am not advocating for letting LLM's write for you, to be clear. Sentiment wise, I largely agree with you. Just not with your total writing off of the possibility that it could serve.

It's easy to imagine an LLM aiding the communication between a mentally disabled person and their parent/caretaker.

Or, perhaps, some day, between animal and man. Who cares if the mediating component "hallucinates" some particulars of expression if it achieves the goals both want, which were previously impossible?

Communication between people is much more complicated than that.

I get vague statements thrown at me with people expecting me to understand it.

Same with writing, setting up whole context to properly transfer 300 bits is always orders of magnitude bigger then just additional 700 bits.

If you read anything you're reading superfluous bits.
Saying the hard part out loud: we give more value to the people with 1000 bit brains than we do to the 300 bit people. There’s a strong incentive to write like you’re the former and not the latter. Freely available LLMs give people the means to match the incentive.
That's assumes a couple of things that are trivially not true:

- The assumption that both parties know about the same as an LLM does. An LLM know orders of magnitude more.

- The assumption that the output of the LLM is not refined over a few cycles.

The point is that you might give 300 bits of semantic information to an LLM, it fills it to a 1000 with perhaps 400 wrong bits. You correct it half a dozen times. It's now 950. You do the final touch ups. It's now at 1000. And it still took you 20% of the time to do it.

- An LLM has more knowledge, but it doesn't have information about what you specifically want to convey. Anything you give to the LLM may be good information. Anything the LLM adds is not additional real information, because whatever it adds can be inferred based on whatever you wrote. Or the LLM adds additional information that can't be inferred based on what you wrote, which is even worse because that's basically just misinformation.

- If you're giving additional prompts to the LLM to refine its output, then you're the one adding real information, not the LLM. The LLM is just rephrasing the information and adding noise.

- That's outside the scope of what I claim. The claim was that if you give an LLM 300-bits of information it cannot add more things that are relevant that the reader could not also do themselves. It certainly can. The LLM may add information you do not want to convey. Which is easily mediated by removing it or additional prompting.

- You are adding real information. The LLM is also adding real information. That's the entire point. It happens very often that an LLM suggest something to me that I did not know or simply did not think about. An LLM solved Navier-Stokes recently. That was most certainly not just adding noise. That's real information purely generated by an LLM. Information that was worth a million dollar price. Information that man centuries of mathematicians were not able to do generate.

Keep in mind that I'm using "information" in the information-theoretic sense. One could argue that because that particular solution to NS is provable, it is therefore implied by the propositions they started with and adds no new information.

And I'd still rather read a human's interpretation of the solution to NS than read whatever the LLM wrote.

> You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

Exactly this. Just send me the prompt! ;)

I’m not sure the 300-bit → 1,000-bit framing applies in all instances. The 300 bits may be a compressed cue to a much fuller idea. The AI can combine that cue with its prior knowledge to help reconstruct what the prompter was trying to express, with the prompter then verifying whether it’s right. Without the relevant prior knowledge for reconstruction, or the prompter for verification, it becomes much harder to know whether you’ve reconstructed the intended idea.
Unless that 700 bit was transferred on a separate occasion the inferred 700 bits is not true information, anyone could have reconstructed it from the 300 bits.
No, the 700 bits come from the sender verifying and vouching for the information before sending.
Unless that takes 700 attempts on average I don't think that actually works.
That assumes each bit is a coinflip, doesn't it?

Even Markov chain autocorrect tools do better than 50% odds*, and even GPT-2 was significantly better than that kind of autocorrect.

* at the word level; IDK how redundant/efficient language is when it comes to bits-worth-of-fact-claims-per-word. But "your cat is sitting on my" -> [mat, laundry, roof, head, belly, laptop, microwave, …] clearly has many bits of information, and a Markov chain will encode the most likely next word even if the user doesn't know what the most likely next word is. Verifying where the cat is sitting is also very easy, as is correction.

I assumed a coin flip, indeed in practice you would likely achieve far less than 1 bit per attempt.
Far more than 1 bit for most attempts: they need that just to be able to write coherent sentences, and a lot more to be coherent sentences on the right topic.

Some specific conclusions would be far less than 1 bit.

The average will depend on both the question and the AI.

This is about information from the sender to the recipient. It is fundamentally impossible to transfer more than one bit in one binary decision.

A LLM adds noise, not information. At least in this framing.

Frankly, it's a stupid framing.

LLM is not a random symbol generator (hint: training data is not random), and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check (at the very least so that blatantly stupid hallucinations don't paint the sender as inconsiderate or incompetent).

That check alone can add bits to the final signal.

> and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check

I'm reminded of an old quote:

  The reasonable man adapts himself to the world: the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man.
I'm not sure we're working in the same framing here. For one LLMs are random, sure you can fix the seed but you don't have to, you can even replace the RNG with true random noise.

So there are 2^300 possible ideas, only 2 outcomes from the cursory check, how do you get 2^700 outcomes? Most of those are just random variations the LLM added which is not a transfer of information. You would be lucky to even identify which of the 2^300 ideas was being conferred.

"Random" is too loose a word, but they were responding in a context where it meant coin flips.

LLMs are not even odds on all possible outputs, they are biased towards patterns which are upvoted by the training mechanism (at a minimum: the source material, RLHF, and synthetic data).

The information any trained model transfers to output, is information it gained during its training.

No single human is capable of having consumed all that training data.

There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?

> There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

Indeed.

The best case is a P vs NP situation: can the claims from the AI be easily verified, or not?

This does not excuse people too lazy (or overly impressed*) who fail to attempt the verification.

> It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?

Mm.

Thanks to a philosophy course I did half a lifetime ago, I think there's a fundamental problem defining "information" in this context. It feels like it should mean "knowledge" because the discussions about Shannon entropy and transmission channels assumes there is an actual source-of-truth, but my conclusion from discussions about why "knowledge" can't just mean a "justified true belief" means I don't believe we can do better than "belief"; an LLM can generate tokens that change your beliefs, but ultimately neither you nor I nor some annoying colleage who has made themselves redundant to the LLM, can be an oracle with definitely-true knowledge.

(I have of course tried asking an LLM about this thread; I don't feel it illuminated anything new for me).

* In the early days of LLMs, I was overly-impressed. Then I realised we were doing the same thing with LLMs today that we did with 3D graphics in the 90s, where every new engine was hailed as "photorealistic" only to be dismissed 6 months later when something better came along: https://archive.org/details/nextgen-issue-26

Only now it's every 11 weeks rather than 6 months.

In information theory, each bit is a coin flip by definition
I realise I phrased this poorly.

The first sentence in this comment contains 32 characters; from the point of view of a naïve channel with no compression, that's 256 bits (given none require breaking out of UTF-8).

It did not take 2^256 attempts to construct the first sentence in this comment, because the generation process was not flipping coins per bit.

LLMs also do not emit bits chosen with a [0: 0.5, 1: 0.5] probability distribution.

From a compression point of view, the surprise can be reduced such that likely messages use fewer bits. However, this requires the receiver to agree with the sender what the probability distribution over tokens is.

Intelligence is a compression algorithm. If I can predict your next token, and we both know this, we can agree in advance that you don't need to actually send it.

No single human brain is able to predict the output of an LLM anything like well enough to do that.

Not anyone, no. From an information theory standpoint, that it's possible at all to complete these 700 bits, only implies that anyone logically omniscient could. It's entirely possible that an LLM is capable enough to infer these 700 bits, and the human reader isn't.
Right but where is this information coming from, it can't come from the sender because an idea that can be conveyed in 300 bits cannot contain 1000 bits of information. You're just using the LLM to translate the idea into something more legible.

Alternatively those 700 bits are information the LLM added, but where is that information coming from? Is it noise? Random facts? Random lies? And who is the receiver even talking with if most of what they read is something the sender didn't know?

This is true - but also, models are absolutely terrible at writing articles and I don’t want to read them.

The issue isn’t that a 300 bit idea is padded with 15 KB of content. You can take any human-written article and reduce it by 90% with next to no information loss. What you lose is what makes the article a compelling read instead of a fact table.

I think the reality is that we will see quality long form AI-written content at some point. It doesn’t even feel like labs are particularly interested in chasing that now; code sells way more tokens. Right now the trend is that models degrade in writing quality as long as that pulls them up on coding benchmarks.

> You can take any human-written article and reduce it by 90% with next to no information loss.

You've unintentionally circled the error here. The "purpose" of an article extends beyond "convey this essential information".

By analogy, a textbook contains far more words than a spec sheet, but attempts to train the human to be able to easily interpret spec sheets. The so-called "information" content of both might be equivalent, yet one does a better job of teaching students.

Yup, it helps when the prompter reads and edits the AI-generated text, but usually they just skim it and send it unedited. Worse, it's usually not 300 bits -> 1000 bits (= a concise message covering all the relevant topics), it's more like 300 bits -> 3000 bits (= typical LLM verbal diarrhea hiding the relevant points in a wall of text).
This is a great analogy, thanks.

That means folks using LLMs starting off with 300 bits KNOW that they lack the full payload of information to transfer to you. IOW, they know they need to transfer much more than 300, so they use LLM to fill those gaps. That's the crux of the slop universe out there. Folks are using LLMs for the 700 bits on top of their 300 bits and passing off the full 1000 bits as their own.

I echo the writer's sentiment. "I don't want to read the clanker's 700 bits. I want only your synthesis." (I can get the clanker to generate those 700 myself. Unless ... unless this whole LLM slop market is all about saving you the time to get an LLM to generate those 700!)

Just use LLM to filter out those extra 700 bits, and you'll happily end receiving 100 bits or original information and 200 bits of slop.
> Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

Sure you can. LLM doesn't know what those 700 bits are, but you do. You may not realize it, and may not even know it at the time of prompting, but you do by the time you're sending.

Typical case is like this: you have 500 bits of semantic information to transfer. You give 300 of them to LLM, and get back the 500 bits you knew you have, and extra 500 you can quickly confirm are correct and relevant. Some of them are just dereferences of your input - where you recalled a pointer, but not what it pointed to. Some of it is information you never had before, but are able to easily validate.

You send that to me. I likely immediately realize the message was AI-assisted, but I trust you to be a decent human being, and not an asshole that lobs unverified LLM vomit over the fence for others to deal with. End result: you communicate 1000 bits of information to me, instead of planned 500, and you yourself learn extra 500 bits.

This is the optimistic scenario, but it does happen when LLM operator is not an asshole.

(Excuse the strong language, but I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself, so it's a topic close to my heart.)

> Some of it is information you never had before, but are able to easily validate.

This is exactly the use case an LLM might (huge emphasis on might, depends on workflow, agentic vs. relying on contextual which can hallucinate) be good at and yet humans are notoriously bad at, because we are swayed by emotional responses and it is easy to have an emotional response to text that is programmed to look good for you and you alone.

> you communicate 1000 bits of information to me, instead of planned 500, and you yourself learn extra 500 bits.

Extremely optimistic. If this were the ideal scenario, you would USE the LLM to garner information ABOUT those 500 bits and then reframe them in a way that you yourself would put it. If there is insight, your "word" in your mental register now expands from the original 1000 bits to 1500 or 2000, and then are "processed" by your human brain that includes subconscious choices that are meaningful to the end result. There are tons of hidden semiotic data in your diction and wording (think resource forks in classic MacOS/HFS, only visible to the filesys) that is lost when you rely on another source to put together words for you; it's as if it is a game of Telephone. These are subtleties which you may intend for your recipient to receive and which are crucially important to your recipient and are irretrievable, it is intrinsically lossy. You have an alphabet soup of words, they cannot be put together by an LLM in exactly the way your brain did. We must rely on the fact that we ourselves put this together, the "aha" moment when an LLM does it for you is illusory and does not itself provide meaningfully important confirmation that you indeed say what you mean to say. Of course, humans say things and put things in way we do not intend to all the time. I still fundamentally believe this is more honest than relying on a third party that is not capable of understanding human emotional nuance to put together language for you, when language is and always has been a manner in which to dictate human emotional nuance.

> (Excuse the strong language, but I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself, so it's a topic close to my heart.)

Does not negate the fact that LLM output itself, at least when used to convey human emotions or thoughts, is lossy. A very highly compressed JPEG with added interpolation from an upscale algorithm might come up with cool details that were never present in the original, and may look cool to both you and the recipient but are not honest to the source material. You receiving 240p JPEGs on a day-to-day basis is irrelevant to this. For the purposes of communication, it is a massive error which has the potential to compound, regardless of whether or not you or the recipient believe this to be the case.

> If this were the ideal scenario, you would USE the LLM to garner information ABOUT those 500 bits and then reframe them in a way that you yourself would put it.

Yes, but at that point in practice we're getting into over-optimizing territory. In this optimistic case I presented, you could learn those 500 bits yourself and formulate a clean message yourself, with all 1000 bits in it, but since LLM already gave you the text, and you feel you vouch for, you may as well send it over and save yourself the effort.

In reality the numbers are probably lower, and writing the message yourself is IMO also a good way to be truly sure you vouch for the "extra" 500 bits, as it forces you to actually pay attention. There's a chance you'll find inconsistency in output, or in your own understanding. I don't begrudge people for eventually cutting the process off here, for practical reasons - it's the fuzzy line between accuracy and perfectionism.

> A very highly compressed JPEG with added interpolation from an upscale algorithm might come up with cool details that were never present in the original, and may look cool to both you and the recipient but are not honest to the source material.

Again, I think it's a wrong take. LLMs aren't pulling the information out of their asses, and you are also not able to express every information directly. LLM can "upscale" information and you can take a look and recognize, "yes, this is exactly as it was", even without being able to write out that "upscaled" version by yourself. Verification is often easier than direct recall.

How often do you modify the LLM output before sending it? If it’s less than 50% of the time, it means that by not modifying it, you have added at most one bit of information to what you originally wrote. (If you don’t understand why, think of it this way: Instead of sending the LLM response, you could send the prompt and one extra bit indicating whether the LLM response to the prompt should be modified, followed by the modifications.)
> How often do you modify the LLM output before sending it?

Me specifically, I never send anyone LLM output I haven't give at least a quick read (not skim, read) to make sure it's reasonable and there is no obvious bullshit there. And then I still mention it's LLM-sourced.

> If it’s less than 50% of the time, it means that by not modifying it, you have added at most one bit of information to what you originally wrote. (...) Instead of sending the LLM response, you could send the prompt and one extra bit indicating whether the LLM response to the prompt should be modified, followed by the modifications.

It's not the case, though. Prompts are not interchangeable with output. There is no guarantee that if you send a prompt, and recipient passes it to their LLM, they'll receive anything similar to what you did. It may have mistakes - different mistakes - or just spend focus differently.

The extra bits I claim LLMs can add to the message hinge strictly on you vouching for the response. Of course, you can just prompt an LLM, learn from the response, and then write your message clean, containing both the bits you originally had, and the bits you gained. But at that point, the LLM already gave you text containing all those bits - if you can vouch for it, you may as well copy it over and save yourself the trouble.

> It's not the case, though. Prompts are not interchangeable with output. There is no guarantee that if you send a prompt, and recipient passes it to their LLM, they'll receive anything similar to what you did. It may have mistakes - different mistakes - or just spend focus differently.

I’m not saying that the response is interchangeable, but that due to the data processing inequality, it cannot convey strictly more information than the prompt.

> The extra bits I claim LLMs can add to the message hinge strictly on you vouching for the response.

My argument is that if you vouch at least 50% of the time, the vouching only adds one bit of useful information – either you vouch or not.

> due to the data processing inequality, it cannot convey strictly more information than the prompt.

Only in the case where the LLM message is not reviewed before sending, and only if we assume reliable LLM (so that the receiver could recreate the same output if given the original prompt). This is not a realistic scenario.

> My argument is that if you vouch at least 50% of the time, the vouching only adds one bit of useful information – either you vouch or not.

The alternative to vouching isn't "not vouching", but "correcting and cutting out wrong bits and vouching for the rest", which means the single "vouched for it" adds all the bits that are in final message but weren't there in the prompt.

I think you missunderstand. If the llm can "add back" 200 bits of information, these 200 bits are superfluous by definition. They can quiet literally be inferred from the starting 300 bits. And the llm has proved it.

If the receiver wanted these 200 added bits, she could infer them either herself or even use an llm to do it.

> If the llm can "add back" 200 bits of information, these 200 bits are superfluous by definition.

No, they're not. Getting those bits takes energy.

SOTA LLMs know way more than any individual on approximately anything there is to know (and what they don't, they can look up faster than people can). It's very easy for them to make the "missing" 200 bits explicit, rather than implicit, which in practical terms is the same as adding 200 bits that weren't there before.

Theoretically, an idealized omnipotent mind / AGI could derive the unifying theory from reading your HN comment on a phone screen. There is enough information there, if you were able to extract every bit of evidence available from it. But you are not. Neither am I. It would take us practically infinite work to try, solving this most cruel mathematical riddle.

In principle, all of mathematics can be inferred from the axioms. The field of mathematics is not superfluous and something a receiver could infer themselves if they wanted it.

My brain does not contain all the information that can be added by an LLM; a human brain could contain it, even the biggest LLMs are about 1% of the (if you approximate synaptic count ~= parameters) parameter count of a human brain, but none actually will.

What my brain may actually contain is the information necessary to verify the (in this example) more than 200 bits the LLM claims to have added and trim out the parts which are false, retaining the (in this example) 200 "new" bits of new information added by the LLM.

Concrete example: I am a software developer by training, though not a web developer. If someone who does not have any developer experience asks me to make a web app, I am forced to use an LLM as I do not know enough JS etc syntax to get it done myself. But as we all know, LLMs are only "ok" but not "good" at making software, so there are a lot of rough edges and outright mistakes. My experience as a software developer extends to detecting such failures and I can usually correct them.

The original person, someone who has no developer experience, can also prompt the LLM. Right now, this would result in something that retains all the errors, because they didn't have someone like me intermediating between them and the LLM.

I add bits by removing noise, the LLM adds bits but they contain noise.

I do not know for how long this will remain true, but today it is true.

> Typical case is like this (…)

> This is the optimistic scenario

So is it typical or optimistic?

> I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself

So why are you so eager to defend your fantastical scenario? It doesn’t matter how considerate you are, truth is the overwhelming majority of people aren’t and won’t be. We’re discussing reality here, not “what could be if we lived in a utopia which will never come to pass”.

> So is it typical or optimistic?

The optimistic case is "having 500 bits, giving LLM 300, getting back 1000, and learning extra 500 in the process". Real numbers are lower. People don't vouch thoroughly and don't catch all mistakes.

But reasonable people don't send every output from LLMs to others without giving it a cursory glance (obvious hallucinations or nonsense would paint the sender as incompetent or inconsiderate), and that alone eliminates the worst levels of noise. A cursory read and cutting out obvious bullshit before sending is enough to make the message carry more bits of information than the propmpt.

> the overwhelming majority of people aren’t and won’t be.

In my experience, the "overwhelming majority" are giving something between a cursory glance and cursory edit; whether the resulting message has more or less information than prompt then depends on how much noise LLM added on top. The inconsiderate people I deal with, they often send "net more bits than in prompt" outputs, but those outputs are also verbose and not fully filtered for bullshit, thus it's effortful to tease out the signal from noise.

I have 300 bits of information in rude direct form with some obscene language as well. I ask LLM to wrap it into a nice polite message. Resulted 1000 bits are sent in the email.

I bet no one would like to get direct rude "source" instead.

> I bet no one would like to get direct rude "source" instead.

You bet wrongly. Rudeness carries information.

“Fucking hell, how many times have I asked you to XYZ” is different from “G’day gov’nor, terribly sorry to bother you. May I remind you to XYZ? Would you mind doing so at your earliest convenience? My deepest regards, toodeloo”.

The former conveys urgency and annoyance while the former conveys that you can keep ignoring it (and straining the relationship).

The former also conveys that you are rude, which is not what you usually want.
Brit here.

The latter (with its twisted mix of Australian, Cockney and Kings English) carries a calm sarcastic tone which indicates ones displeasure far more than the former, more vulgar statement, could ever hope to achieve.

One is however, reminded that Americans simply don't get our sarcasm, frequently leading to some amusing cultural clashes.

gosh...this twisted indirect way of communication makes me nuts. Just be open and say what you mean and stop being offended so easily...
If you equate sarcasm with ease of offence, you are very much mistaken.

Regardless, in Blighty, such sarcasm is by no means twisted and indirect. It is as plain as the nose on ones face. As I said earlier, many Americans simply don't get such sarcasm, leading to comic clashes of culture.

> with its twisted mix of Australian, Cockney and Kings English

This is a great way of summarizing a lot of the TV we’re getting in the US that have British characters, and it’s annoying once you notice it.

See how the accent changes in Lie to Me between season 1 and season 3:

S1: https://youtu.be/bWyhsqh_e9s

S3: https://youtu.be/oPqOET_xCKw

The latter one conveys a sense of urgency in the sense it burns more tokens on the govenor's side.
Then you should push everyone to talk like a caveman. It really saves time and tokens.

But there is a reason no one is talking that way...

I would much rather get the rude original message.
>If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer.

A math teacher only has "this class is about math" to transfer, the rest is known. ;-)

Jokes aside, I think we're in a weird transition now where AI is used to generate text that looks good but is bad. In a few years people will know that and be more critical.

I think we went through a similar phase when DTP had it's breakthrough. Suddenly school papers were laser printed 300ppi times new roman and got more attention than better papers written by hand. But eventually that became the baseline.

I think Ai will make it so that well-written texts with clarity, good layout, correct illustrations, callouts etc become the norm, and will no longer impress anyone unless the information itself is actually good.

And if the information is good, it won't matter if it's AI generated or not.

Maybe things will change in the future, but I know where things stand right now. Would you rather learn about math by interacting with the teacher, by using the internet (including public LLMs), or by having the teacher tell an LLM "teach a math class" and then copy-paste its response to you? Option 1 is the best, but option 2 is at least better than option 3.
>Maybe things will change in the future, but I know where things stand right now

It's not an either-or though, LLM generated text already provides lots of value in many situations right now. It's also used for fluff, yes, but much of what humans write is fluff too, reporters often get paid by the word.

The point is that the value in a text has nothing to do with whether it was generated by an LLM or not. What matters is if it's useful or not.

Oh it very much depends, your opinion there is certainly not always true and not shared by everyone; you only know where you stand right now with any certainty.

It depends on the teacher. I’ve had good and bad math classes, and LLMs today are quite a bit better than the bad ones. The worst human teacher in my memory didn’t offer interactions. He walked in, turned his back to the class, wrote equations on the board for 45 minutes, and then left. The best math teachers, the ones better than LLMs, are the ones who share the joy and sense of discovery and history of math, and not just the mechanics. But there aren’t that many teachers of that sort.

Learning by using the internet without LLMs is rarely very good, but more often than not in my experience sucks much worse than using LLMs. If you include using public LLMs in internet usage, then it’s not very different from just using LLMs that search the internet. I have heard that a huge swath of today’s high school and college kids are reaching for chatGPT before Google (which is incidentally OpenAI’s goal), and that many of them would rather talk to chatGPT than talk to a teacher. I’m going to refrain from making any claims, but I believe there are a lot of people who disagree with your ranking of the options.

Foreign language learning is one case where I love using LLMs, because it’s not typically an option otherwise. You can practice non-stop and have conversations with someone fluent in a language who will be infinitely patient with your mistakes. This is true of math and other subjects too; using LLMs to practice, so that the human teacher isn’t the bottleneck, to supplement and reinforce the human interactions, is usually better than using the internet without LLMs. The other reason many people prefer talking to LLMs is the lack of judgement. If you aren’t getting it and ask the teacher one too many basic questions, they treat you differently. Sometimes it’s necessary and helpful, and sometimes it’s harmful and takes a long time to change. LLMs don’t do that, they just explain and explain. That lack of judgement is a big reason many people prefer LLM interaction to human interaction.

Okay, sure. The difference between option 1 and option 2 is completely irrelevant to this discussion and I gave what I thought would be an uncontroversial take on it because I thought it didn't matter. The point is at least one of option 1 or 2 is almost always better than option 3. Option 3 is what I and others are arguing against.
Irrelevant? How so? And why did you present two different options if you thought it was irrelevant? I thought you were trying to make a point about human interaction. From my POV, your options 2 and 3 are much closer together than 1 & 2, so if you’re not suggesting that human interaction is preferable to LLM interaction, then what’s your justification for claiming any option is better than any other? I still completely disagree with your clarified point, and suspect many many other people do too, and are demonstrating that preference by using chatGPT for their classes instead of using either the teacher or Google. It’s a simple fact that interacting with an LLM can be preferable to interacting with some teachers some of the time, and LLMs currently often make interacting with the internet more pleasant and more efficient. Your ‘almost always’ still sounds to me like one person’s opinion that is not shared by all.
We're not talking about humans vs LLMs in general, we're talking about humans copy-pasting LLM output as if they wrote it themselves. It's less "LLMs are bad" and more "if you want to communicate something specific to me, LLMs are a bad substitute for your own writing".
Oh I see I’ve misunderstood. Your analogy is restricting the scenario to no interaction with LLMs. In that case, there are still some reasons LLM writing might be prefereable for students: for non-native speakers, and for math teachers new to match developing a new curriculum (sucks, I guess, but it certainly happens), just to name a couple.

One question to ask is why is your contrived LLM scenario any different than a math teacher using a textbook, or district/state worksheets? Math teachers typically do not write the course text or exercises, and never have. Very few math teachers do their own writing.

> you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer.

It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't. As an existing example, consider formulating a conjecture and having an LLM prove it. But I agree; if I wanted to read an LLM's output I'd simply ask it myself rather than read someone's supposedly-human writing.

>It doesn't actually follow, because maybe the LLM is smarter than the original writer (at least in the domain the writing is about) and hence really is able to complete the ideas in a way the writer can't.

Then what's the point of the original writer?

I think there are a lot of situations unaccounted for here

- people writing in a non native language

- people insecure in their writing

- people not used to writing in industry terms

- people with the curse of knowledge that are aware that they can’t write for a general audience well

Surely others too. None of those mean you have to read it, but I have gotten immense value from reading some things people have had ai write (and I’ve seen a ton of junk as well)

> people writing in a non native language This category of people are going to be most betrayed by LLM output: 1) the receiver loses the signal that the sender might need to be queried to find their real intent and 2) the sender doesn't have the ability to determine if what they are sending is what they mean

> people insecure in their writing These people can grow up, I don't care. Not a good enough reason to send a slop grenade.

>people not used to writing in industry terms Similar to the non-native speakers, but slightly less in magnitude. They can educate themselves though.

- people with the curse of knowledge that are aware that they can’t write for a general audience well These people probably can get some value out of it but they should take care

Still not really a good enough reason in the end

That’s like, just your opinion man.

I disagree, but I think neither of us are going to benefit from continuing this discussion.

I don't think that necessarily holds, but it needs nuance. I tried to convey this within my company by giving a "guide" as to how to use LLM's for writing, across three modes:

1. Transliteration - roughly keeping the number of characters or bits, but translating to a different lingo, language or mental model (e.g. metaphors). Roughly the safest mode, but can still yield catastrophic results - it's safest if the author still provides taste and editing.

2. Compression - taking out redundancy to make the text more dense and more salient. The LLM chooses what to take out - and might take out the wrong things. More dangerous - but if you're happy with the salience and you believe the reader won't have time to read the uncompressed - it's probably safer than having the reader LLM compress without the benefit of your editing process.

3. Decompression - using the salience of your idea to add detail to the reader who wants to understand it fully, by utilising knowledge that is common to you and not common to the reader. This can be very powerful when there's no time to fully write the thing by a human - but it's the easiest to get wrong and to create slop. As an example - you could try explaining concept X + illustrate it through 3 examples. You know the examples are in public memory and easily retrievable - so you write your explanation of concept X, list the examples you want - and the LLM can take all of them, synthesise and bring the full package from your 300 bits to 1000 bits.

You are right that those are not the exact 1000 bits from the original brain, but they could contain 900 of the 1000 - which is still better communication efficiency than transferring 300.

I am however, more and more in the camp of fleshy brains writing everything, as my slop allergy rises.

How do you explain professional speech writing?
Professional speeches are usually equivalent to LLM slop in terms of information content. They're written to sound nice first and foremost, and communicating information is a secondary goal.
> If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

LLMs do inference or computation among other things, so the remaining 700 bits can be something like that. The hidden implication in your claim is that computation adds no information content, which leads to an interesting philosophical discussion.

So for example, if I ask an LLM to give a proof or derive a new theorem from a set of axioms, according to your assumption, if it answers correctly, then I haven't learned anything new.

I am not really sure how to resolve this paradox in information theory.

I like your premise but I will point out that communication isn’t just about conveying information.

It is often about persuasion and that sometimes benefits from framing effectively, which I think an LLM can help with given the key points you’ve got.

Using an LLM to write something persuasive is a good way to persuade me that you don't care about the topic personally, and therefore I probably shouldn't care what you or your LLM think.
So I spent a full workday investigating an issue with an agent and then I spent another hour prompting the agent to summarize the findings so that I can send them to a colleague. How does what you're saying come anywhere close to describing my work?
Correct.

The real benefit is that the use of an LLM allows me to convey to it 1747 bits of scrambled information in an order that fits what's actually inside my head, and then have it unscramble that and convert it down to the 1000 you need. It can do that better and faster than I can.

It massively reduces the time it takes to make a short letter.

No, thanks. Depending on the type of information being transferred and just how uncertain the recipient's knowledge is, I may or may not need to add some amount of information to whatever I have to share. That information would be extra context you may need to have, on top of what new thing I am trying to explain. If I know I'm talking with someone who's all caught up, great, no need for context setting. Though even in those cases, there might be some terminology I want to define, or to provide a list of acronyms or whatever. That is boilerplate work, but is at least useful, if not needed. I know people tend to view this sort of comms as "technical" - they shouldn't be. So much misunderstanding happens because people are just not aligned on their shared context or medium. Of course, AI is not fully reliable to generate that context, you as a writer need to put in the work to review and align the result with your own understanding, or drive it thoroughly before.

All that being said, I acknowledge that people (me included) love to be sloppy in their comms (with or without AI) and then blame others for misunderstanding. It's also unlikely we'll change soon. What can you do.

What if the bits I'm transferring are "I don't remember the precise syntax, but here's a link to the relevant documentation."
I think you need to take information compression into account. Compression naturally created by shared understanding of concepts and acronyms.

Can you explain to a child what it means when "GitHub is down" with less characters?

For a given information model, compression isn't a useful term, because "information" refers to a concept at its maximum compression. By definition, it can't be compressed further.
While I acknowledge that this is meant to be framed in the context where the message is fully compressed, I think particularly when the medium is language (usually pretty compressable), then actually anti-aliasing the transfer of information is still 100% relevant! After all, AA is really about making better use of the information that you actually transmit. Compressed or otherwise, 1000 bits of aliasy garbage is not in any way the same as 1000 bits of perfectly antialiased signal.

I like to think of this idea of transfer of information (via words, let's say) from speaker/writer to listener/reader, as actually a fairly normal signal processing situation. One where the constructor of the message during it's 'rasterisation' from continuous thoughts to discrete words, can do the job well (ie: band-limit and 'antialias' the message, fully considering the target sampling domain), or badly (ignore the target domain and just speak/write from the source context, leading to 'aliasing' in the listener's received message).

I know this isn't how many people think about words, but I think the idea of antialiasing is relevant and something that people should think more about (including with respect to AI-generated text).

It's tricky because I don't want to write what no one is going to read (but that for some reason I'm required to write).
Nobody is going to read anything regardless of whether you used a LLM to write it or not (but if you use a LLM they'll use that as their excuse for the virtue signalling aura farm), so honestly why bother writing.
If it's written down in a form that can't be repudiated, a succinct message to any stakeholders covering any decision of substance is good practice. They might not read it, but they cannot pretend I did not make a good faith attempt to inform them.

(And I try to be judicious about what counts as "of substance" so they can't accuse me of flooding them with messages)

I appreciate the direction of this article, and commend the author on publishing work, but the first sentence in the first paragraph is exactly what they lament.

> A pattern I see is that people use AI to build something new, then they use AI to retrospectively summarize what they have already built into a design document. Reading a document like this isn’t just difficult—it is punishing.

FWIW Pangram [0] "We believe that this entire text is human-written."

I've only just discovered pangram, but I've seen it referred to a few times in HN recently wit nothing obviously pejorative about it. Take it with all required grains of salt, though

0: https://www.pangram.com/dashboard

I believe that to be a very obvious joke. The kind that the british are fond of.
The not this but that is bad when it is out of nowhere combination. In here, it has actual meaning.
I was about to comment on the exact same thing. Stopped reading after this. AI writing patterns has started to show itself in my AI allergy. Don’t get me wrong I like AI but it’s a pain to read the same style over and over again.
(comment deleted)
These last few days, there has been a small stream of blogposts on HN expressing similar things, and I have enjoyed them all.

My question recently has been how to broach this subject with colleagues who really enjoy producing prose with AI. There is not yet a better cultural shorthand for this sort of thing than "slop" which is a harsh-sounding word and itself sort of a thought-terminating cliché. "I don't want to read what you didn't write" is maybe closer — but it needs a pithier and somewhat more encouraging encapsulation, like "I want to hear it from you".

Has anyone had good experiences setting up professional boundaries or team norms around AI-written docs?

im sensing about a 20 month lag between hn sentiment and normy-coworker sentiment. managers and business-types pointing out the benefits and challenges of working with an llm i told them about early 2025. im just politely nodding along waiting for them to catch up. theres some emperors new clothes happening at my company so maybe im complicit
"Hey, the comments are too wordy. Can you clean it up so its only important information?"
It's something that my company has been struggling with. Developers are generating mountains of code, documentation, and Jira ticket comments. It's incomprehensible and overwhelming.

For the team I lead, my guideline is AI generated is fine but it needs to be human-edited and/or summarized. You want me to read what you're offering? Put some effort into it and meet me halfway. I don't want AI generated gibberish with made-up terms. You'd better also understand what you are presenting as your work. It's been fairly well-received though we're still working on it.

I have some co-workers on other teams who use AI to generate responses to literally everything. Ask a simple question? Get pages of AI generated nonsense in response. They are proving a tougher nut to crack.

In my current position(DevOps) we are getting this kind of messages or issues a lot recently. While forwarding these issues to my team lead, i don't have to deal with them which feels great but once its assigned to me... Its another long process to understand what issue actually means and what i need to do.

Thats said, i don't know any professional way for this but currenlty trying to communicate with peers in order to work without a problem.

I’ve seen the problem crop up in several places.

At a broad level, the issue isn’t generated content, it is the ratio of verification capacity to generation capacity (V/G). Your pain is because generation capacity has increased significantly, while verification is laborious and capacity has not (and can not) catch up.

Unlike spam which is from external sources and can be ignored, messages from other employees have to be responded to. I guarantee this is creating bottlenecks all across the firm, outside of the individuals who are feeling productive.

For fixes, theres theoretical approaches that might work?

If you need leadership to help you, then this issue has to become something that is on their radar, which means that something needs to go wrong or costs need to be registered.

The shortest conversation for that is to make people aware that generation has improved individual productivity, while moving the costs of that production to the rest of the firm.

If leadership is not at the stage to listen, then you need to move the costs you are incurring to the people who are sending them to you. Maybe set time aside to sit down with whoever sent a PR and then read what they sent together, to understand it.

It also makes a difference if tokens are being subsidized or not. If the firm doesn’t care how many tokens are being used, then you are naturally going to have over production.

Ask questions about how the audience is taking the blog posts. Say that many people seem to really not like LLM writing and ask how are they listening to their readers on this point.

You may find that brings a better and more human focused discussion. Who is the audience and what is best for them? Don't frame it about your preferences but about the real persons that your colleagues seek to talk to.

This sentiment gets echoed over and over. While I can certainly see the value in human-created creative works, if the subject is pure facts, there are only so many ways to state the facts without being caught in the weeds and missing your entire point.

Moreover, people are finding it hard to differentiate what is an is not AI-generated with newer models, often attributing original work with those of LLMs. It has just become an easy scapegoat for lazy comprehension and a desire to do less. You are jumping at AI boogeymen.

Just about the only thing here I can level with you on is, yes, AI is far from perfect and will continue to advance. Otherwise, so much of this reads as fruity prose to excuse apathy.

The lazy ones are the people spending 60s prompting then hitting copy/paste. They dump it on me as is and call it a job well done, which leads to me doing their job for them or calling them out to stop the behavior.

I don’t like doing either of these things. I don't want to do your job and I don't want to lecture about how it isn’t “your work” if you don’t touch the content after an LLM spits it out. It is rude and selfish to put me in that position.

The title is the conditional that underpins the claim.

Do you read our AI output end-to-end and decide that it’s the accurate content communicated concisely? If yes, this is not an issue.

If you are sending people generated text you haven’t read, how do you know what you sent doesn’t fall in the category of writing the author describes?

“Reading a document like this isn’t just difficult—it is punishing.”

I’d be curious whether the author composed this sentence himself or it was the output of AI. Personally I often find myself “it’s not X it’s Y” and then recoiling in disgust and rephrasing it simply because AI has made it so grating from overuse.

Same here. I also used to love writing with em dashes, but don't anymore out of fear people will call my writing AI and stop reading. It'd not hard to do in most word processors, typing two hyphens "--" does it in Word and Pages, so I don't get why people think it's such an obvious AI tell. It's not that exotic.

Maybe we will get to a place where we normalize AI authorship or co-authorship in our writing?

I appreciate what the Palm Springs Post does with its news articles. It has an "AI Assist" author and gives it credit in bylines as appropriate:

https://thepalmspringspost.com/author/ai-assist/

That's nice. So when I'm reading I know whether it's a human author or not. And if it's a dry municipal meeting, seems like a perfect place to send in the bot!

It was very clearly on purpose, if you look at the surrounding prose.
It's funny, i push back on pull requests because there is too much description now - a 20 line change has pages and pages of generated description, rationalisation for why it is safe, defense of each design decision, analysis of risks and side effects. People are indignant, you're rejecting my change because there is too much documentation? And my response is, I don't have time to read it and you put me in the position where I can't afford not to - because approving the PR implies I did and accepted it. The investment to read all that for the value of a code change that I'm one prompt away from doing myself if I cared is just not high enough. So it's rejected.
I'm naturally verbose due to my ADHD. I'm trying to learn to be more succinct.

In the meantime, for business communication, I use AI to shorten my text, to make it more concise.

I also have ADHD-driven blah blah blah. I find achieving brevity laborious, but find LLM results to be unacceptable. A journalism class taught by a well-known professional food critic helped most, but consistently using the free “Hemingway App” editor was almost as beneficial. Blindly following every suggestion pares your writing down to a nub, but it’s great at highlighting sentences you need to rethink without blurring your voice and ideas with bland suggestions.
That's a solid recommendation, thanks. I also did journalism. So I can write succinctly, I just have to put effort in like you. Which is much more than I'm willing to put in after I just wrote a treatise on the human condition to send to the lead developer at 07:30 on Monday.
Maximizing the amount of documentation was never the goal, and people who assume that have entirely misunderstood what the point of documentation is.
I’ve started just asking it to touch up text to make it sound more human and follow technical writing guidelines and it converges on convincing pretty quickly. It also makes it easier for me to verify that it’s correct because all of the stupid LLM noise and formatting goes away and it gets distilled down to the important bits.
You might be surprised how many organisations/ teams now don’t even review pull requests. Claude does coding, Codex do code review, and if developer did 5 round of this before pull request, we could very well just let CI merge.
a whole organization just pressing "OK" intermittently
I'm very curious how this goes long term. I guess we will find out.

My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.

But this may be all my human-biased fantasy that justifies still taking a role in software development.

> ... these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher ...

wow this is a beautiful way to put it

That is exactly how one could describe the relationship between heroin and the users health systems. It’s okay for a little while, until it isn’t but there is no going back. When the risks start telling, lawyers will have a field day. Assisted by LLM’s obviously, which then puts the burden on judges, who then need LLM assistance.

This thing called attention economics has captured my attention quite hard. Digital opiates are everywhere since 10 to 20 years. And in my view, the part that they implicitly of explicitly strive to capture your attention is new.

Wasn't this always the case - in enterprise software? At some point engineering pushback collapses, and the thing just gets rewritten.
(comment deleted)
Well articulated. I tend to think that the outcome will vary hugely by project. We already observe this with some reporting certain models behave in a certain way, whilst others see the opposite.

To use the wooley term “quality”, the top 20% might stand a good chance of making huge strides. But the remaining 80% of projects (in particular the bottom 20%) will atrophy extremely quickly. Yet, these will be the project that many push LLM’s too as their domain/technology is complex and/or outdated. Digital transformations that can be done quickly will be tantalising but ultimately unsatisfactory long term (as you describe).

A poll of estimates would be quite helpful in figuring out what the heck is actually going on.

How many are you seeing / estimating?

[delayed]
They may have done you a favour if that's what they sack you for, good luck finding another job!
Others are just implicitly doing this but pretending review still exists.

Everyone is fatigued by endless code review which you get no credit for and has become massively more of a burden.

All PRs are superficially fine now. There are no typos, there is unit test coverage, but there are deeper issues that require massive amounts of effort and time to spot.

Unfortunately, a lot of previous PR reviews already were just gatekeeping, or "presenteeism". People would leave comments about class names or method names, like they couldn't understand what an `apply` method, the only method, meant on a class with that was already named appropriately and did one thing only (to give an oop example). No, the method needed to be renamed `PriceChecks.applyPriceChecksWithTimeConstraints`.

Lots of review comments about various conditions that wouldn't feasibly happen (same shit with claude now).

But then I'd see these same reviewers approving PRs where the bigger design was just fundamentally broken. Oh, we're adding a blocking call on our hot path, but at least the method name makes it very clear that it is blocking.

In general I agree that the current AI reviews are creating too much noise and it is masking these bigger design issues.

Source? Or are we just upvoting wild speculation based on vibes on HN now?
IMO, PRs were never the right process to use in tightly-collaborating teams such as most companies. PRs got popular because orgs started using GitHub, and GitHub had made that workflow to fit the needs of open source (and modeled it after what open source was already doing in the days when patches got sent around via mailing lists).

I'm old enough to remember using CVS and then subversion in companies. People would commit straight to main (which was then called "trunk"), because making feature branches and merging them was cumbersome. And, on regular intervals, the person responsible for some corner of the codebase would do a show-and-tell presenting it to peers, but without the sharply defined boundaries of what the code looked like before vs. after some recent set of changes. People might remember some things from the previous show and tell or from first hand experience with that code, but that kind of memory is necessarily fuzzy, and diffs weren't an artefact that was typical to look at. So, these reviews didn't block people, and any comments that came from reviews defined a direction that things should go from here on out. If a corner of the codebase was deemed to be in a bad shape, the blame around that was equally fuzzy.

Summarizing large PR documentation sounds like a job for AI...
(I assume downvoted because people didn't get that this was joke)
The one thing that gets with those PRs is that they want you to read stuff they haven’t even read themselves! Yeah, no.
I just have my LLM read the PR, then give me a summary of what the PR is about, how important it is, and how good the PR itself actually is. Then I have a conversation with the LLM about specific points, especially things where I get that feeling that I don't have a 100% understanding.

Until I fully understand what's going on, the PR doesn't move and my interrogation of the LLM doesn't end. My interaction is littered with "Explain X" and "How does this square with Y?" and "What if Z happens?"

The interrogation is the point, without me having to wade through hundreds of lines of irrelevant code to get at the meat of the matter.

At this point, what is a human dev even there for?

We have this at work : fully AI-generated code and description. People will give review comments generated by AI which the "author" replies with an AI-generated response, all with LLM wording full of jargons no one understands not even the person who sent it. When you ask them what they meant, yeah idk Claude said so

The human is there as the control valve, the one who keeps the overall context, and the one who injects actual creativity.

Every time the LLM throws jargon around, you call it. "What do you mean by gated wedge?" You call its bullshit, check what it's saying against your understanding of the overall system, and keep it on the straight and narrow.

It's a lot like supervising a junior dev who happens to be very quick at absorbing lots of info, but not so great at the big picture.

Why would I do that? It's the job of the person sending me the PR to explain what they're trying to do. I'm done wading through AI slop.
Lately on my hobby projects I've been lazily committing work with the prompt 'add the relevant files to git and commit with a message that explains why' - only because I remember 15 years ago reading an HN comment from someone complaining that too many commit messages answer 'what' but not 'why' haha

So far I haven't had a reason to go back through commits to isolate any issues but if I do hoping the 'why' messages may come in handy for my LLM lol

this is intended as a silly observation and not an actual accusation, but for an article about not writing with AI, the following is one AI-ass sentence:

“Reading a document like this isn’t just difficult—it is punishing.”

(comment deleted)
There's nothing AI about that sentence.
It's the stereotypical example of AI in my experience. I.e it's not this it's that.
It's a super common writing and speech pattern.
You are oblivious to AI smells then. This is a very common AI writing pattern.
AI didn’t invent the em dash nor “It’s not just X, it’s Y” and sometimes those are appropriate for the situation. Prose should be judged accordingly.
LLM didn't invent it themselves, they are plagiarism machines. Obviously these kind of patterns had to be popular in the first place to be overused by LLMs

Also we have reached the point where LLMs are pubkishing more contents than humans so we have probably reached a point where humans will inconciously start copying and using LLMs style too.

c'mon, it's literally the top 2 AI writing tropes in a single sentence. i googled "AI tropes", and here are the first 2 tropes listed in the first search result:

1) negative parallelism: "It's not bold. It's backwards."

2) em-dash addiction: "The problem -- and this is the part nobody talks about -- is systemic."

again, i'm not saying the author wrote this sentence using AI. i was just tickled by how stereotypically AI it sounded. for better or for worse, those characteristics are just going to set off the AI detector in people's heads now, as evidenced by the fact that there are like 10 separate comments on this post that picked up on the exact same thing.

Also, writing, indeed any making activity, forces you to think clearly. Often that is the deepest purpose of prose, even code - forcing the author to confront the problem and sharpen their own understanding of the world we share. That it then communicates something to another person (or to a machine), is a happy bonus.
How do I type this character on my English keyboard? — also not reading what you didn't write.
On Windows, use wincompose. Space followed by a hyphen gives you an emdash with the cute half-width spaces around it. On Android, I can get an emdash by long pressing hyphen. There's also a keyboard shortcut on the mac for it. No idea about Linux/BSD or iOS, unfortunately.
On Windows I also use QuickAccent, now part of PowerToys. I could never remember the Alt-code for the em dash. Apparently it's Alt-0151. With QuickAccent I hold a letter and right or left arrow and up comes a menu of related symbols. It's pretty neat.

On Mac, it's option-shift-hyphen. Option-hyphen gives the en dash.

On Linux, you may be able to type Ctrl-Shift-U to get unicode entry and type 2014 and Enter. 2013 is the en dash. Again, I will never remember that.

In Linux if you enable the compose key it’s the compose key followed by 3 dashes.
on a Mac, it's Option+Shift+-

— (em dash)

Option+- gives you an "en dash": – which is longer than a normal hyphen.

On Linux, it's whatever you say it is, after reading through `man xcompose`.
How do I type this character on your English keyboard? — also not reading what you didn't write.
I'd like to know how to type it with voice dictation.
On Macintosh keyboards it's ⌥⇧-. On the iOS soft keyboard, it's a long-press on the hyphen key. People often point out that various software will turn a double hyphen -- into one by default. We're overflowing with different ways to type an em dash, for those with any curiosity to learn.
> I don’t want to live in a world where you use AI to summarize something important into unreadable text, and then I use AI in an attempt to decipher it. I want to hear you, imperfections and all.

It doesn't fundamentally change the equation if I use AI to prepare and then write it myself. If I'm using AI effectively, it's likely that you won't be able to tell.

This sort of post is increasingly coming off as high and mighty, where the user thinks they are being exceptionally creative and other people who are using AI are using it mindlessly.

The other comments here were written by LLMs that have a hard time detecting meta-irony.
Apparently most readers have read too much AI output that they don’t recognize a bit of irony as a joke.
> Apparently most readers have read too much AI output

You’re giving them too much credit.

Just imagine the quality of text we will have if we re-feed this into an llm a thousand times...

The juxta-positioning of the ambi-dextrous personification of the meta-sematicism is going to be both rich and soul transpiring....

It's gonna be basically a fingerprint in your soul, from my soul...

"Reading a document like this isn’t just difficult—it is punishing."

I can't tell if this essay was written in earnest or as a subtle troll.

Whichever it is, I stopped reading regardless
I don't disagree with the main points of the article. But I feel like soon with all the writing that's been hating on AI writing recently on HN, the LLMs are going to be really good at writing articles about how bad AI is for writing...
If it isn't worth the time for someone to write something themselves, it isn't worth my time to read it.
On the note of imperfections... My preference is now strongly towards reading someone's stream of thought thrown down at speed over than that same stream of thought shat through the digestive track of an AI agent. At least that feels human.

To me, it also seems like they AI digestion is getting actively worse? As best as I can tell, all the agentic nature and reasoning for code is now making writing actively worse, as the agent pulls across your whole knowledge base and will take that one thought and eagerly join and context it thinks is relevant, with the reasoning spread throughout the page.

> At least that feels human.

I feel the same. I've actually started to appreciate things I used to dislike. Like typos, or grammatical errors. I used to see it as a lack of attention to detail. But more and more it now feels like "hey, something written by a fellow human!".

Same thing for video voice overs. Things like a bad quality microphone, or someone who doesn't pronounce things very clearly. Now I go: for sure a human!

For better or for worse I’ve always preferred this, and IME a large portion of people (especially younger or who don’t write/communicate much for recreation or their work) will not read it or assume it was AI.

IMO it is a costly/goodhart-resistant way to “show your work” and help other people understand or challenge your mental model. (IE a justification for something you believe to be true).

To a certain extent are all wrong or ignorant about almost everything because our knowledge/time are very limited. But it is really important to understand what other people think in order to coordinate with them/align human goals and understanding.

It’s good that the average person taking the lowest-friction path to using an LLM in bad faith is easy to identify now. The more obvious and disliked it becomes, the more they’ll be hit with the stick to actually know things and not bother people. It’s so much worse to be “bad and stupid and not care” than “possibly cringe or wrong”

I’m fine with reading something you didn’t write if it has the information I need. I’m not fine with reading something you didn’t read and will just waste my time.
And I don't want to read the millionth iteration of some old man complaining about the kids these days, written by a human or not, yet here we are.

If you don't like AI slop. Don't read it. But wasting your time generating human slop to complain about AI slop is so obviously futile that it immediately identifies the writer as lacking the capacity for reason or emotional clarity, or merely seeking attention for their self-promotion with clickbait.

The problem is, some of us work with colleagues, or even bosses, who throw AI replies at us in nearly every conversation — so we don’t have a choice to NOT read it.
I am sorry, I didn't take into account people in literal slavery. I am very sorry to hear about your condition. Perhaps you should contact the UN or some human rights organization, they may be able to help.
Just throw it into your LLM and de-slopify it. I sometimes just ask "wtf is he on about, 2 paragraphs max". Tends to work. Sometimes you still have slop afterwards but at least it's 2 paragraphs of slop instead of a wall of it.
You don't have the choice to read or not read except when you read for leisure
(comment deleted)
There's a lot of misunderstanding here. The quality of LLM writing has not plateaued, it has dropped significantly. They've cut it to the point where people start noticing.

Do a search on "Claude Sonnet 4.5" on Reddit and you'll see lots of disappointed users [1]

It was better with GPT 4.5, 4o, and even gpt-3-davinci. 4.5 was probably the most expensive one we've seen, so my theory is that good writing is very expensive.

[1] https://www.reddit.com/r/claudexplorers/comments/1ta6f9c/i_s...

I'm curious if there have been any formal studies on this. I suspect that people have become better over time at identifying LLM writing styles, and are also combining that with selection/confirmation bias. Dunno what fraction of the perceived drop in quality comes from that.
Yeah, I think I'll start making my own creative writing benchmark. It's a little late to start data collecting, but better late than never.
Have you seen the episode "Darmok" of Star Trek The New Generation series? It is about a people with an impenetrable language aimed at packing as much information as possible by using references and metaphors, not descriptions. I guess this is where LLM language output is heading.
AI-generated writing does not bother me, unless it is intended as prose writing, where the flow communicates as much as the specific facts that are described. For technical writing, I usually just scan for important information anyway, and in that sense the only thing that bothers me about AI-generated writing is if it's too verbose and obscures the important information by leaning ineffectively on jargon and high-minded verbiage. I find that Claude Opus is very guilty of this, whereas GPT 5.6 variants are not.
>the only thing that bothers me about AI-generated writing is if it's too verbose and obscures the important information by leaning ineffectively on jargon and high-minded verbiage

But so many READMEs and design docs are degraded to this level by uncritical/unmonitored use of Claude. You make it sound like it’s an exception, not a rule.

In my experience, that brain damage inducing Claude-style technical writing is everywhere. It sucks not only because of its verbosity, but because it includes completely pointless tidbits extracted from random extended LLM sessions, with no attempt to prune redundant or gratuitous details.