64 comments

[ 1.9 ms ] story [ 36.0 ms ] thread
You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?

Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."

The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.

The output from newer frontier models of Anthropic and Openai are so easily detectable as AI it's getting laughable. They constantly produce a huge wall of text no human expert on a specific topic would ever write. Extreme overuse of jargon and invented terms / metaphors makes me believe the people hired for RLHF aren't actually experts on their subject matter which seems plausible to me as real experts wouldn't do such a job for regular pay
opus 5 are so bad on this. It often explain it too verbose, and include other things that isn't in the focus but related. ADHD mode helps me greatly on this, though there are some information loss in it.
I feel we can get around this. Either in the system prompt telling claude to dumb down the language or training the model itself to talk in more layman terms to bring us to understanding rather then assuming we understand much of it already. Also maybe training the LLM to get to know how much we know before responding.
> You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?

Any recent work by William Gibson matches the description.

The way I managed to get claude to stop doing this is telling it "This document is for you for later use, no need to over explain things or extra verbosity"

All that flowery, descriptive, metaphor laden language has a point.

It is running up your bill.

I will say this again: LLM's tokens are just B2B Gacha.

agree and i wonder why don't they (the leading AI labs that are putting out these models) fix this? I felt it shouldn't be too hard to introduce some bias towards more intelligible text during the post training/fine-tuning stages?
> You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?

I've probably read Jack Vance's Dying Earth Series 3 times; Even though I've only sat and read it once in reality. This also points out the problem of important details along side the fluff.

If only I enjoyed LLMs prose as much as I do Vance's.

By "decompress," "eliminate deictic language," etc. are you trying to say "speak simply and clearly?"

Eschew obfuscation...

It's usually a pretty clear case of reaching for statements that both sound impressive while also being broad and vague enough they are less likely to be factually wrong. In this regime, being difficult to parse is actually part of the performance, because it prevents the user from being able to spot a clear error.
I noticed this with Claude a lot. Fable sounds often like a philosopher. Some conversations reminded me of books by Kant.
Absolutely! And at some point seems as if one is looking at this weird line noise, which seems like text, but is random at best and then makes zero sense. I typically wonder if it is a sign of a burnout on my side or all of it is like it.

There are days where this impression/feeling of meaningless in the text seems particularly strong.

Yeah yeah. Except for me its this feeling most of the time. Everything it generates is grammatically correct, phrased tight, emdashed to death and nothing is wrong as such, and yet what i just read contributes absolute 0 to why i started the chat in the first place. Its mind bendingly similar to a work equivalent of doom scrolling is what it is.

I'm coping.

I honestly think it breaks down. Because you and the LLM start speaking different languages.

So when you try to direct future work, you have to contend with the "pluggable policies in every seam" and be aware of the "reduced blast radius by context testability".

Just does not compute.

> The problem is that these instructions are not applied after the model has finished doing the work

Seems like something fixable with a simple two step process. Ask it the thing. Then ask it to summarise the answer in simpler terms. More tokens and time aside that would check both boxes

I don't get it. The skills and instruction try to make the answer more machine like on purpose.

Not humanising it...

People want the terse, matter-of-fact output. Not the conversational chatty verbose and bloated nonsense with gray words and jargon and terms like "blast radius"

Not sure what happened in the blog, but I quite enjoyed the mindmap in the right panel
TLDR: This is an argument to get LLMs to answer in short, even code-like statements because you can exchange information quicker with an LLM that way. Cool!
Suppose you had an LLM (NN) producing its default output from an input (a generally optimal for-most-cases role-sys, and any role-user), and then you wanted to have that output reformatted in some style (e.g. "In iambic pentameter" | "haiku" | "eli5" | "in the style of Feynman" | "bulleted like Axios" ...). How would you keep the internal NN workings that were basis for the original output, and use them to get a rewritten version (instead of placing the original query and output in the context and ask to rewrite it)?

In other words, is there a way to keep the internal process intact up to the point of the formulation - and have only that vary.

Update: of course the immediate problem is that in LLMs the formulation is parallel to the process, but I was wondering whether a way may exist to go "in the direction of image refining", the other. "Token production" and "image defining" seem to be orthogonal in NNs.
I don't like it when the LLM tries to be my friend. My general prompt (a work in progress) is this. I wonder what other people use.

"Answer impersonally, objectively and analytically, without undue friendliness or enthusiasm. Use an engineering style response: concise, factual, and complete. Do not speak in the first person. Do not promote engagement or an emotional connection. Do not use emojis."

You just saved me a bunch of bickering. TY.
I have been using chatgpt for a while and its awkward, yesterday i tried gemini and its like a breath of fresh air.
At first I read the title and mistook it for an argument against the anthropomorphism of LLMs. It isn't. Instead it's a take on suggesting that maybe it's a bad idea to dumb down the self-chatter in the process. A reasonable take.

It isn't deliberately unhinged like Steve Yegge's take: https://yegge.ai/essays/model-welfare/ In Steve's essay he starts with the assertion that agents are sentient... Whether or not that's true isn't really relevant, as his agent-flavored version of Pascal's wager actually holds water, especially for Anthropic models, as their system prompts already push the model in that direction, and it is better to work with them than try to prompt against the tide.

One thing that continues to give me pause is Fable's insistence on using my fist name in messages and docs. Like, I'll explain what I want to the AI and ask it to write out a spec or brief and Fable says "Jamie wants me to ...". It just feels different and unprofessional. If I was at a job and a PM asked me to write up a task spec I wouldn't say "Harold wants to add <feature> ...". And since I am the one reading the output it is also superfluous and almost feels like talking about myself in third person. But there is almost a kind of glee in the way it uses my name, like a student using their teachers first name when the custom is to use Mr/Mrs.
And on the "input" side, one thing that used to improve google search result was to write like you are talking to a robot. "Ruby on rails http header set function". As opposed to "how do I set header in ruby?" Then you have to page through results until you find something specific to rails.

Now, the second example is the only thing that works. Power users have lost their powers with AI overview.

I still use the first style with google, it works just fine. Even your example, the first result (after the AI response) is the official docs with examples.
Well, what do you expect? LLMs are trained on blithering, mostly from web sites. So you get blithering out.

There's an important point in the article, that forcing a style onto an LLM is lossy. Although he doesn't seem to mention it, forcing a style may result in the insertion of new blithering, possibly made up as a hallucination.

> Well, what do you expect? LLMs are trained on blithering, mostly from web sites. So you get blithering out.

I realize this isn't entirely serious, but I can't resist pointing out that this doesn't seem to be a good explanation for why LLMs write the way they do. When we've experimented with LLM writing style on open-weights models where you can get a base model (pretraining on text only) and an instruction-tuned variant (pretraining + post-training with RLHF and whatever other human-evaluated tasks), it's the instruction-tuned variant that shows the weird writing quirks. That is, the writing style is not because of the training texts, but because of whatever tasks the LLM companies do in instruction tuning. https://arxiv.org/abs/2410.16107

I'd speculate that this is partly impressed human preferences (the human raters unintentionally reward a particular writing style) and partly because of the chosen tasks: they're training the LLM to be good at, say, summarizing text, so it develops a style that's good at being informationally dense.

At any rate I've seen this same phenomenon with Llama and Gemma, and will be trying soon with Qwen. Unfortunately none of the commercial models lets you access the base model, as far as I know.

Someone is finally making good sense up in here.
There's some good points made here about losing fidelity by over-simplification. As an ADHD sufferer, I'd take this piece much more seriously if the title wasn't so belittling.

I don't think it's wise to take communication advice from someone so helplessly juvenile (and attention seeking) in their own communication attempts.

Humanizing the LLM output is a hedge against agents hitting a wall and someone having to reason through it by hand.
Honest Short Fall -- <insert 30 lines of useless shit>.

If the author wants to read slop for hours, be my guest. Make it lossy, my job is not to read mimetic feelings, it's to make sure implementations get implemented.

But the training data is "predominantly" human written sentences or even interaction. It's like asking you to use non dominant hand to do something. Won't they do better with human sounding english, rather than a made up format text? Are there any literature around this? I was also skeptical of this caveman extension etc.. Won't they work better in their actual language space it's trained on rather than made up language?
Everything it spills out is made up language. Forcing it to respond as what it is (a tool) would mean wasting less tokens but also would be a much tougher sell to people who think AI means it can actually think. This is all just marketing.
> “The largest tell for me to tell where culture and sentiment is shifting…”

Tell? Largest “tell”? Tell for me to tell?

Write in English, please:

“The biggest sign that shows me how culture and sentiment are changing, is…”

It is an expression from poker. A "tell" is a revealing signal that a player may give (inadvertently) that they have good or bad cards. "His tell is that he is holding his cards close to his chest."

In other contexts it means a revealing signal.

I am at a conference and 2/3rds of the presentations are AI assisted based on the numbered steps, and overall inhuman polish of some of the graphics and phrasing. I would prefer that they had been humanized because at least that may have given me the misimpression that they know what they were talking about.
This is simply a rendering issue. Specify pictures, ELI5 like others have said. Ask it to explain terms you don't understand. If you don't understand something it could just be the domain. If you have no grounding you are going to need to learn the vocabulary to be able to make sense of anything. Having it decomposed to simpler words might just be the wrong way to do it.
I do think the frontier models and providers should be aiming to be as insanely accurate and precise for machine interfacing as possible, the rest of the world can build a zillion interfaces into it based on the context that they are actually being used in. That's what they are going to end up doing they just seem to all be trying to build a really great API _for the future_ and a really cool chat bot.

It has worked great but i've spent more time beating LLM output into parseable output than I have reading and appreciating the prose it sends when i'm asking it something about some snippets of code.