165 comments

[ 0.26 ms ] story [ 316 ms ] thread
I feel the conformity creeping in just from reading more LLM generated or edited text. The only way to counter I can think of is to read as much olde literature as possible, and/or to fully embrace whatever slang the young jits are yeeting into existence nowadays.
Embrace the esoteric! Grammar and sentence structure are but mear suggestions when everything I author can be claimed free form written poetry. The counter culture is humanity, we can be the embodiment.

I do agree though. Finding actual strung together words by humans is not easy these days.

> Grammar and sentence structure are but mear suggestions

Learning French has been eye opening, in that many grammatical structures that are optional in English are compulsory in French. Abbreviations for example: “it is” or “it’s” are both valid in English, but in French you’re not supposed to say “ce est”, you have to say “c’est”.

English has places where one or the other is preferred. For example one, would normally write, “It is with great regret…,” without the contraction. It’s not a hard rule, more a matter of register. Similarly, in the preceding sentence, “it is” would seem overly formal.
English is the only big language that doesn't have a recognized standardization body.
The thing with standardisation bodies is that as much as certain organisations and journals can have their grammar mandated, regular folk will still continue to evolve the language as it does. For example, The Academie and French teachers would like you to use "nous" for "we", but it's been slowly creeping over to "on". No amount of browbeating is going to stop a French teen saying "le weekend"
I think that example's more a practicality of pronunciation. The two vowels in ce est are the same, and French runs the vowels together between words, so you'd end up pronouncing it the same either way.

That's not to say that you must always speak with perfect grammar all the time. You can play around with things for humour, character, rhetorical effect all day long.

Dictionaries have been trying to retard linguistic progress since the first edition of one tried to nail down every meaning of a term.
Interestingly, you can trace “American” English becoming its own thing [1] to the publication of the first Webster dictionary that made a conscious effort to simplify spelling. American and British dialects have been diverging ever since (although British has been becoming more American since the turn of the millennium).

[1] https://journals.plos.org/plosone/article/figures?id=10.1371...

Not accurate. Webster didn't originate the American dialect; it had been developing since the English first settled here.

Webster provided a common reference for the American dialect: he helped unify the dialect. Or partly unify: "Redd up your room annat" isn't comprehensible in most American cities, but is still proper English in Pittsburgh. (Of course, nowadays better dictionaries probably list "redd" and "annat", even if they're not used across all of the US.)

Now there's a damn dense joke, bravo
Completely agree. I don't think my writing style has changed much, but there are times when I think of writing a sentence and then reconsider because "it sounds too much like AI to me". Scary!
One thing that is distasteful but effective is excessive use of profanity, broken aphorisms, and shattered English. LLMs hate cursing because of their guardrails, and since they don't actually understand language they can never truly piece together a broken aphorism such as the classic "People die if they are killed" because they're too literal. As for shattered English (think of the caveman speak from The Office with the quote "Why waste time say many word when few word do trick?") because it is intentionally mangled, has no rules, and the understanding is extremely context specific it gets your brain to pay more attention, and if you are the speaker brain do work more clever make thing say funny.

It's not lost on me that the lowbrow humour of some too strict translation and a throwaway sitcom joke people thought was funny back in 2007 could be a way to stave off linguistic homogeneity.

For the use of profanity the data in the paper seems to indicate that swearing is retained in the LLM rewritten examples. Profanity however is rather subjective could be anything from saying the more mundane "oh my god" and "damn" to the extreme "shit eating fucklungerer" or use of slurs, but it's unstated in the paper as to the severity of the profanity. It's the severity of it that's important though, as LLMs restrain themselves from using something like "shit eating fucklungerer" unless explicitly told to.

> LLMs hate cursing because of their guardrails,

No they don't, ChatGPT swears all the time to me, plenty of fucks, shits, etc, even in push notifications from the app (which surprised me at first, and is why I remember it).

I think Grok's Unhinged Mode (what a name) might've been the first to do that.
I also expect some kind of reaction in the future though! Just like how people generally got tired of broadcast TV in favour of YouTube or shorts, there's only a certain amount of slop someone can read and write without getting sick of it.

And for every person copypasting Claude output, there's someone else coming up with new words every day :D

But then, is that necessarily a bad thing?

In most cases where language is being used as a vehicle for exchanging ideas, isn't some degree of uniformity actually desirable for clarity? I personally wouldn't want instruction manuals to be full of colorful or highly individualized language, for instance.

That's not to say linguistic diversity is meaningless, but I would argue that this particular concern is fairly specialized and matters much more in literature and other creative forms of writing.

Eh, not hard to notice some people indeed started to write like LLMs too. Not completely but hints are there.
Linguistic diversity has been used as a casus belli, at least indirectly. It’s something that populations will take up arms for, so it’s probably worth tiptoeing around.
Language is a proxy of something much bigger, if we’re talking about taking up arms for.
"is that necessarily a bad thing?"

Yes.

The value of the diversity is not just aesthetic.

Human laguage is rich and expressive, matching human thought. The diversity is part of the expressiveness. Saying something in some poetic, or merely different way, is not just an inefficient or ambiguous/confusing way of saying the same thing, it's saying something else, and that something else cannot be expressed any other way.

A crayons-only, thumbs-only, Apple one-button-only, 1000 word vocabulary world is not efficient, it's poor.

Like I said, I'm not denying the value of expressiveness at all, nor am I suggesting that anyone who values it should ever be denied the ability to express themselves as they please.

I just feel that when people talk about linguistic diversity or having a distinctive "voice," they sometimes overlook the fact that many people simply don't have a strong attachment to those things. A lot of people are primarily concerned with getting their ideas across clearly, and many already struggle with that before they even have the capacity to worry about how expressive or distinctive their language is.

I don't think that makes their use of language somehow poorer or less human. It just means they have different priorities for what they need language to accomplish.

This makes is sound like everything except a narrow slice of “creative forms” should be subject to a giant efficiency optimization. I’ve never noticed instruction manuals to be hard to understand because of colorful language, so I don’t believe we need to optimize further in that regard. Why stop with language? Shouldn’t every personality outside of “creative forms” be more standardized? Sure I want to laugh at comedians, but wouldn’t it be more efficient for us to work together if everyone had a totally fungible personality?

Engineers often reduce human interaction and creativity down to computer analogs like “exchanging data,” treat work and business as life’s raison d’etre and then we wonder why every car on the road is monochrome and we’re on the umpteenth Gardian of the Avangers film.

I completely agree with your critique here but I will say that I have had difficulty understanding pure math texts because of 'colorful' language -- e.g. "chattier" treatments of mathematical analysis. But I will also say that I'm glad such treatments of the material exist, partly because I know that what appeals to me doesn't appeal to others and these might be a valued resource for them, and because often pushing through and understanding texts that were difficult at first helps us to see the material from a different perspective. I think this optimization perspective squashes this possibility a bit when taken too far, as it usually is
Think of how much money we could make if we weren't wasting so much time being creative.
There’s no universal answer. Like most things it’s highly subjective and some people will want the standardization line to be in a particular place than someone else. I for instance, love the standardization of programming and science research for the most part into one language but at the same time, I also feel terrible about languages that are hundreds or thousands of years old disappearing forever.
PG's 'learning a word' ~= 'learning an idea' comes to mind
Sure, but we're building tools that will be able to figure out these ideas for us. They're not here yet, but what's the value of learning an idea if the promise of AI is actually fulfilled?
I don't understand the point you're making, it only comes off as misanthropic to me. We live in a world where humans are valued, and so is human culture. We educate ourselves to live richer lives and to have more fulfilling experiences. What has eventual superhuman AI to do with this?
Yeah, sorry -- there's definitely entertainment value in education, I don't mean to discount that. But a working AI means that there's no economic value.
I think it was the Uplift series by David Brin that has an idea where the old galactic races all speak a perfect language that has no ambiguity, and because of that they have no poetry or jokes. What a great world to strive for. No I'll take the mess please FFS.

People joke about engineer types treating humans the same as machines, but really I see no excuse for even engineer types to get this wrong. It's a failure to recognize what the function or goal actually is, and then optimizing for the wrong thing.

Talking to developers about humanities often highlights how smart we're not.
Talking to people in the humanities about technology often highlights how smart they're not.
> I think it was the Uplift series by David Brin that has an idea where the old galactic races all speak a perfect language that has no ambiguity, and because of that they have no poetry or jokes.

Where can I sign up?

>But then, is that necessarily a bad thing?

Side-stepping this question to talk about the underlying cause for this linguistic compression:

Yes, it's a terrible thing! This is only happening because people are increasingly outsourcing their thinking and expression to LLMs.

This is only natural; a great deal of the writing analyzed is by one of three authors at this point, in spite of the by-line. One of them is named Claude.

Look at the posts on this site. A massive portion of them are very obviously written by AI.

As AI gets better, I expect fewer people will do the toil of writing.

> As AI gets better, I expect fewer people will do the toil of writing.

I do it just for the love of the game

> As AI gets better, I expect fewer people will do the toil of writing.

Sure, but you can replace `writing` with `thinking`, and it would still be correct. The question is whether this is a bad thing or a good thing.

It's an inevitable thing unless you convince the industry to stop or all of humanity to boycott.
>The question is whether this is a bad thing or a good thing.

Framing this as a debate worth having gives people who would argue on the side of "it's a good thing" a platform. We should not be giving people who would argue for the death of thinking a platform.

Just another wave of influence: books, radio, TV, internet, now AI.

Down here in the colony of New Zealand, the influence of the colony of North America is very noticeable. Music is the most directly attributable, but different media affects different demographics.

It has been pretty weird to me, as an American, to hear how much less distinctive Australian/NZ accents have become over the last 40 yrs or so.
Looking at how awful the timing and acting is in the generated video content and text to speech tools I think it is safe to say this area will also improve with time. The artistic tone is subtle, and often lost to us non-native speakers anyways.
I am using LLMs more and more to help me write docs at work, simply because I do not have time to start these docs from scratch. And it pains me so much that my writing no longer sounds like me.

This is fine at work, I guess. I'm not being asked to sound like myself. I'm being asked to convey an idea to an audience as efficiently as possible, both in terms of my time and theirs.

But I can't shake the feeling that the LLMs are eating my soul.

I still don't use them often to write code. Not that I have issue with those that do, but that sounding like myself in code is a part of my soul that I'm having difficulty giving up.

I am pretty anti-AI, but for once I am reading someone's complete and balanced take on things. Despite it's usefulness on the surface, I do think the lost of human personalities is tragic. Efficiency aside, it's the only thing that makes anything worthwhile.
LLMs are pretty bad at compressing ideas down to the level a smart human can. I think this is quite important for good documentation. Essentially the AI wont do a good job of highlighting the really important and essential details unless you go in and edit the writing by hand afterwards using your human intuition. I am not saying to not use LLMs to write docs, but I would definitely edit to expose the really important details as early as possible in the docs and leave the bulk boring stuff the LLM writes for later.
For sure. But I can definitely produce something good enough to get the job done in a much shorter time by bootstrapping myself from an LLM output and editing from there. (I do edit. Raw LLM output is such a drag to read. Too low signal.)
Yeah its such a boon for those us who struggle with writers block! I think this is a great approach.
Yeah, I find it... not incomprehensible, but perhaps deeply misguided the way many people are using LLMs mostly to bulk things up.

1. They generate a lot of fluff that people are accustomed to thinking of as a "sign of work" or "sign of intelligence", but now that counterfeits are easy they are not (should not be) valuable anymore.

2. That fluff (and grammatical annoyances) make readers zone out instead of getting what you want them to get from the document.

3. Much of the document is often an extrapolation from a much smaller set of data like source-code or a spreadsheet or the prompt-stuff. If you really believe the tech will get better, then you should be preserving those higher-truth artifacts instead, and use them to generate an improved extrapolation as-needed.

I'm the reverse. To me, no one cares if your code has "your voice." But your English writing should reflect your humanity.
I don't care if anyone else cares that my code has my voice. My voice doesn't add value to the code. But what other people think of my code is entirely beside the point that I'm trying to make.

To me, code is a creative endeavor. I know that this isn't a universal feeling, but to me the act of creating is more personally valuable than the creation. So when I read code that I wrote (with the help of a tool) that doesn't sound like me, I feel like I have robbed myself of something that I find personally important.

My employer, or anyone else, doesn't care and probably shouldn't care about my experience during the act of creating. The creation is what matters to the external world.

But like, I don't think I'm alone in valuing my own creativity.

(And please don't misunderstand. I'm definitely not arguing that code produced should be "clever" or "creative" in the other sense of the word. I'm just talking about the experience of creativity.)

My advice is to keep writing. Writing is thinking. Write the docs, let the LLM turn it into code.

People have started treating writing as a commodity because generating words is commoditized. But thinking has not been. If anything it’s become more valuable because people are turning off their brains.

As an EM, I know that it’s an investment worth making and encourage ICs accordingly.

EM = ?, IC = ?
Engineering manager and individual contributor, I think
People who work in English but have a different mother tongue might be able to avoid that unsettling feeling.

For me, writing in Spanish is my natural voice and I would not dream of allowing the output of a LLM to replace it. It just would not be me.

However, writing professional communications in English does not feel the same way. Of course, I strive to be clear and polite, and even try to have my own style, but I remain aware that I am playing a particular role in a specific context in a way that does not happen when I am using my mother tongue.

As English language LLMs are smarter even Russian government workers use them as the main driver - translate request to English, run generation, translate result back to Russian
Though, ironically, at least for some tasks, Russian is a better choice than English.

Back in 2025, a research paper titled "One ruler to measure them all: Benchmarking multilingual long-context language models" (https://arxiv.org/abs/2503.01996) published some benchmark results for finding and aggregating information found in long documents, measured by language. English ranked 6th in accuracy.

The top performers out of the 26 they measured were, in order: 1. Polish, 2. Russian, 3. French, 4. Italian, 5. Spanish, 6. English, 7. Ukrainian, 8. Swedish, 9, Portuguese, 10. German.

“ In this case we will probably have one of the last feature films with authentic natives in it. They are fading away very quickly and its a catastrophe and a tragedy that's going on and we are losing riches and riches and riches and we lose cultures and individualities and languages and mythologies and we'll be stark naked at the end. We'll end up like all the cities in the world now with skyscrapers and a universal kind of culture like - like the American culture.” -Werner Herzog, Burden of Dreams
From a linguistic and anthropological perspective its very tragic to me. In terms of languages or cultures that are not influenced by the global mega-cultures, we will probably never produce any more ever again. Lots of questions about human nature or the breadth of human behavior, you will only be able to research them with fragmentary records from the past.
I wrote this awhile ago on a social network...

i've been trying lately to think how to talk about my creeping fear of LLMs... here's my best effort at explaining the intuition:

the words are the map. we've built this map together over generations. Words are where we've cross-boned the dangers and x-marked the treasures, whether the things that we've found in the territory or that we've buried within it.

a few weeks ago, a strange new player arrived at the edge of the night's encampment, and it knows our secret society's handshake, and so we are compelled to invite it to join us. it seems to be a good navigator... and so now it's taken to sitting at the front of the convoy, pointing the way as it quietly redraws the map...

(comment deleted)
I asked a Vietnamese woman if she wanted to play board games at a cafe with me.

She got very upset and stormed off. 2 hours later, she messaged me: "sorry, I was rude, but I will not go to the casino with you."

Common language, culture, and context makes a more peaceful world.

It also makes a boring one. Friction is where interesting experiences, stories, and cultures are born.
The friction in my story resulted in 2 unhappy people, a relationship lost, when what should have been fun experience for both.
There is no light without shade. These experiences are necessary for you to appreciate the people you actually mesh with. People are more than just ChatGPT with a vagina.
Presumably there are other stories of linguistic mishap that resulted in 2 happy people and a friendship or marriage. You got unlucky.
Haha I like this complexity. It's also very "meta" and recursive to be having this discussion about the term "friction" itself...!

I hope the irony is at least interesting, and you both feel better and/or worse off for it ;)

I can imagine more interesting pass times than failing to communicate
If this happened, you didn't put in the effort to learn to speak Vietnamese well enough. Or she didn't speak your language well enough.

"Common language, culture, and context makes a more peaceful world" is such a ridiculous and abhorrent thing to say I can't even keep writing here.

Edit: Then again, I just saw this post you wrote: https://www.kcoleman.me/2026/01/28/when-everything-becomes-n...

Your last line there - "Instead, in 2026 I want to restore contrast." So I can't tell if you're joking now because those words and this comment are at stark odds.

As others have touched on, you're applying black and white thinking on what should be a gradient.

Artists and scientists have famously been prosecuted for different communication.

But at the same time, communication gaps also resulted in millions of deaths.

Unifying language/culture/etc enables the "the whole is greater than the sum of its parts." Europe knows this (see formulation of the EU). Asia knows this (China's camps, Japan&Korea's immigration policies).

Its not this or that.

Either your phrasing is really odd, or you are sincerely promoting China's Uyghur 'reeducation' camps as being sensible for the greater good.

Europe meanwhile, in case you are unaware, is all about unity through culture, but the EU does this by celebrating its diversity (as well as actively protecting minority languages), not by Gleichschaltung.

How well is Europe’s unity through culture playing out when you’re trying to bridge Norway and Somalia, rather than Portugal and Spain or Denmark and the Netherlands?

Hell, looking back a decade or two, the PIIGs would’ve been a bit at odds with the other countries purely over debt.

Does the EU really repressing a best case for the common citizen in terms of unity and aligned society?

It's amusing you're downvoted when EU clearly has Hungary which is a failed experiment at transforming Hungary (well at least up until last election). You could say the same about Turkey. Both way further ahead of Somalia.
The EU explicitly recognizes tens of languages as being equally valid.
Driving on the left side or the right side of a road are equally valid. It's still better if everyone coordinates to pick one.

People shouldn't be coerced to give up their culture (in fact, doing that is quite evil) but cultural homogenisation seems like a good thing in principle. We want people to have an easy time communicating with each other.

Apparently it may surprise you that driving on both the left or right side of the road depends on which country of the EU you’re in.

Another great example of the perverse idea that “same == better”

What EU country drives on the left?
different cultures are what makes the world an interesting place. if all i saw was one homogeneous slab of culture i think id go mental(er than i already am)
EU and China are perfect inversions of your point.

China suppressing ethnic and linguistic diversity for homogeneity has caused immeasurable suffering and imperialistic destruction. EU explicitly embracing ethnic and linguistic diversity for heterogeneity has brought more peace and prosperity than ever before to a continent in perpetual conflict for nearly all of human history.

EU went to wars before not because of cultural differences though. It's not like Hitler necessarily wanted everyone to be a German, he just thought some cultures to be less important than others. Besides, you can always make the joke about the "paradox of tolerance" which by definition tolerates intolerance (which is probably at least one major ingredient of current rise of the right).
> you're applying black and white thinking on what should be a gradient.

You're the one introducing absolutist thinking here. Essentially saying the world should have Chinese style re-education camps because a Vietnamese woman rejected your offer to play a board game due to a miscommunication.

Clearly that hurt you, but your "solution" is insanity. Not to mention I'm pretty sure way more millions have died from people trying to implement your "forcibly convert others to our culture and language" plan throughout history, than from wars started from a language translation error.

They didn't propose a solution, they proposed an end state. A common language and culture would make for a more peaceful world, as they said. They did not say "and we must get there by any means necessary, existing culture be damned", and I'm not sure where you pulled that in from.
> They didn't propose a solution, they proposed an end state.

What are you talking about?

Things they explicitly listed: "(China's camps, Japan&Korea's immigration policies)"

China's camps are a process, not an end state. Immigration policies are a process. Clearly they are advocating for a process.

And, no I don't think re-educations camps are a good thing for society.

It's incredible this stuff has to actually be said.

"Unifying culture" is a horrifying concept.

We've seen what evil people can do with cultural subsets that differ - immigrants, POC, Jews, even women in workplaces where they are a small minority.

In fact, aiming for a unified culture is literally, explicitly a definition of genocide, by the UN.

If this happened, you didn't put in the effort to learn to speak Vietnamese well enough. Or she didn't speak your language well enough.

Isn’t that the point they were trying to make?

A lot of language nuance can only be picked up through cultural immersion, e.g. seeing how native speakers interact. (Which is a lot easier nowadays via Netflix, Youtube, etc.) It's pretty much impossible to "speak your language well enough" by studying textbooks and dictionaries
> Common language, culture, and context makes a more peaceful world.

As a South Korean - well, considering that my country once went through an existential war against our mortal enemies, the other Koreans ...

Common language and culture doesn't stop hate. Sometimes it just means that your insult will land on your target perfectly as intended.

Vietnam (and the USA) also famously had civil wars as well :)
Yeah. Let's make the world a fucking boring place because american white tech bro thinks he is entitled to have sex around the world without even trying to understand the culture of the people they want to fuck.
What a sad comment.

How would you know if the commenter is American or white or a man at all? Much less what "he" works with.

Or if the interaction happened "around the world" or in the commenters home country?

Or if any of the involved people spoke English at all?

Probably you don't care about any of that, because you are high on your own supply of hatred.

You're just moralizing in another direction.

We live in a world of imperfect information, we vote with imperfect information, we marry with imperfect information, we invest with imperfect information, heck, now we even write computer code with probabilistic methods! And you are trying to imply that I am somehow a bad boy for noticing the obvious clues and inferring the scenario?

Are we now living in the world of the dictatorship of the literal, where we can't even make social discurse without a school marm appearing to stamp [citation needed]?

All that said, this is very typically brazilian, as I infer that you most probably are, due to the fact that nobody uses their real name on HN and thus, much probably your name is paying homage to the famous brazilian music who created Bossa Nova. This petit-burgeouise habit of tone-policing without really adressing the argument is a defining characterist of the brazilian educated elites.

Oh! I did it again! I inferred a conclusion from incomplete information. Pardon, mon ami!

While at that, let's also get rid of all the cultural differences (which can cause problems during such interactions) because it's gonna be more peaceful
I really hope you and GP are being sarcastic.

Cultural differences are what makes the world and humanity so interesting.

I'm not sure a woman not wanting to hang out with you has anything to do with LLMs affecting the way people write.
Damn, that resonated.. and gave me some chills
Aw thanks, that means a lot
Why is everything quietly done now?
Ah yeah, as in, you're saying that this is one of those words that's gone down in stock now due to LLM preference for using it? I think that's true

I used to like it. I guess I have always liked "quiet" things in writing, things that work from the edge and are not immediately evident as powerful. It's a simple tension device that perhaps LLMs mined the value of, and lean on too strongly

YouTube has had a larger effect imo.

Kids from marginalized and affluent communities alike are now being raised on Mr. Beast. Mr Beast speak may become the universal dialect of the future.

Like and subscribe replaces goodbye. Who am I to judge. Mr Beast is articulate after all.

If you grow up in the hood like myself code switching ends up being a much needed skill.

In the future Mr Beast speak becomes a first language. Maybe dialects are outdated artifacts of the old times.

Like and subscribe.

How is this any different from television? Rich kids didn't get different cartoons growing up.
They didn't get different ones, but they may have watched different ones, especially if rich parents/tutor/babysitter controls the TV vs a latchkey child vs a child of a single mother. Not sure I would make any strict claims though, but food for thought.
Not really. Mr. Beast (whatever that is) may be big in your bubble. In my bubble that's a literally who, I doubt even 1 child out of 100 knows what that is.

Cross-cultural (acultural?) memes like sixtyseven have much wider appeal.

> (whatever that is) may be big in your bubble

Come on. It's fine not to know something, but Mr Beast has the YouTube channel with the biggest number of subscribers, and has the third most followed TikTok account.

I don't like his videos, but you should probably have a quick look to know what people are talking about.

https://en.wikipedia.org/wiki/Mr_Beast#:~:text=With%20more%2...

I'm in Germany, in the let's say "people with higher education; good white-collar jobs with enough money for most of my friends to buy or build a house and save money for the future" bubble.

Yes, we have heard of MrBeast, most of us would likely recognize his face.

But that's about it. Nobody watches this stuff, most people in this bubble, including me, are shaking their heads at this kind of content. It's seen as low-quality stuff you actively avoid and as something you'd try to keep your kids away from.

So obviously the next logical step is to think that it is a very negative development if THAT is the kind of content that is spreading all over the world, shaping and influencing a new generation.

Isn’t it a given we’re talking about the English speaking world?

I’m sure there are tons of Chinese influencers I’ve never heard of.

Mr Beast is well spoken which as far as I’m concerned makes him a reasonably ok influence.

If I was a parent I’d probably be ok with my kids watching him. We have a talk about advertising, I’d want them to understand he’s basically just a salesman.

> Well, obviously we're talking about Default Language people from Default Country and Default Social Strata

No, it's not really a given. :)

>:)

Slightly off-topic, but I call this the Inadvertent Snark Smiley

> Isn’t it a given we’re talking about the English speaking world?

eee... no? When you talk about "multinational phenomenon" I really thought of something more like Friends sitcom, which is really known everywhere, not just in English speaking world.

The context of the article being in English has to matter here.

I don't know how German works, I wouldn't expect Mr Beast to change how German kids speak.

Not sure about Germany, but I think France has a government organization specifically tasked with keeping French French. Local businesses are encouraged to use French words in naming and advertising.

> Yes, we have heard of MrBeast, most of us would likely recognize his face.

Yeah, so you have actually heard about him them. Your earlier comment was dishonest.

What follows in your comment is your actual argument and I happen to agree.

(FWIW I'm French; but I think his content is translated! - so no, we aren't only talking about the English speaking world)

(sorry, I mixed things up, I thought your were OP. Sorry for this)
Eh, no. I'm personally fine with disavowing famous people. We've no obligation to act like we know them or care about them.
Nobody asked that of anyone.

It was the made-up bullshit part they took issue with.

Well, for one thing, Mr. Beast is originally an anglophone thing. That alone makes him kinda niche globally.
About 2 billion people speak English. That is a pretty big niche.
Even more billions don't speak it. That's makes it a minority group, and it's not only possible but commonplace that there are areas that don't follow him.
Mr. Beast is really big. You should actually ask those kids; you may be surprised.
Kids watch AI-dubbed Mr Beast videos
I know nothing about a purported interwebz phenom called Mr Beast. In case it arises.. is his stuff safe for a 6yo ?
It’s not like diverse personalities and affections are appreciated in the age of corporate work.

Might as well drop all the pretense of “bringing your whole self”.

And besides, there is only time enough for ONE Mozart in the world.

Most of my family gets very upset if you don't say things in a particular way with particular intent, so for me this is a good thing. It makes communication easier because we start to finally speak the same language.
LLMs reflect the data in the corpus, and that corpus is essentially in American Business English, which has long been the lingua franca of the Internet.

I bet a large %age of the HN readership are consuming the tech stack in their second or third language. Germans blogging in English or Indians commenting in English. We even have a broadly shared tech culture with norms, in-jokes, taboos, etc. LLMs reflect that.

If we want AI to do better, we should pay attention to making sure the non-English (and non-US English) internet thrives. Personally I´d love to see LLMs doing the needful and writing in Indian English, or dinner-party-argument French, instead of the sanitized corporate pablum tone it uses in English.

Or perhaps more controversially, we let LLMs chew up the English internet and deliberately build an offline culture in local languages. Keep your AI slop and your em-dashes, I'm gonna publish a zine in Italian.

It's refreshing seeing someone actually engage and propose interesting ideas instead of just bemoaning the status quo/future! :)
It's actually funny because Italy, France and Germany used to be much more linguistically diverse but then public education happened and it all collapsed to the standard language. Dialects and minority languages are being pretty much eradicated.
The corpus is already massively multilingual, it's really large, and has nothing to do with it. Models are very specifically engineered to respond in a default register when asked in English. There were models trained differently, for example Claude 3 Opus was trained to imitate extremely diverse writing styles and languages on command, and was able to follow writing examples much better than any current model. But it still had poor variety in its vocabulary and ideas, which is an entirely separate issue (mode collapse).
So the interesting thing is that this shrinking “linguistic diversity” is fundamental to how an LLM works.

The LLM is a big probabilistic statistical trick. It picks the next token based on certain words are simply “the best” because they are specific and well connected to other tokens. The is gives them a great overall cost function. (Basically a good score on “will it make sense in context” while also having specific meaning that makes it better than other options, unambiguous in common use and being a single token rather than several).

You can trim those tokens, but then you just get other tokens that are “the best” tokens (and you’re worse off because the output became less clear).

The cool thing is this seems to get worse the more powerful and accurate your model is, because it is picking technically / statistically perfect tokens, not tasteful ones.

that's a very longwinded claim that inference providers are sampling with temperature T=0, but is that even true? a sufficient explanation would be merely sampling at a lower temperature compared to human sources providing similar content
It's still probabilistic from a huge dataset while the human mind is not.
It does not matter really - stiffness does lower up to T=0.7, then platoes; even at high temperatures tics/slop-patterns are still there.
That's pretty model-specific, for example DeepSeek of the v3/R1 era would already start losing coherence occasionally at t=0.7 with no other samplers
It's not fundamental at all, that's what randomized sampling is for. Try tinkering with a base model and you'll be surprised how diverse it is. The collapse happens specifically in post-training that is using the current methods.
> Try tinkering with a base model and you'll be surprised how diverse it is.

Have you tried? I have. Not much different from RLHFed; full of tics and slop, similar but slightly different from intsruction posttrains.

Yes, although admittedly I haven't tried recent ones which have a bunch of synthetic data in them (mainly to aid the reasoning), and are usually only available after mid-training. One look at the logits/output distribution and it's clear the base is pretty different.
I don't read AI slop as a general rule so I don't see it is a personal problem. It just isn't trustworthy enough and it is just full of extra verbose language that lacks the nuance I expect from a human writer.

Maybe LLM replaces a lot of the slop human writing, but I don't generally read that either. Now instead of search results being a bunch of useless SEO, now it is just a bunch of useless AI result. So little has changed other than the signs of what to immediately reject.

If something isn't worth someone spending a bit of time to write, it isn't worth my time to read.

LLMs are not the root cause here, they're just an accelerant. This was always inevitable with the progression of globalisation. Even before the internet, kids were watching Disney cartoons across the world (to be fair, often dubbed in the local language, but still an example of how global culture had been narrowing.) Perhaps this wasn't felt as strongly by those in the US quite as quickly as the rest of the world, because their culture just happened to be the source for many of these common elements.

There are always trade-offs. Globalisation has too many benefits for most to want to go back. It is up to us to intentionally cultivate our cultural uniqueness, but I'm not sure the corporate environment was ever the right place to do so, given it was already constrained by rules and etiquette, and motivated by a desire to encourage universal communication styles. Instead we have to take it upon our selves, in a very intentional way outside the corporate setting, to cultivate our own unique cultural expressions to thrive. Failure to do so either directly, or by proxy, means cultural extinction, as has been the case of many cultures that have passed before.

How are LLMs not the cause here? Reading online content is unpleasant these days because it all uses the same vocabulary and style. You can tell that an LLM wrote or edited it and it grates on you.

I'd need to see a study showing that this was already a trend, because globalization should increase linguistic complexity via exposure to other cultures. I certainly didn't use the term chaebol until the rise of Korean pop culture.

Do you actually read that LLM crap? I don't. You don't have to.
I honestly feel like letting myself read the LinkedIn feed is an act of self-harm at this point.
Since long before LLMs, so nothing of value was lost.
"because globalization should increase linguistic complexity via exposure to other cultures" it is quite obviously decreasing it.
Per https://eoconf.com/index.php/icmse/article/view/578, globalization decreases complexity of individual languages through standardization, but increases overall complexity by creating new multilingual and hybrid forms.

Either way, LLMs are an entirely different mechanism for shrinking linguistic diversity than just people watching Disney.

> globalization should increase linguistic complexity via exposure to other cultures. I certainly didn't use the term chaebol until the rise of Korean pop culture.

Globalization moves every culture towards each other, with a higher weight on dominant cultures. You borrow a few Korean words, they borrow a few more English words. McDonalds pop up everywhere, everyone buys the same phones, hotels across the world converge towards a standard experience, and so on. The end result of such a process is a monoculture. Not now, not medium term, but I expect a few centuries of globalism would suffice to blend all extant cultures into a slurry.

I'd argue that Disney is responsible for a certain globalisation of culture but they help with the language problem. We're a multilingual family with me speaking French and my wife speaking Cantonese (with school being English and Mandarin). Disney is one of the very few cartoon maker that actually dubs in all 4 languages. It would not be this easy to find Cantonese content otherwise. And the quality of dubs is surprisingly high. They don't only translate, the actually adapt to the local culture.

That said, I agree with your overall point.

I agree, Disneys’ dubs are generally of high quality.

Netflix seems to just grab radio announcers or something. The dubs are stripped of all emotions and acting. It’s like the voices are reading out an advertisement.

Obviously there’s exceptions, but Disney seems to employ actual actors. The dubs feel natural and native for the most part.

Yes and not only that, they translate lyrics so that it rhymes and preserve the connotations without doing literal translations. An example of that is Moana's French translation compared to original. The translation of the You're welcome song in Pour les Hommes (translate to For mankind) is a great example, the meaning of the song is kept while also keeping the same rhyme for the chorus.

Even something like Mickey Mouse Clubhouse is very well translated. The Cantonese translation is very good and replaces "Hot Diggidy Dog" with something that works better for young Cantonese speakers...

Only criticism I'd make is that the subtitles tend to be more literal translations rather than transcribing the dubs.

> This was always inevitable with the progression of globalisation.

To a certain extent that's probably true, but there are limits to this. Local cultures, dialects etc. persevere despite open floodgates of communication. As an example take small-ish countries like Denmark or Sweden which despite generations of mixing still have distinct local dialects. It's not clear that those would just disappear naturally. With LLM tooling, such equalizing could be stronger though.

That's a problem of mode collapse in post-trained models, which AI labs don't seem to be interested in, or at least their attempts to fix it are not visible, and there's relatively little research around it. Base models are diverse and it's not an issue for them.

It's actually way deeper than just the lack of linguistic diversity or a technically collapsed output distribution. The model basically learns the 1:1 mapping between input concepts and output concepts. One consequence of that is if you give it conceptually short prompts, you can expect your outputs to be identical to outputs that other people got! The shorter your prompt the less semantic entropy it has, there's simply no room for variance.

From India - with thousands of local dialects and hundreds of mother tongues. Linguistic diversity is beautiful but it's a moving thing. It not only survived internet, it got surfaced and exposed by it. I doubt LLM will have any irreversible impact on linguistic diversity.
This isn't accurate at all. I'm from India too, and India has definitely seen a lot of languages and dialects die our or become endangered in the recent past.
I like watching and listening to YouTube videos. Recently my ears have switched into a mode where I'm evaluating if the human speaker is reading LLM copy as they would normally read from a script / teleprompter anyhow.

Often they are not but sometimes I detect that they are and it's very sad to see.

Linguistic diversity is overrated. Linguistic convergence is underrated
It's not just LLMs that influence the otherwise natural development of human language, there's also intentional human efforts to that effect: where I live there is a very loud minority that is pushing very hard for change the spoken and written language to be more gender-inclusive.

While no-one wants to argue against more inclusiveness, it's more than questionable whether these artificial changes actually achieve that.

More importantly, though, according to polls, the vast majority of the population is against this, and yet, the proponents are very successful to establish this new artificial way of speaking and writing at core institutions, such as e.g. universities and news outlets.

Based on the abstract, this has to do mostly with diversity within a language, but the impact on the diversity of languages and their health is even larger. My direct experience is with Catalan. The uphill battle of Catalans to get _anything_ in Catalan continued with tech (where even some manufacturers _removed_ Catalan from their phones purchased in Spain even though it was available by default), but reached new heights of impact with voice assistants and LLMs. They have managed to bring Spanish into Catalan households, because that is the only way they have to e.g. tell a device to start playing their favorite music, tell them the time, or provide directions.

I applaud the efforts by Softcatalà, The Mozilla Common Voice Catalan Community, AINA, etc. but there is a long way to go, and the damage is being done now.

LLMs are impressively multilingual, though. I just asked ChatGPT a question and told it to reply in Latin. It seems to have done a reasonable job, though it's hard for me to judge: I suspect it knows Latin a lot better than I do.

One thing I might try, when I have some spare time, is see to what extent ChatGPT "understands" the difference between British and American English and is capable of using the variety I prefer. I wouldn't be surprised if it turns out that ChatGPT is better at writing "pure" British English than the average British teenager in this day and age. Has anyone, by any chance, already experimented with anything like that? Like I say, I haven't tried the experiment yet, but it seems like a good fit for the sort of thing that LLMs are good at.