183 comments

[ 0.24 ms ] story [ 34.4 ms ] thread
Workaday copywriting is dead (was moribund, now has been shot in the head). Prestige literary writing is zero-sum and therefore eternal, in the same way that equities trading (not formation and not business-building) is zero-sum and therefore the actual success of LLM has not given anyone a signal advantage anywhere there, because everyone else has LLM too.
Copywriting struggles with the issue that businesses know they need it, but they don't value it. It's absolutely shocking to see how poorly companies, and governments, communicate and no amount of AI/LLM usage will change it. For decades it's been neglected and relegated to "cover my ass"- writing.

The majority of AI deployment in businesses could have been avoided, if more attention had been placed on good communication. LLMs don't yield better communication or corporate writing, just more of it, because no one in business seems to understand the value or have the ability to recognize good writing.

> It's absolutely shocking to see how poorly companies, and governments, communicate and no amount of AI/LLM usage will change it.

Thank you for saying the first part. It truly boggles the mind to see how badly most organizations communicate, even those with professional staff tasked with communications.

Regarding the second part: AI/LLM usage will lead to improvements for some people/organizations, if only for the simple reason that it removes barriers to communications. All of a sudden city hall or a third-tier supplier can quickly update their websites in multiple languages, process email communications more quickly, and empower staff who weren't good writers.

> and empower staff who weren't good writers.

That's the part I question. You have to be a really good communicator to put yourself in the place for the reader and adjust your language to your target audience. So much of "professional" communication fails because the language assumes that the recipient is knowledgeable or even just interested. If you can't communicate effectively on your own, I doubt that you can prompt the LLM to do it for you. It's going to be very much like development, if you're a skilled developer, LLMs are a massive boost to your productivity, but if you can't code, you're just producing garbage.

I am not a writer but I do have to write a lot for work. The person who used to edit things for me has not been replaced, and I am expected to use AI. It is very good, but it's not all the way there. I am now left to decide whether to take the time to do it right or to go back to doing what I actually get paid to do. I wish my organization realized how important both sides of that decision tree are: we really ought to be careful about how readers perceive what we put out there, AND this isn't my job and we should have someone who's even better than me at this.

But this is an old trend with technological developments. Thanks to e-mail, most people handle much or all of their own correspondence. There used to be specialized staff for that. They were very good at it! But they were shown the door once they ceased to be strictly necessary, and now we all waste a bunch of time fiddling with Outlook instead of doing our actual work.

> The person who used to edit things for me has not been replaced,

I assume you meant "now" instead of "not": "The person who used to edit things for me has now been replaced,"

No, that is not what I meant. No replacement for the person who left the organization has been hired. It would be inaccurate to say that they were "replaced" by AI, insofar as we had no plans to get rid of them if they had not decided to leave.
Thank you for the clarification.
"Humans are going to be significantly better at writing things intended for other humans to consume" is completely at odds with "And organizations care enough about the gap, and are willing to pay for the difference".
It’s orating.

Yes write, but then orate.

Look at all of our AI feeds.

They want to be us so bad.

Has writing ever been a safe job, even before AI?
If you are Steven King, sure.

From what I heard it's a bloodbath at the bottom.

It's a pretty reliable career for hundreds if not thousands of mid-list authors who generally stick to genre fiction. It's entirely possible that AI could core out the middle of the pack as well as the bottom, which would threaten not just jobs but a wealth of institutional knowledge that those authors share among their peers and especially new up-and-comers.
If you're good at it, which extremely few people are. It's my belief that not enough good writing exists for the tech companies to train an LLM. At least not a generally applicable on.
Tech companies are training off the shadow libraries, which have the bulk of English-language human writing in book form from the twentieth century, and that traditionally published stuff with an ISBN usually got a proper editor of some kind. So, it is hard to believe LLM companies don't have enough material; it's not as if they are limited to low-quality stuff on the web.
It might come down to that they don't care (enough). They could probably do it if they had to (hire a lot of professional editors and authors to train models), but AGI seems more profitable and also a better "story" for investors than models that are good at writing.

Thing is, LLMs don't seem to be the right vehicle for AGI, but even if the nerds understand that, see above about profit perspectives and investors. These companies are under insane pressure to deliver the next super hit.

I think it's likely to be lawyers who are the safest. Lawyers are the ones primarily in charge of the government, and I think it's likely that as soon as they feel like the legal profession is threatened, they will pass laws making it illegal to use AI for legal issues. It's not all sunshine and roses (paralegals are probably hosed), but the lawyers seem like they'll do ok.
> I think it's likely to be lawyers who are the safest. Lawyers are the ones primarily in charge of the government

Given what seems like an increasingly inevitable deprecation of these outdated, lumbering nation-states, it seems to me that these two assertions are mutually exclusive.

Eh, I think we're going to run into the same issues we're seeing in software. That is, junior and lower status positions (so SDETs in software, paralegals in law) are rapidly being downsized, and that's going to have a big issue in the future where there isn't the same talent funnel. And then in the future those that work their way up will basically be dependent on working with AI.
It is great time to start business as downsized paralegal. While ago there was an app that objected parking ticket fines in NYC. Anyone can find similar niche to specialise in, and just produce boring repetitive law work using LLM templates and mostly automated work.
If that was the logic then software engineers would stop using AI to "protect their jobs" but that hasn't been the case.

There will always be people that work against their own profession and colleagues for short term gains

I know of a few software engineers who have stopped using AI beyond using it selectively as a search / troubleshooting / stack overflow replacement.

If you are taking steps to protect your brain maybe in a way you are also protecting your job?

Who knows how this all turns out years from now.

I am one of those engineers. I just don't want to be dependent on it. I want to still do the work and learn how to be a better programmer
Unless the model watermarks text, it will become increasingly difficult to tell the difference between human legalese and AI legalese. Especially if they make it illegal, that will create a strong incentive to find-tune to imitate human lawyers more convincingly.
There's already been one case where the prompts were allowed to as a subject of discovery, and lying about it is perjury. Lawyers really are probably safe.
Pretty sure it’s too late. Lawyers are lazy like everyone else. They don’t want to go back either.
Safest job from AI is also the oldest job.
If you thought all of the money being poured into data centers was excessive, wait until people can order up a proto-sentient sex bot to their front door. Oh boy.
I think there is good chance that AIs have already replaced parts of the certain aspects of industry. And with some tuning could probably replace lot more of at least virtual parts...
The reason LLMs use the same cadence and cliches is that their RLHF does not emphasize writing well, the way it emphasizes, say, coding. And the reason for that should be pretty obvious: coding is where the big money is at.

So the author may be correct, but for a different reason: unless writing starts being very valuable as a profession, it's unlikely the labs will spend significant resources making their AI models better at it.

No, they use "the same cadence and cliches" because they inflate a short and ambiguous prompt into long and specific prose by making statistical assumptions about what best fills in the gaps. It's not a training problem, it's an information theory problem, and it's not really surmountable.

Any given model will always have some distinct implicit voice that its biased towards for that infill content, and so a popular model will always become exhaustingly common, painfully familiar, and cliche. Users can use more elaborate prompts that shift the voice away from the most normative and towards some other nodes, but they need to put in special effort for that, and what people-at-scale specifically want from these tools is to put in very little effort, so we can expect that overwhelming number of casual and naive users will always be generating cliche slop with them.

Code escapes this problem not because of training but because it specifically benefits from cliche (boilerplate, patterns, etc) and so an model whose code "voice" reflects your own taste as a coder (or your toolchain's taste as a vibecoder) is going to feel like productive output rather than slop. But it's still cliche.

No. Reinforcement Learning is doing a lot here. Anyone who played with these models before the Davinci intstruct-tuning (completion) era can tell you the same. In some ways, SOTA models have gotten better at writing, but the neuroticism of instruct-tuning has still not been resolved.
> No, they use "the same cadence and cliches" because they inflate a short and ambiguous prompt into long and specific prose by making statistical assumptions

Even if LLM output has to largely follow some statistical rules, yet, first of all, some amount of randomness is normally injected during token generation, and, secondly same true for human speech.

> about what best fills in the gaps. It's not a training problem, it's an information theory problem, and it's not really surmountable.

This is not true, as LLM has internal knowledge storet in its weight. Unless you force it to produce 2000 words doc out of 3 word prompt, you would end up adding some sense information.

>Users can use more elaborate prompts that shift the voice away from the most normative and towards some other nodes, but they need to put in special effort for that, and what people-at-scale specifically want from these tools is to put in very little effort, so we can expect that overwhelming number of casual and naive users will always be generating cliche slop with them.

True, here I agree with you. But using finetuned or simply less popular models like Kimi, Hy etc. should take care of that.

"The reason LLMs use the same cadence and cliches is that their RLHF does not emphasize writing well, the way it emphasizes, say, coding. And the reason for that should be pretty obvious: coding is where the big money is at."

Software can be checked for being 'written well' by compilers / linters etc. There is no equivalent for well-written natural prose. Spelling and grammar checkers haven't a clue about prose semantics.

They just haven't got to it yet. It's not a high priority but eventually they will make a solid effort on easier style guidance and more humanization. They will probably drop the most worn ou constructs like "not X, not Y, but Z" that have become tells and target a lot of known issues in reinforcement training.

It's slightly weird how confident writers are that it won't get improved.

The safest jobs are the ones that can't be done behind a screen. I'm a technician and there's no robot that could replace me and there won't be for a long time.
> Since LLMs lack an active mental model of a specific human reader... They cannot empathize with the human reader, as they don't have the human lived experience. And they have zero skin in the game.

LLM is perfectly capable of empathy, it just never told to do so.

Most writers today lack empathy and have no lived experience. Young californian uni graduates have strong opinions on everything, but produce repetitive boring preachy cringe stuff.

I will take well prompted LLM generated writing anytime over thst!

It’s just a guessing next word program.

It has no feelings, therefore no empathy.

Best come up out of the rabbit hole for some fresh air & sunshine brother.

> LLM is perfectly capable of empathy, it just never told to do so.

They absolutely cannot exhibit empathy. The definition of the word shows that:

> the action of understanding, being aware of, being sensitive to, and vicariously experiencing the feelings, thoughts, and experience of another

They don't experience feelings. They don't empathize.

The real thing being said here is that the author can tell good writing from bad, and assumes everyone can... I'm a visual artist and at generative art is blindingly obvious.... To me.... But not to many folks around me! Including a few artists and art adjacent folks.
LLM prose has degraded with each model update, the models are now being RL'd into wall of texts that only makes sense to other agents. I think its going to become more obvious as we move forward that is llm generated because labs seem to only be focused on tool-calling/agentic-coding environments, as that's the only thing that drives revenue.

There may come players who focus on models that are good at writing for technical writing/docs , copyrighting ect but I think people will lean towards not using them and will rather have the "human touch" for the things that directly impact brand perception.

Keep in mind, every single AI company that is selling the idea that you don't need to hire designers and web design is "solved" have $100k retainer designers crafting their landing pages.

Yeah, this tracks with my experiments. About every 6 months or so for the past year-and-a-half I've tried using the "frontier" LLMs to as-near-as-possible autonomously write novel length stories, because I find it fascinating.

When I had Gemini 2.5 write a novel, it wasn't really objectively "good" by any stretch of the imagination, but while the prose was very purple and full of cliches and, well, bad writing I guess, it still felt ... subjectively good, at least for what it was.

Last week I did a run with GPT-5.6, and wow. On the one hand, it managed to produce 110,000 words that were "shockingly" coherent. The model was able to maintain state and plot lines and background details extremely well, much better than older models.

But I just don't like the prose. I haven't really liked _any_ prose that GPT-5.6 produces. It's significantly better at "instruction following" and keeping track of things, but, wow.

> “The sequence is consistent with their voluntary choices.” Mara enlarged the uncertainty field rather than the result. “It does not prove what happened to anyone we can’t observe. It does not prove contact did this. And it does not turn the Shard into treatment.”

GPT-5.6 in particular becomes so fixated on certain ideas like "consent" and epistemology, that by the end of the narrative, the prose and dialogue are all just "agent speech", despite the prompt/harness specifying that it's a _novel_ with narrative prose and such.

Interestingly, the model itself produces an accurate critique of its own output:

> The draft has become a *consent-centered medical, legal, and logistical procedural*. The important drift is therefore not that many events were omitted. It is that the retained events now prove a different thesis.

Which begs the question of if it would do better with a couple rounds of output -> critique -> revision. But I think I've had enough LLM prose for a bit...

I'm curious if you've tried Claude. Subjectively, I've always preferred Claude's writing over ChatGPT and Gemini.
Claude is the one model family I've not really used. Which yeah, feels like a backwards thing to say in a world where seemingly everyone using LLMs is using Claude Code.

I've used some Opus 4.5/4.6 via Antigravity and Sonnet by the web chat. I'm torn because as far as LLMs go, it does feel more ... "literate".

But maybe too literate, judging by how many people are complaining about "Claudeisms". I suspect Claude would be just as susceptible, if not more, to the sort of ... "moralizing" that GPT seems to gravitate towards (for lack of a better term).

Claude was really far ahead of GPT in writing from 2-4, but the later models have started to get an overly distinctive style, wheras these days GPT tends to be coherent but fairly concise and dry.
Nobody uses that shit stop lying.

Nobody is using fucking CLAUDE get real. Anslopic.

Stop advertising that “me too” shit no one cares

The entire output of an LLM is also part of its input. Appending to long stretches of LLM-generated text, it will continue to get more and more robotic as the style of the input gets replicated and enhanced in the output.

It's possible to work around by generating short passages at a time with carefully constructed setup. But it's a real pain.

It's that, yeah, but coming from multiple orders of abstraction.

In this case, part of the experiment was to see what "oh-my-pi", a "fat and feature rich" LLM harness, could do when coupled with modern GPT, given a 6k~ word overview of a story, and told to come up with a plan to write/review/audit it, making use of subagents and all the fun new groovy LLMisms...

Part of the problem was just "it was basing its style off the last scene/chapter", but part of it was also that its instructions were constantly being "compressed" through repeated compactions. Even with the use of subagents, the "top level" agent's prompt was getting muddied, and in the "review" phase, it began to focus more and more on creating increasingly complex ledgers.

You can see this happen in the "plan" files it created for each chapter, looking at word count:

   1304 d1-ch-01.md
   3701 d1-ch-02.md
   5151 d1-ch-03.md
   6462 d1-ch-04.md
   9587 d1-ch-05.md
  10605 d1-ch-06.md
So it wasn't just that the prose was being based on an increasingly compressed "style" of the prior context window, but the planning for writing each scene was, itself, becoming fixated on the "continuity error correction" process itself, to the point where by the end, it had mostly forgotten about the prose part, and was completely fixated on ensuring maximum state continuity.

This could definitely be fixed, but honestly, I've about had my fill of the "autonomous writing agent" goal. The idea was to make a model that could generate sufficiently interesting stories based on "vague premises" for my personal entertainment, but, "surprise", getting LLMs to actually produce both "new" and "coherent" content beyond what you specify is _hard_.

It seems like you really do need to just stay "in-the-loop" with every scene, and constantly provide correction/feedback, to correct the "semantic drift".

Or, gasp, I could just try writing things by hand again... :-)

> write novel length stories

This would never work. Anything longer than 1500 words gonna be bad. To get proper quality you should generate piece by piece then stitch.

Sure, one shot generation doesn't work, you have to split into sections and generate each section independently, then usually do a "seam" pass to check continuity between two bits.

Though, I'd also agree that if you're not providing feedback between each piece, the result is gonna suck, or at least, it's not really going to be "more than the sum of its prompt".

I experimented with "introducing randomness" in the form of web search + using older LLMs like EleutherAI's GPT-{J,NeoX} to try and inject novelty into the generation, but I never really got that to work either.

Generally true, yes, but the absolutely best prose I ever got from a local model I can run on my PC (and I've used many, many of them last two years, explicitly for prose) was from a finetune of Qwen-3.6-27b from July 2026.
> generative art is blindingly obvious.... To me

Have you tested this systematically, or is it possible that you are experiencing survivorship bias? If there were any generative art pieces that you didn't notice, you would have thought that they were human-made. Therefore, all the pieces you identified were "obvious" to you. Not to mention false positives.

I imagine that they haven't done a peer-reviewed scientific study of their ability to notice generative art.

We can replace their statement by "I'm a visual artist and at generative art is often blindingly obvious.... To me....". In other words, they can notice some or most generative art but other people familiar with art can't, or at least can't as often as they do. I think what they are trying to say isn't too changed by that.

Even if it’s obvious to some at this time, as time passes artificially generated writing will become mainstream and what becomes the normal writing style. If you read IX century literature, it reads differently, the vicsb is different, the sentences are built differently. It was totally normal back then. It’s not normal now. In ten years the new artificial style will become commonplace and accepted -I mean unless we artificially always make it deviate.
It extends well beyond creative stuff too. It makes me think of Gell-Man amnesia [1].

People can tell when something they are experts in is being done poorly, but others can't, and it frustrates the experts. Then, those same people think something completely different is being done well despite what experts in that respective field say. It's like when someone from your family reads an article about your profession and then proceeds to tell you how your job works. lol

It's a really pervasive issue in society IMO. And a few prompts in someone's favorite LLM just reinforces it to people that don't know (any better|what they don't know).

[1] https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect

I think you are too close to the trees though to see the forest.

I have made digital art for 30 years and this all just sounds like what people use to say about digital art in general.

The main problem I see with generative art is not that you can tell it is generative. It is that most the art is shit. The same way if you gave a 1000 random people a blank canvas and paint, most the paintings would be shit too.

The counter example is there is a billboard that I see driving sometimes that is obviously AI generated graphics. It is so eye catching compared to any of the other billboards because most billboards are boring.

You are just puppeting the standard gate keeping bullshit to a new art form and personally I sick of reading this.

Who the fuck are you to say what art is or what art can be?

I don't think there's any chance people will be employed to sell the output of their brains in 20 years, outside of "athletes", or whatever we end up calling people who think for sport, even though machines will outdo them in all practical uses.
There will always be creative superstars.
Of course. But they're going to be machines. I don't think there's magic that makes the human brain impossible to automate. And engineering is going to lead to faster improvements than evolution could.
lol yeah sure. AI sucks and always will.
People will say that until it doesn't. We're currently in the largest research investment in history for this, and have made a crazy amount of progress since even five years ago.

I'm not cocky enough to bet against it. Especially since now AI is solving frontier math problems, debugging better than humans, and writing most of the content posted to Hacker News.

You're ignoring a much more mundane problem.

What makes you think that there are economic incentives to build an AI that has the same capabilities as the human brain?

The current scaling story is to intentionally take a system that objectively doesn't and sell it as if it does.

What makes you think that offloading expertise and making fungible whatever human operators remain necessary has no economic incentives?
And why would we accept this passively? For one, I will be in the hills with my rebel "compañeros" should things get that bad.
If that's how you feel, you'd better get started -- we've already got AI writing most of the posts on this site, handling frontier maths, writing most of the commercial code at most startups, and its capabilities are growing real fast.
This feels premature. The thing about the LLMs people use today is that it's a handful of super expensively trained models serving 1000s of use cases ranging from frontier math to recipe planning.

We get to learn the foibles and language of Claude and ChatGPT as a result. The slop is almost detectable if you provide no steering prompts about story structure, narrative structure, or stylistic cues. And most writers are not finetuning the weights to their LLMs explicitly.

If you invest time into doing all that (not really trivial stuff), the results will be better.

The safest jobs are the ones that don't feed a training algorithm with data, any work that remains more of a mystery.
I don't think writing is an "AI-complete" problem. It's just that the models are inefficient at it right now, due to training and/or architecture. Good writing requires thinking, reflection, and doing multiple passes. This currently translates to using a high reasoning effort and burning lots of tokens, but big AI labs like Anthropic are struggling with handling the load, so they're doing the exact opposite and finding whatever cheap trick they can to reduce token usage. One such trick is condensing ideas into as few tokens as possible, which in my opinion is one of the reasons Claude sucks so much at writing.
There may well be some writing jobs that are safe for the reasons given in this post, but that doesn't help all the writers I know who have already lost their jobs and are struggling to find work.

It's not enough that humans can tell the difference and feel an ick, there also need to be enough organizations willing to pay money for that difference. From my vantage point, there are not. It turns out that for a ton of the writing produced by companies, the quality of the prose wasn't really "load-bearing" as Claude puts it. That writing is there to occupy a space and look professional at a glance, the same way elevator music is tolerable for the duration of an elevator ride.

I have worked in the publishing industry, translation and editing (got out to a more dependable career just in time), and what is astonishing is how even "load-bearing" text is not safe from the cost-cutting pressures. That is, even some respectable publishers for the last several years will publish a book one is expected to pay good money for, with minimal editing and proofreading compared to the old days, and if it is an originally foreign book, with machine translation and minimal post-editing.
I've got some heavy reader friends and they keep complaining that many of Spains's fantasy book translations are ridiculously bad and blatantly machine translated now: Senseless terms that require looking up the original to understand, e.g. "fall" translated as in falling when it meant autumn; proper nouns, including character names, sometimes translated sometimes not; and many sentences that mix up the grammatical gender because the machine wasn't given enough context to infer it.

But these publishers don't get punished because they keep licensing popular foreign franchises, which aren't quite fungible, so consumers don't want to miss out and thus don't vote with their wallets.

They DO vote with their wallets, though. They're still buying that crap, don't they?
Unfortunately the answer is no. Reading or just owning a book by a certain author can be a matter of fashion or prestige. That allows for really bad work still be published.
If it's just for prestige, they should buy the original (English?) version and get even more prestige.
Proficiency in English in Spain is still rather low, and someone carrying around the original could be challenged to show they actually understand it.
Sometimes the drop in sales is less than the cost savings
Since you can't really go buy the same work translated by a different publisher, voting with your wallet means boycotting it. I agree it would be the right option to punish publishers, but people find it frustrating to miss out on both the work they were hyped for and its fandom; the latter is a really big deal for those deep into reading subcultures.
I know, but really, in the 21st century, aren't there many substitutes for whatever pop-hype we're talking about? Not enough worlds with wizards, or swords, or dragons, or whatever you like?
How is it possible? Even if you just put text to ChatGPT it'll translate it much better than that. And with actually decent harness, the translation should be quite good. Do they even save on tokens by choosing cheapest model or something?
It's impressive how bad it is even for machine translation, I've seen examples. You're imagining them using a good LLM and a decent harness, but I'd bet it looks more like an intern copy-pasting through Google Translate or some equivalently bad workflow.
I appreciate what this piece is trying to do directionally but it shares a logical flaw with lots of other pieces about AI and knowledge work. It assumes that in order to disrupt a particular field, AI needs to print serviceable work unassisted; that it's "vibe-shipping".

AI today is most effective when it's not vibing, but rather copiloting a skilled operator. I don't necessarily want my agent to build an entire system for me, even if it's ultimately the author of almost every line of code; I'm actively making decisions throughout. There's obviously a spectrum here but at most points on the spectrum the amount of assistance available is still a step change in the economics.

So too with writing.

The first rule of accelerating writing with AI is that you're not allowed to use a single word the AI suggests. Even if what the model comes up with is great, better than what you could have done, as soon as the AI suggests it it's poisoned. At least with current models, readers can detect LLM prose in the parts per trillion, and as soon as they do you've lost them.

The second rule of writing with AI is that AI encouragement is toxic. A structural consequence of RL is that models are exquisitely tuned to generate responses that make their users perceive value. We recognize this in a gross sense in "sycophancy", but the problem recurs fractally in at finer-grained levels, where stuff like "this part is really strong" will subtly allow the model to set a course for your writing and you'll confidently ship crap.

With those two rules in mind, models are incredibly valuable for writing, more valuable in my experience than the professional copywriters I've worked with. The trick is to get them to make suggestions at a higher level than just writing alternatives:

* Do the sentences in these paragraphs end with the new idea or information?

* Are the real actors in each sentence the grammatical subjects?

* From paragraph to paragraph is there a clear flow of topics, or are things jumping around?

* Is this piece crudded up with metadiscourse like "it's important to note"?

I've had a stack of notecards for ages that I took down from Joseph Williams "Style: Towards Clarity And Grace", the most programmer-brained writing book ever written, I love it very much. For the past year or so I've been feeding them through GPT and Claude one by one, and it's drastically increased the speed at which I can knock out a completed piece.

I think it's pretty hard to argue that AI isn't going to have an impact on the writing profession. It's just not the most obvious impact everyone assumes it will have, where it, like, writes whole op-eds or whatever. At least not yet.

Mmm, I doubt it.

I have a weird background (product, development, writing + devrel). I write a lot of code and a lot of articles.

I think there is a lot of overlap in how people who write code or articles (documentation, books, etc.) use AI. On one side of the spectrum, you have people who just blindly input some prompt, accept the output and move on with their lives. You can likely predict how that is going for them (not great). On the other side of the spectrum, you have people who outright reject all AI and are continuing to plod on with how they have always done things.

In the center is a more reasonable approach that leverages AI to create without blindly accepting the output. This applies very much to writing.

The workflow that I've adopted over the past two years or so has been to leverage AI to help with the research and outline process. Once I'm happy with the structure I go and I write what I need to write.

This maps pretty closely to the code that I write. It's fine.

Yeah I had GPT5.6 Sol act as an editor on my new chapter and it did find decent issues with PoV and narrative distance but it was so bad at prose. The best way I can describe it is, it was too clean. Often times humans write about things unsaid and left for reader to interpret or deliberate awkward sentences to tease character psychology. The AI just straight up corrects these without second thought. This will likely be the case with general purpose models unless we see some genre specific fine tunes.
> This will likely be the case with general purpose models unless we see some genre specific fine tunes.

Are we at the point in the hype cycle to go "there's a skill for that" yet?

A lot of the replies are insisting AI will get better at writing with more development but I don't see it. Even if you have a mathematically perfect writing AI you still run into the same problems you would have if you handed off your writing task to someone on fiverr or something. It can't magically know what you want to say, it only has the information you gave it. A prompt complex enough where it won't get any wrong ideas has to contain as much information as the output would have.. so just write it.
One of the biggest lessons in life to learn is there are no shortcuts.

Doesn’t stop people trying.

> It can't magically know what you want to say

I think for this argument to be true, the axiom that supports it is that the models have just as much context as they will ever have, and you cannot see being able to give them more / enough to be able to understand your perspective. That feels unlikely to be a position that doesn't change. As a society we're giving more and more context each day to this, and that makes this a valid opinion now, but one that erodes over time.

Context went from 8,192 tokens on GPT 4 to 1M tokens currently with zero improvement. The latest models got even worse.
Size of context is not the entire story here, it's ability to properly feed and index the context that's needed on this sort of thing. E.g. your entire slack/discord/email/github/jira/zoom meeting/coffee chat ... history is the context that you bring to the table on this sort of thing. Most of this is unindexed. Much of this will not be in the future.

> The latest models got even worse.

Which models? This is one of those things that likely has both model and domain specific aspects that impact your experience. In my experience with OpenaAI models predominantly (I previously worked there), they've improved significantly over the last 6-12 months. My experience with Claude is worse, but I haven't spent as much time getting into a mechanical sympathy there. They're still not perfect though and I have many steering docs that help avoid the biggest problems in the models I use when generating docs.

(comment deleted)
Claude writing quality got unbelievably bad with Opus 4.7, with no improvement in Fable. Opus 4.6 was fine. Im starting to see it as a security risk - my brain just can't process its word vomit, so just tell it to go on, implement whatever
its not about dumping more and more info into the context, its about the intention. whats not in the context is just as important as what is. and i dont see how that can be automated.

also were seeing models become worse at writing as they get smarter.

I also believe this. Post-training LLMs with vague metrics can only be achieved with RLHF, which is not impossible, but extremely costly and difficult. Instead, companies will opt for RLVR, focusing on math and programming tasks. This pushes objectives away from writing quality; often far away. That is why older models, in my view, actually read better than newer ones. It's by design.
You can brute force it by making it try random stuff then judge itself on it. You don't have to always use an LLM's output. Sometimes you can use that plus other things to add flavor. An LLM is actually really good at judging if something is good or bad. It just has a really hard time coming up with new things. But if you had unlimited compute you can throw in some rng and whimsy and get something resembling what humans do.
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly worked well. The hope was that if RLVR was used quite a bit, and the general performance crossed some threshold, that it would then be possible. However, since it has been shown that RLVR only concentrates the distribution of outputs rather than truly shift it, I doubt this will ever be a viable strategy.
> A prompt complex enough where it won't get any wrong ideas has to contain as much information as the output would have.. so just write it.

This is only if the output is fully compressed. Writing is not just about encoding the writer's ideas but also about how the reader will ingest those ideas. The writer needs to consider when to put in rests in between complex ideas to help the reader flow through the text. This suggests the LLM could be prompted by a dense complex idea to be presented with the boilerplate needed for the human mind read smoothly and without unnecessary effort.

I recently have been trying out frontier models on writing. Just for laughs, to bring an idea I have into the world acting more as director than author. I have a multi agent setup and am getting different models to argue about rating the story for clarity and whether the plot is complete.

I agree with the article. Even the frontier models call out that all the characters tend to sound the same. It also started at some point making huge changes to the core premise, and also adding characters willy nilly. The issues it called out with the plot (the ones it actually consulted me on) also made me realize how terrible of a writer I actually am.

Overall, it has been an interesting experience. I look forward to reading my own book!

I understand a lot better now why people are bemoaning KDP being filled with absolute garbage AI slop.

I was thinking about this yesterday, while going through some slop documentation generated by deepseek. If you compare claude fable 5, side by side with the deepseek, the differences in prose are glaring. It's not even close. Deepseek prose reads like fragmented shorthand, where claude fable 5 comes fairly close to human, certainly not superior to human quality writing.

I watched a podcast with a cognitive scientist and one of main contributors to the theory of linguistic relativity, Lera Boroditsky.

She said something to the effect that, "in this very moment, we are speaking in ways that were never spoken before. We are saying things that no other person has said before...."

Language models are not sample efficient and cannot adapt to evolving language, unless it's documented in large amounts of examples.

So whatever isn't documented, whatever isn't in the training dataset or the rag corpus, the model will always be incredibly different in expression from humans.

> ... and model checkers can instantly catch errors.

Lol. If that were true software would've been a lot better historically... Model checkers don't scale to 90% of the software we write. Typically you have to 1) heaviy abstract the program and 2) put it in some sort of harness to specify how you want to model the outside world (which will always fall short of practice). Not saying they're not tremendously useful, but that bullet doesn't hold up at all.

(comment deleted)
Even safer jobs might be the ones don't get automated because nobody really cares about, e.g. fishing lobster on a boat