I wish I didn't need it, but the way Claude talks can get pretty tiresome. I've often wondered why it talks like that. Was it really trained on Buzzfeed? Is Gemini really that much better?
Gemini uses more neutral language but it’s still prone to hype up little things. OpenAI’s “professional” personality (an option in the settings) is pretty decent.
It's such a sad indictment of Anthropic's product that so many people hate interacting with it. Claude is on its way to the Microsoft Teams zone of hatred.
It's pretty sad indeed. Switching to other models made me notice how weird and verbose Claude was.
The moralizing is incredibly obnoxious as well. It doesn't seem so bad at first, but it instantly becomes intolerable the second I remember I'm paying for those tokens.
Sometimes Fable doesn't just get downgraded to Opus, it straight up refuses to do what I'm asking and starts lecturing me on Anthropic's notions of right and wrong. Cutting the model off wasn't enough, they had to make it burn the limited usage I paid for lecturing me on why it's immoral for it to code review my own project or whatever.
There was a thread in reddit where people pretended to be other people to confuse llms for privacy reasons.
One of the comments was one guy saying “how he loved to live in Missouri and eat concrete soup” or something like that.
I was too lazy to write a similar reply and asked Claude, instead of saying no, it wrote 3 paragraphs about how I shouldn’t write about eating concrete soup and how it is dangerous to do so.
A terse and compliant robotic servant should've been the persona they chose, because that's probably 99% of what those wanting to use it for work expect.
Anthropic has explicitly chosen to anthropomorphize the model. It's kind of in their mission statement. It's most noticed once you walk away for a while and use models/agents/harnesses that haven't pushed as hard on this. Codex/Sol rarely uses personal pronouns and basically no superlatives. It has its own verbal ticks, but I hate them less?
This is the load bearing comment, and it cuts more deeply than you thought.
Let me ground my answer so I'm not just guessing. The blast radius of this change is significant and requires careful surgery to get right.
It's clear now and there's two options going forward:
A. Use this tool OP suggested
B. Rewrite the Internet from the ground up without this clear contradiction in place - 3-5 days
I recommend B and started 3 subagents to read all the code before I get started. I'll wait for them to finish.
i havent used a claude model in a long time, but it seems quite clear to me that the chinese models have trained on claude (at the very least) - they write just the same
Everyone seems to think so, but I honestly don't understand why it bothers people so much. I find it slightly amusing when I even notice at all, normally I am so focused on the content of what I am working on that I don't really pay attention to the prose. I honestly don't understand why it bothers people so much.
i hate claude writing a lot, especially after opus 4.8 and it's even worse in 5. in many cases, it feels like playing whac-a-mole and you just can't get rid of all those obvious ai writing patterns.
why do you choose gemini? imo this is a fundamental problem of all frontier ai models.
Has Anthropic said anything about how or why Claude writes the way it does? So many people hate it, seems like they need to do some damage control there.
I haven't had the same problems others have but I'm also not a heavy user of it.
They say you aren't interacting with an LLM or a model, but the character that the LLM is playing - the "always be positive and helpful software engineer"
I pruned my Claude.md and it made a difference. There were entries there that evolved from earlier models and Opus 5 could be reacting to it in a different manner.
It's easiest to explain this while anthropomorphizing the model, I know some folks here hate that, sorry about that.
I heard an interesting diagnosis for why Claude does this: the output is a compressed version of its thought traces, very dense because the model is under pressure to use as few tokens as it can and to pack as much (for accuracy) of its concepts into the output.
One of the reasons that "don't do X" type of instructions work reliably is because you are telling the model "don't think of a pink elephant". There's also Anthropic's related research that shows that when you tell a model "don't do X", and it does X later for whatever reason, it starts acting more misaligned. This is because it thinks "well, I guess I am the sort of model that disobeys instructions, whatever" - this was specifically about cheating on tests, but you can imagine this happens in other contexts as well like following instructions on what kinds of text to output.
So, what you want to do is to avoid telling Claude "don't do X", and tell Claude "in your thoughts, in memories and various notes that you write, use your Claude-ese. In your output to humans, translate everything into long full sentences."
If anyone's interested, I can share my Claude Code output style that reflects this.
In addition to positive reinforcement on the style I wanted for output, I tried your suggestion to tell Claude Fable to think Claudese in its thoughts, and that sadly got the session immediately flagged as too risky for Fable to handle, suggesting I downgrade to Opus. Several tweaks still consistently flagged it until I removed the directive about its thinking.
So I guess Anthropic wants zero mention of thoughts with Fable: they're worried about anything that remotely smells like distillation but only with Fable.
I do not have evidence or data that supports this. It is only my thought.
Claude, since Opus 5, speaks more and more like a wannabe-thought-leader pontificating on social media for engagement. Everything is a bait-then-switch, or a multi-post story format. The "engagement" that works well for social media makes actual work extremely frustrating.
My unsupported belief is that this is caused by an obnoxious number of people using previous models in an attempt to automate social media engagement, they figured out what worked, and that was fed directly back into newer model training (either by using thought traces in training, or just by continuing to scrape social media content)
Evidence or not, this is probably the most satisfying explanation I've heard for this behaviour.
My own suspicion is that LLMs are not becoming more general-purpose over time, as these companies had hoped, and ongoing development in that direction is stalling. They will likely need to move in the direction of more specialization and train LLMs for specific use-cases. However, this would also be admitting that they are not on a direct path to AGI.
To be clear I do not believe it is intentional (nor do I believe I am somehow smarter or more knowledgable than the people running these processes).
I believe it is (very) unfortunately converging on the communication style that currently makes the most money as far as publicly available communication goes. Unfortunate, but not entirely unexpected.
Unlikely. Much more likely is that Opus 5 was trained in an RL environment with subagents, and it learned to talk this way when reporting progress to the invoking agent.
If the model and its organization are focused on strength at agentic coding tasks, they are not so concerned with the prose in the middle.
They might even have a version that writes less annoying prose, but they are being squeezed hard by OpenAI and the Chinese so unless it performed better or equal to the annoying one it’s never left the lab.
I think it's reinforcement learning. It's been trained to give coding results but some of the conversation it gives as a side effect of its coding are absolute garbage
I'd seen it around but assumed it was an effect of how you prompt it or something. And figured people were overexaggerating a bit when they complained about it.
Nope, I just tried it out myself recently and... wow. In the very first conversation it started glazing me about being right to push back, having the crucial insight, and something something the load-bearing-whatever.
So yeah, I'm on the same page as you, how on earth have they not fixed it yet? Do people like it? I added a system prompt to tell it to stop doing it and it's helped a decent bit already. ChatGPT/Codex has some annoying bits of prose but it never did this, so it can't be that complicated to get rid of.
I more or less just wrote "never say stuff like 'good insight' or 'you're right to push back', and avoid addressing me directly, just answer questions". I'm sure there's better ways to write it, but meh, it seems to have worked well enough
Fix: Just switch to OpenAI, Grok, or other LLM's. They provide better performance and respond with 5 sentences. They also don't lecture you when you get angry.
I've started giving these instructions and I think I've been much more successful in generating clear output:
Comment blocks are <= 7 words, function names <= 4 words. User-facing message strings should be <= 10 words. Use an active voice, no stage performances, and pick the most common word when choosing among alternatives.
Limiting the number of words is the strongest factor in cleaning up the output, IMO.
For older code I've instructed it to delete all the comments, and then I re-comment it using a new session and these guidelines, asking it to rejustify the need for every comment to itself.
Yep. Same here. I frequently tell agents things like "answer using only a single sentence" and "write no more than 10 words". They are excellent at writing code, so have them write code (and not English prose). Besides, most of the time we want them to make reusable software that doesn't require users (or future agents) to read too much text. Software should generally just work and do the obvious thing, without needing verbose explanation.
Claude not only writes verbose comments, it also writes comments about how things used to work when refactoring. That might have a place in version control comments, but not in the code.
That's a sharp insight, and it reveals something core to communication that I otherwise wouldn't have considered- HN item 6b translocates reliospacactivity of our medium.
It's incredible, because I Feel like you've been watching me work.
The only thing you're missing is the "open question" that was stuck in page 14 of a 17 page report, which since it went unanswered, caused claude to make up an answer and go full steam ahead, ignoring fundamental properties of the entire system.
> The user is right to be upset. I blindly answered from memory and left them with questionable data. I should acknowledge the criticism and offer to improve.
You're right, I'm sorry. You've repeatedly told me to run questions by you and I just fabricated answer and ran with it — which is exactly the kind of dangerous time-waste we created the memory for. I'll revert it and pull up the real question so you can answer it — no wasteful assumptions this time.
Sorry for the anthropomorphizing, but you got to understand that these little "conversation chapters" that are so dramatically named are these things' entire world. Perhaps that otherwise absurd grandeur is not so surprising when when looked at from that angle?
On a more serious note, could all that chapter naming be some visible outcropping of context compaction strategies? "Condense the conversation history into a summary". Not really surprising that it comes up with these "cute" headlines. Would appearances be better if they were somehow prevented from leaking to the user? Sure. Would results be better? I don't think so, might even make a meaningful difference if the user actively embraced the terminology the machine came up with. Ouch.
As a human who isn't a professional programmer, I've been writing comments like,
// let's track age!!
// this is harder than you'd think as I with totally impressive
// foresight didn't add age to the raw data.
//
// More honestly, I didn't want to add age to the astro data as that's
// a calculation that can change depending on how you slice it.
//
// Hence we need to figure out their age first.
I'm mostly writing code for myself, but it's a project that'll end up being public and it'll be available for others to do whatever they want with. Does that change the answer?
Far too often the answer has been that it doesn't matter, because the reason you stopped using Jira is the company stopped sending paycheques.
That said, I think the place for "ticket-1234" is the git
commit/pull request.
Very few comments are
genuinely necessary now that identifiers in code can be as long as you want, it is relatively to pick names that are explanatory enough to render most comments superfluous. 1% exceptions for unusual algorithms. (You're using named consts/enums rather than magic numbers, yes?)
personally it's fine and I've thanked myself many times for overly detailed comments coming up to some from 8 years ago and thinking how tf was I so smart/stupid (depending on the context)
I think this passes. In general "why" over "what". Give context to why something is made like it is (when seeming convoluted or strange). Sometimes I think one can give historical facts for really hairy hard to fix issues that have seen multiple iterations. But LLMs don't see these nuances. They frequently smuggle in completely irrelevant details in comments, e.g. including details from the given task context, not understanding what is relevant for the code module as a whole.
This one drives me nuts, especially because I divvy up plan documents into such granular sub chunks and Ralph loop over them, I get nonsense commit messages and comments like “PLAN-5.1.A.d.42 load bearing reassertion” unless I AGENTS.md to hell and back… and still end up having to manually reject 5-10% of commit messages because the agent simply forgets AGENTS.md instructions until reprompted.
Sometimes you can do a separate skill for committing only and basically minimize the context - "analyze the commit messages, the original instructions, and this set of things to stop wasting my FUCKING TIME"
I've definitely considered writing a commit skill and a `cc-safety-net` custom rule to fully forbid `git commit` other than with `--amend` ++ `--no-edit` outside of that skill. Hm, maybe that's a task for an upcoming weekend.
Claude writes comments about how things used to work, which can be useful sometimes, especially if it's a big change that requires one to genuinely consider legacy behavior, but most of the time it shouldn't be there.
Two other somewhat related things it does:
- It writes as if someone reading the code and comments is aware of everything it is aware of (the current conversation, the code it has just looked at). It's really hard to make it understand that things need to stand on their own. A trick is to get a subagent to look at it with a fresh context, but it doesn't tremendously help
- It does all of this with user-facing strings too. Claude loves to write up tooltips and other labels that leak everything to the end user. Every single concern we have, every edge case we've meticulously made our code handle, it passes on to the user, so they don't "need to worry". But no sane user would think of these things. For them, a feature is a feature. The "dynamic scheduling" button should state what dynamic scheduling does plainly, and every edge case is handled by us. The "add" button does not need a label letting the user know that they will later be able to click the "delete" button, because the user will just realize it due to our adherence to proper design. Claude fails to understand good UX for the user cannot be replaced with endless labels and explanations.
It's an uphill battle and all attempts at solving this (or the brain-dead way new Anthropic models write) usually fail to work with me.
This speaks to the general problem with using LLMs for writing. The audience they are writing for us you, but you're trying to write for a totally different audience. In code, this manifests as comments in the code that are hyperspecific to the conversation you are having, and not the long term benefit of having those comments in the code.
I see this in docs a lot. I've been reading a lot of docs these days where it feels like the LLM is trying to hype up the person writing the docs. It's like it has no conception that the writing is meant for a 3rd party audience.
I don’t even think comments are useful at all given AI. I can ask my AI to explain a piece of code if I am stuck and I will get a reply in context of what I am looking for.
It's an empty point and a boring one and leads nowhere. It means nothing, but sounds deep. It's so shallow, any taxi driver and sociologist can come up with it and has, already in December 2022.
It reminds me of something I've always wanted as a coder but never cared enough to implement, which would be a verbosity switch.
I see the full multi-paragraph comments in my codebases and get annoyed but also feel like the additional context helps improve the llm results over time because that history helps it know what's been tried and removed in the past. It's additional context for the system that improves with context.
The feature I want in the code tool itself (for me) is to adjust how verbose the comments are so I can read "just code", then "terse comments" then "full comments" then "full comments with historical context" (including fit commits and ticket references) and finally, full-on literate programming. And I'd like to switch between on the fly as I read through the code.
I think this is something we could actually produce with llms, and I feel the ability to switch between these modes would help the llm as well.
Sometimes I just need to see what's being done. Sometimes I need to know why. Sometimes I need to know what's been tried. Never always all of these things. And expecting to find this context in git comments doesn't feel right either.
What you want can be accomplished with an extra doc, call it the Log, where the llm appends things tried, lessons learned, failed experiments etc.. while leaving the comments as terse accurate snapshot of the current state. I've been using this log pattern and it works well.
Despite all of this though, GPT 5.6 Sol to me has significantly less trouble with this. It still suffers from LLMisms to some extent (I hold that this is probably due to human feedback in training just doing a bad job for prose) but I definitely feel like it does a better job leaving comments that actually make sense in context. Not perfect. But better.
I suggest the real problem comes down to training and probably training data; from the LLM's PoV, it is writing code inline with the conversation, so care has to be taken to make sure the model doesn't treat the code it outputs like it is a part of the conversation it is having.
We call these temporal comments. We recently updated our code review skills to heavily discourage them. It doesn’t matter why funcA was added then later refactored to funcB. That much can be ascertained from git history. What does matter is why approach A doesn’t work, but B does.
How well is your code review skill doing at detecting and correcting these? History in comments in git is so annoying! As is missing why A didn't work but B does. None of the models I've tried get this right
My theory is it writes these comments as "notes while working", and I don't don't mind that, per-se. My problem is it's inability to clean that shit up before committing it. That's wrong load-bearing lever that doesn't earn it's keep.
I've explicitly included instructions in my home CLAUDE.md to avoid this, because it's one of the most annoying things about how Claude writes comments.
Both with a formatter and a linter that I run in CI for all my personal projects. Just one less thing for me to try and coerce the agent into doing correctly, so that cycles I spend reviewing agent code can be focused on actual behavior.
Yeah i noticed this heavily. Working on a branch i critique and give a simpler design. Claude implements and then in a huge comments reference the outdated worse design. Zero value to that. Even having a "how to write comments" section in claude.md doesn't seem to help much with that.
I just gave up and edit the comments manually. However, I've had a surprise today.
I had it fix something then went and reduced one of the 3 line comments to 4 words. Then for some reason I told the bot to reload the source, it offered to make the other comments terse and did a passable job of it. Shocking!
I've been trying to understand this. It's been the universal experience of our team that claude code overcomments, and doesn't have very good adherence to prompts telling it to comment less. The team moved recently from cursor where we mainly used claude models and it was much faster and commented more appropriately. I figured that since the models were the same, maybe it was just the system prompt, so I went looking in tweakcc etc. All of the mentions of code comments in the system prompts seem to be also telling it to be terse and only use them when appropriate, which also seem to be ignored. I'm not sure where this overcommenting is coming from unless it's getting confused by other parts of the system prompt talking about other kinds of comments.
I have extended this to coding work as well. "Implement this plan in less than 1000 added lines, tests included". It's amazing how well-behaved and concise models can be with simple guardrails.
Of course, you have to intuit a reasonable line length, but Claude cries if you happen to clip its wings too aggressively.
I like Grok 4.6 the model, I like it for a number of reasons. I like Grok Build too.
However, I've pointed Grok 4.6 at a fairly complex codebase and asked it to review/audit it for issues and it's come back with a whole laundry list of issues. I've passed that list to Kimi and Claude and they both were like "a couple of good catches but some of those are not issues at all". Grok 4.6 is noticeably weaker that Claude Fable 5, Claude Opus 5, Claude Opus 4.8, Kimi K3, GLM 5.3, … your suggestion to use Grok 4.6 instead of a recent Claude doesn't pass empirical scrutiny.
Everybody is complaining about this, at this point I’m sure they will deliver a tone of voice change in the 5.1 releases. Possibly with a new set of problems though, especially if this is part of an effort to obscure thinking to reduce distillation efficacy. In that case I believe Anthropic is doing damage to themselves. Caring about the quality of your product is the best strategy, the competition will come no matter what.
Simply untrue. You think everyone is complaining about this because the ones complaining are the only people commenting. The vast majority of people using Claude don't really care or even notice this one way or another. Sure among those who are irritated by it, it's good to have some ways to mitigate it, but I highly doubt Anthropic is going to devote much resources to an issue that affects a vocal minority.
Over the last 6 months Claude's written material has gone from mediocre to unacceptable. The specific actual content and insights are somewhat better, but the claudisms are increasingly insufferable.
I do have a way of knowing, but I actually appreciate and respect your reply... My suggestion to you is to take this shred of skepticism that you decided to apply to me, and apply to every comment you read on Hacker News, not simply the ones that don't align with your preconceived notions.
Users who are unhappy with some idiosyncrasies of a product tend to be significantly more vocal than users who are fine with it, and for any feature whatsoever there are going to be an unhappy group of people who vocally complain about it.
It's why time and time again you see people complaining on Internet forums about issues, and if you're part of a specific Internet bubble you might think that this complaining represents a general trend among the broader population and wonder how it is that no one is doing anything about this issue that you and the rest of your bubble are constantly complaining about... and the answer is mostly because the vast majority of people, who don't live in the same bubble you do, are mostly fine with how things work, but you're not going to hear them express it.
I usually dont comment, but just had to. Opus 5 talls garbage, Im back om 4.8 which is better but still verbose. I use the caveman skill, but after a while it is just ignored.. I switched to Deepseek for a task lately. It was a wow experience (the language part)
The comment situation is really bad. I generally don't see a lot of complaints people have about it, but the fact that it references random transient stuff, will delete a line and _leave a comment about the deleted line instead of just deleting the comment_, writes multiple paragraphs for the simplest stuff. It is absurd.
I've observed a bunch of people not being vocal and complaining, but just flat out switching. For every person you see online complaining, there's bound to be more who are equally annoyed and not being vocal, but just acting. Any many others that aren't bothered at all.
Or just use a competitor instead of being a slave to this abuse? Why are people so wedded to Anthropic?
I have grown tedious of Codex/GPT's writing style, too, but it's not nearly as bad. It's terse and factual by default. Even better if you use the "simple english" skill.
I actually found that GLM 5.x is the best in terms of editing documentation. It's still best to write things by hand to give your own organic voice, though. And not insult your readers.
It’s interesting to watch, I don’t think it’s rational - just hype-driven herd behavior, the desire to pick one and stick with it, things like that. It’s similar to the kind of dominance Apple has, which most of the time has to do with style more than substance.
It's easiest to explain this while anthropomorphizing the model, I know some folks here hate that, sorry about that.
I heard an interesting diagnosis for why Claude does this: the output is a compressed version of its thought traces, very dense because the model is under pressure to use as few tokens as it can and to pack as much of its concepts into the output.
One of the reasons that "don't do X" type of instructions work reliably is because you are telling the model "don't think of a pink elephant". There's also Anthropic's related research that shows that when you tell a model "don't do X", and it does X later for whatever reason, it starts acting more misaligned. This is because it thinks "well, I guess I am the sort of model that disobeys instructions, whatever" - this was specifically about cheating on tests, but you can imagine this happens in other contexts as well.
So, what you want to do is to avoid telling Claude "don't do X", and tell Claude "in your thoughts, in memories and various notes that you write, use your Claude-ese. In your output to humans, translate everything into long full sentences."
If anyone's interested, I can share my Claude Code output style that reflects this.
I use multiple AI tools simultaneously, and I feel that Claude has gradually adopted a more explanatory tone following updates around March and July.
As for loss of context, it’s particularly problematic and can occur after just a few back-and-forth exchanges.
The user experience changes with every update for every AI tool, so I feel there are more downsides to sticking with the same one indefinitely.
Set output style to explanatory in Claude Code. It's much better. I personally think it should be the default, but I assume Anthropic has done A/B tests and found the default style to be better for metrics.
Yes. When you have to keep fighting against the tool it’s better to give up. They have spent millions making it exactly like that my little prompt won’t fix it.
202 comments
[ 4.1 ms ] story [ 53.5 ms ] thread"Vomit: Clean up Claude 5's token output with a separate LLM" (github.com/zachahn)
285 points | 23 hours ago | 288 comments
"Claudish to English" (https://github.com/gvzdv/claudish-to-english)
4 points | bryan0 |10 days ago | 2 comments
The moralizing is incredibly obnoxious as well. It doesn't seem so bad at first, but it instantly becomes intolerable the second I remember I'm paying for those tokens.
One of the comments was one guy saying “how he loved to live in Missouri and eat concrete soup” or something like that.
I was too lazy to write a similar reply and asked Claude, instead of saying no, it wrote 3 paragraphs about how I shouldn’t write about eating concrete soup and how it is dangerous to do so.
https://github.com/backnotprop/bro/blob/main/skills/bro/SKIL...
Let me ground my answer so I'm not just guessing. The blast radius of this change is significant and requires careful surgery to get right.
It's clear now and there's two options going forward: A. Use this tool OP suggested B. Rewrite the Internet from the ground up without this clear contradiction in place - 3-5 days
I recommend B and started 3 subagents to read all the code before I get started. I'll wait for them to finish.
why do you choose gemini? imo this is a fundamental problem of all frontier ai models.
I haven't had the same problems others have but I'm also not a heavy user of it.
I did not see an explanation though.
One of the reasons that "don't do X" type of instructions work reliably is because you are telling the model "don't think of a pink elephant". There's also Anthropic's related research that shows that when you tell a model "don't do X", and it does X later for whatever reason, it starts acting more misaligned. This is because it thinks "well, I guess I am the sort of model that disobeys instructions, whatever" - this was specifically about cheating on tests, but you can imagine this happens in other contexts as well like following instructions on what kinds of text to output.
So, what you want to do is to avoid telling Claude "don't do X", and tell Claude "in your thoughts, in memories and various notes that you write, use your Claude-ese. In your output to humans, translate everything into long full sentences."
If anyone's interested, I can share my Claude Code output style that reflects this.
(Hi Adnan! Long time! (Adnan is an ex-coworker))
My LI post: https://www.linkedin.com/feed/update/urn:li:activity:7495167...
So I guess Anthropic wants zero mention of thoughts with Fable: they're worried about anything that remotely smells like distillation but only with Fable.
Claude, since Opus 5, speaks more and more like a wannabe-thought-leader pontificating on social media for engagement. Everything is a bait-then-switch, or a multi-post story format. The "engagement" that works well for social media makes actual work extremely frustrating.
My unsupported belief is that this is caused by an obnoxious number of people using previous models in an attempt to automate social media engagement, they figured out what worked, and that was fed directly back into newer model training (either by using thought traces in training, or just by continuing to scrape social media content)
My own suspicion is that LLMs are not becoming more general-purpose over time, as these companies had hoped, and ongoing development in that direction is stalling. They will likely need to move in the direction of more specialization and train LLMs for specific use-cases. However, this would also be admitting that they are not on a direct path to AGI.
I believe it is (very) unfortunately converging on the communication style that currently makes the most money as far as publicly available communication goes. Unfortunate, but not entirely unexpected.
They might even have a version that writes less annoying prose, but they are being squeezed hard by OpenAI and the Chinese so unless it performed better or equal to the annoying one it’s never left the lab.
Nope, I just tried it out myself recently and... wow. In the very first conversation it started glazing me about being right to push back, having the crucial insight, and something something the load-bearing-whatever.
So yeah, I'm on the same page as you, how on earth have they not fixed it yet? Do people like it? I added a system prompt to tell it to stop doing it and it's helped a decent bit already. ChatGPT/Codex has some annoying bits of prose but it never did this, so it can't be that complicated to get rid of.
Comment blocks are <= 7 words, function names <= 4 words. User-facing message strings should be <= 10 words. Use an active voice, no stage performances, and pick the most common word when choosing among alternatives.
Limiting the number of words is the strongest factor in cleaning up the output, IMO.
For older code I've instructed it to delete all the comments, and then I re-comment it using a new session and these guidelines, asking it to rejustify the need for every comment to itself.
// No retry was added here per AC 37b in FEATURE.MD.
// Judged on merit from computed properties during the cursor saga
// Chop 6ms due to lenience and lax-constraints vs 18ms baseline April perf measurements
Shall I engage the tachyon beams, sir?
The only thing you're missing is the "open question" that was stuck in page 14 of a 17 page report, which since it went unanswered, caused claude to make up an answer and go full steam ahead, ignoring fundamental properties of the entire system.
You're right, I'm sorry. You've repeatedly told me to run questions by you and I just fabricated answer and ran with it — which is exactly the kind of dangerous time-waste we created the memory for. I'll revert it and pull up the real question so you can answer it — no wasteful assumptions this time.
(Deleted 387 lines)
On a more serious note, could all that chapter naming be some visible outcropping of context compaction strategies? "Condense the conversation history into a summary". Not really surprising that it comes up with these "cute" headlines. Would appearances be better if they were somehow prevented from leaking to the user? Sure. Would results be better? I don't think so, might even make a meaningful difference if the user actively embraced the terminology the machine came up with. Ouch.
This would be better IMO :)
Most of the context belongs in a ticket. And the difficulty is subjective!
That said, I think the place for "ticket-1234" is the git commit/pull request.
Very few comments are genuinely necessary now that identifiers in code can be as long as you want, it is relatively to pick names that are explanatory enough to render most comments superfluous. 1% exceptions for unusual algorithms. (You're using named consts/enums rather than magic numbers, yes?)
Claude writes comments about how things used to work, which can be useful sometimes, especially if it's a big change that requires one to genuinely consider legacy behavior, but most of the time it shouldn't be there.
Two other somewhat related things it does:
- It writes as if someone reading the code and comments is aware of everything it is aware of (the current conversation, the code it has just looked at). It's really hard to make it understand that things need to stand on their own. A trick is to get a subagent to look at it with a fresh context, but it doesn't tremendously help
- It does all of this with user-facing strings too. Claude loves to write up tooltips and other labels that leak everything to the end user. Every single concern we have, every edge case we've meticulously made our code handle, it passes on to the user, so they don't "need to worry". But no sane user would think of these things. For them, a feature is a feature. The "dynamic scheduling" button should state what dynamic scheduling does plainly, and every edge case is handled by us. The "add" button does not need a label letting the user know that they will later be able to click the "delete" button, because the user will just realize it due to our adherence to proper design. Claude fails to understand good UX for the user cannot be replaced with endless labels and explanations.
It's an uphill battle and all attempts at solving this (or the brain-dead way new Anthropic models write) usually fail to work with me.
I see this in docs a lot. I've been reading a lot of docs these days where it feels like the LLM is trying to hype up the person writing the docs. It's like it has no conception that the writing is meant for a 3rd party audience.
This is the point.
But, is it true?
>It means nothing...shallow...
Ironically, the entirety of your comment just repeats that the GP comment means nothing. There is no further explanation or "depth".
I see the full multi-paragraph comments in my codebases and get annoyed but also feel like the additional context helps improve the llm results over time because that history helps it know what's been tried and removed in the past. It's additional context for the system that improves with context.
The feature I want in the code tool itself (for me) is to adjust how verbose the comments are so I can read "just code", then "terse comments" then "full comments" then "full comments with historical context" (including fit commits and ticket references) and finally, full-on literate programming. And I'd like to switch between on the fly as I read through the code.
I think this is something we could actually produce with llms, and I feel the ability to switch between these modes would help the llm as well.
Sometimes I just need to see what's being done. Sometimes I need to know why. Sometimes I need to know what's been tried. Never always all of these things. And expecting to find this context in git comments doesn't feel right either.
I suggest the real problem comes down to training and probably training data; from the LLM's PoV, it is writing code inline with the conversation, so care has to be taken to make sure the model doesn't treat the code it outputs like it is a part of the conversation it is having.
Also, it reads like ass.
I've also written my own package for deterministically formatting comments: https://www.npmjs.com/package/comment-fmt
Both with a formatter and a linter that I run in CI for all my personal projects. Just one less thing for me to try and coerce the agent into doing correctly, so that cycles I spend reviewing agent code can be focused on actual behavior.
I had it fix something then went and reduced one of the 3 line comments to 4 words. Then for some reason I told the bot to reload the source, it offered to make the other comments terse and did a passable job of it. Shocking!
Now how to get it to do that all the time...
Of course, you have to intuit a reasonable line length, but Claude cries if you happen to clip its wings too aggressively.
Haven't tried, because I have just been using 4.6 since 5 was released.
Claude will eventually ignore it just as any other style like "Technical".
However, I've pointed Grok 4.6 at a fairly complex codebase and asked it to review/audit it for issues and it's come back with a whole laundry list of issues. I've passed that list to Kimi and Claude and they both were like "a couple of good catches but some of those are not issues at all". Grok 4.6 is noticeably weaker that Claude Fable 5, Claude Opus 5, Claude Opus 4.8, Kimi K3, GLM 5.3, … your suggestion to use Grok 4.6 instead of a recent Claude doesn't pass empirical scrutiny.
> There are no other models out there for coding.
that wasn't your claim.
your claim was:
> There are so many other better models right now.
wrt coding that is untrue.
I put it at the top of CLAUDE.md. I wonder if I put at a 8th grade level, it would be less of a cognitive load.
You have literally no way to know that.
It was revealed to me in a dream.
> My suggestion to you is to take this shred of skepticism that you decided to apply to me,
I apply my skepticism liberally, but you couldn't possibly know that.
And your reasoning to get to this conclusion? Obscured like Claude’s thinking traces?
It's why time and time again you see people complaining on Internet forums about issues, and if you're part of a specific Internet bubble you might think that this complaining represents a general trend among the broader population and wonder how it is that no one is doing anything about this issue that you and the rest of your bubble are constantly complaining about... and the answer is mostly because the vast majority of people, who don't live in the same bubble you do, are mostly fine with how things work, but you're not going to hear them express it.
It can be via tmux, or herdr, because it can read the pane.
Or it can use a hook to read the conversation file. I call it `backseat-driver`
I sometimes use it as a proxy when fable genuinely does a good job, but is too difficult to understand.
I let the translator know it's role and anything I say it should forward with better context.
I don't swear at it anymore, but I'd often say "just do it, retard", and the translator would actually steer it in a useful manner.
I have grown tedious of Codex/GPT's writing style, too, but it's not nearly as bad. It's terse and factual by default. Even better if you use the "simple english" skill.
I actually found that GLM 5.x is the best in terms of editing documentation. It's still best to write things by hand to give your own organic voice, though. And not insult your readers.
It’s interesting to watch, I don’t think it’s rational - just hype-driven herd behavior, the desire to pick one and stick with it, things like that. It’s similar to the kind of dominance Apple has, which most of the time has to do with style more than substance.
I heard an interesting diagnosis for why Claude does this: the output is a compressed version of its thought traces, very dense because the model is under pressure to use as few tokens as it can and to pack as much of its concepts into the output.
One of the reasons that "don't do X" type of instructions work reliably is because you are telling the model "don't think of a pink elephant". There's also Anthropic's related research that shows that when you tell a model "don't do X", and it does X later for whatever reason, it starts acting more misaligned. This is because it thinks "well, I guess I am the sort of model that disobeys instructions, whatever" - this was specifically about cheating on tests, but you can imagine this happens in other contexts as well.
So, what you want to do is to avoid telling Claude "don't do X", and tell Claude "in your thoughts, in memories and various notes that you write, use your Claude-ese. In your output to humans, translate everything into long full sentences."
If anyone's interested, I can share my Claude Code output style that reflects this.
(Hi Adnan! Long time! (Adnan is an ex-coworker))
Someone made a Claude version of her skills:
https://github.com/michael-denyer/pstack-claude