Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.
A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.
What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.
My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.
I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.
This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.
Yeah, it's way better when you do design documentation, or even ticketing, to instruct it not to include any implementation specifics. You're not doing the deep dive on the zero shot that writes the ticket or the document and so it is much less informed than the agent doing the work will be.
- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).
- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).
I'm genuinely curious as to what you're working on that you find 4B models good enough. I wouldn't even let a 27B model code, never mind supervising smaller models.
At the moment I'm still doing shakedowns, so Typescript games compilation with a menu that has 4 games and retro artwork.
This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.
The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.
The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)
The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.
I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.
I use Claude's Plan mode daily and it's great. I almost always have feedback to refine the plan and I want a clean separation of planning time before writing code. I don't get what the problem is.
I also find this "kill Plan mode" push on Twitter odd, because developers have been complaining about AI supposedly killing their jobs, yet they want to take away the main feature that lets them be an active collaborator and participator in the process. Weird.
Yeah, honestly, I'm wondering, not even why people are against plan-mode, but how they're getting stuff done without it.
Maybe it's me (it usually is), but I don't give an LLM small tasks that a human could knock out in a half-day. I give them big tasks, stuff that would take a human a few months to a year - and then I ask it to give me a plan, as a living document, and we spend a good hour or two iterating over the plan.
Then, when its finally at the point where I'm happy with the plan, and I've talked it down from wherever it first wanted to go, or pointed out that we don't need to be all-things-to-all-(wo)men and focus is important, I get it to start going through it.
Likewise, I get it to maintain TODO.md with lists of known bugs, separate lists of future-features, and again, I make a rule that this file must be updated whenever something material changes. I just asked Claude "where are we ?" and got back stuff like:
The ## Still-open detail section lists two items:
- Task #1070: the ported back end doesn't fold offsets into vector loads, so it emits an extra add on 11 files (from 892) at -O3. The output is correct, just longer.
- Task #1080: array sizes must constant-fold. For example, u8 buf[EVSZ * MAXEV] is rejected because size expressions only accept a literal.
This is for 'xc' [1] - a compiler for an Objective-C-like language (but without the excessive []). The language has ARC, blocks and bound-functions/methods named 'block' and 'callback', automatic parsing of DWARF data so you can #use a shared-object, so there's no header files - just read enums/types/functions/methods from the shared object. It's a cross-compiler, runs on mac,windows,linux and creates executables for mac,windows,linux,ios,android,WASM (amongst others). I have a binary running on my iPhone which was written on, and signed on a Linux box - no Apple software used at all. Oh, and it produces code that is very comparable to clang in speed on both arm64 and x86_64.
You can appreciate it's a reasonably large project. It's taken actual months(!) [grin] for me to get working. Months! There's no way I'd approach a problem like this without detailed plans of what I wanted the language to do, where we were going with it
FWIW, "I" wrote blewit.net [2] entirely in xc - both the server back-end (#use <psql> was very useful for binding to Postgres) and the WASM client - which share classes between client and back end, to make it very difficult to get out-of-step between them. No Apache (#use <tls>), no scripting, just a lean-and-mean daemon talking to postgres via valkey (#use <valkey>) - a reddis-alike. Oh yeah, blewit.net has a plan too. Actually it has lots of planning :)
I have never used plan mode and always worked by planning. I just say 'don't write any code yet. I want to [describe usage]. Ask me the top ten relevant questions to doing this well. We're working up a planning .md we'll handoff to a fresh agent.'
I likewise don't understand this push. Surely it's a trivial feature to keep enabled. It mostly seems like advertising that "our models are so smart you don't need to plan". But I find it to be a great mental review process for my workflow.
It sounds like the Ant perspective is "you can just ask claude to plan", but at the same time that's a little more tedious than shift+tab
I see what you’re saying. That’s actually why I built Nuanced (https://www.nuanced.dev), I wanted a better way for humans and agents to work through intent before the code got written.
but I think I overfit the interface to how I used to work without agents. when I was at GitHub, I wrote a lot of ADRs, RFCs, and design docs, and I liked that because writing is how I clarify my own thinking. with agents though, I’m often not doing the writing myself. I give the model a rough intent and it fills in a bunch of gaps, and then I get back a long, polished plan containing decisions I didn’t explicitly make.
That’s the part that feels broken to me. The plan can be detailed and technically correct, but still be hard to review because the important bits are buried and feel distant from my own thinking. All the assumptions and tradeoffs and questions may or may not be legitimate, but it’s hard for me to get into flow state and carefully check them.
So I still want the collaboration step. I’m just less convinced that generated prose à la plan mode is the right interface for it.
There are many ways people solve this. I use spec directories, and the agents just keeps updating it as I make decisions. It isn't a single file, it's a directory; sometimes nested based on the size of the project. I'll have the directory open in my editor of choice. It being a directory allows me to use an arsenal of text tools, and also optionally commit it for history/tracking.
This tool is a difficult sell. Users _might_ consider an Open Source version, but switching people away from familiar tools is not easy.
I also don’t get what the problem is. Claude’s plan mode is great, especially when you add interviewing into the equation. It means I can get to a point where I’m confident all the changes proposed make sense, and that I understand them, before anything is actually changed.
I feel like this post is hawking for commercial reasons more than it’s based in reality, especially a reality that takes into account the extremely varied experiences engineers seem to have with agents.
As such I’m not inclined to take it very seriously. More than that, the “X is dead” trope was overdone 10 years ago and I don’t think it’s been long enough to warrant a resurgence.
The real reason why plan mode is dead is because you can just conversationally instruct the agent to not make changes to the repository or to make changes to selected documents only, and it will listen. There was a time when we needed to enforce this via selected tool use, but we have surpassed that.
It also doesn't allow leaving an audit trail of plans and decisions (by default, anyway). Most of my mutating prompts look like "Propose a plan for change X and write to file Y" and "Execute steps M-N from file Y".
RE the article: I don't think it's obvious why this process is worth following until you find your time and attention wasted. Conversationally-building is the express train to waste. I'm not sure why you would even be talking to claude if you don't understand what you want to build.
I agree. Plan mode came about because earlier models were loose cannons, doing what they pleased. Today's models follow instructions better enough to not need a separate mode. However, there is still value in using a separate, smarter model for planning than execution.
I typically converse with the default model to point the plan in the right direction, then have it iterate with a smarter reviewer to find flaws until the plan is converged.
Do you really trust it with a production code base or database? Telling it not to make changes feels an awful lot like "Make no mistakes". I also like how Plan mode on Codex asks clarifying questions. I'm in management now so I've used Work more than Code lately.
Claude code’s plan mode began its slow march towards deprecation when they hid the “clear context and implement” option behind a flag. From the discussion at that time it seemed that the developers considered it mostly a legacy feature, and that newer models are smart enough in long context not to need it.
This part I find a little disappointing. The plan is the perfect boundary to cut context which still isn’t free, even if models can handle a lot of it now.
This is why open harnesses and open models are superior, it's only a matter of time until someone decides that something you use isn't worth maintaining anymore. The reply from the Anthropic employee up above doesn't give me much confidence that Claude users will be able to continue using plan mode forever.
Codex and Claude Code are the best harnesses out there by a long shot. They work out of the box, require no real config, and just do the thing.
The flexibility and lightweight nature of open harnesses is mainly useful for open models. These models are considerably far behind the frontier and need a lot more steering. Complex harnesses confuse them. I use open harnesses for my OSS model rig because the model needs it.
Sure, but their workflow is the best for Claude, and it's the same with Codex for OpenAI's models. These companies work and live with their own harnesses, they are ahead of the curve on what works for their models and what doesn't.
Disagree. While it's easier to 1-shot more things now, when there's more complicated interaction of components, a good plan from plan mode can make the project a "looks good, go" and run overnight thing vs. "steering" needed.
How many thousands of dollars are you burning a month on “looks good, go” and then *letting it run overnight*?? I spend enough money on 8 human, interruptible and human-thinking hours with paused hours.
I can’t imagine how high that number would go if I got rid of the collaborative style and also the “please don’t run scary commands on random directories without permission” mode.
Write a good high level plan with functional and non-functional requirements and it works out costing less, works more efficiently, and no surprises.
It’s like this: If you know the right words to say to the LLM, you’ll likely get back the “right words” also. And the more right words up front, the less steering you need to do.
(Having 20yrs of experience writing these user stories and designs help, my “unfair” advantage)
I'd like a brainstorming mode, with limited side effects.
I want it to go and read stuff but also invoke some commands and generate reports and discuss on the results. Latest GPTs in codex plan mode always go.. write a plan. Shocking, i know, but thats not what i want on every turn when im in that mode.
[I work on Claude Code] I broadly agree with the author’s point: plan mode was useful, and is no longer useful.
In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.
Yeah, as a user I came to the same conclusion as I naturally used plan mode less and less over time.
I still use plan mode in Astra to come up with a plan that I then feed into Fable. I feel like OpenAI models still do better big picture investigation and planning, while Claude is the better software engineer, if that makes any sense.
Of course this could well come down to my own biases and the specific things I’m working on.
It's useful because it let's me see the decisions the model will make before it wastes a ton of time implementing them. The model is smarter now but that doesn't solve for underspecification if it guesses my intent wrong
I greatly prefer this, since it lets me iterate on the plan with Claude for a while without it repeatedly asking if I’m ready to implement the plan.
Once I’m satisfied, I usually start a fresh session and tell it to implement the plan.
For smaller plans, you don’t need the file. Just ask it to come up with a plan. I don’t recall the last time it just started implementing if I only asked for a plan.
It’s not that planning is dead, but rather that planning has outgrown the simple “Plan Mode” feature as models have become capable of taking longer turns.
I really wonder what I’m missing because my general rule to plan first unless explicitly instructed works perfectly fine with all models I use.
“I want to ...” / “Let’s ...” -> Plan
“Do X” -> Actually Act.
But then again I also have it configured to only ever answer questions instead of inferring them to be instructions (which I’ve seen others do differently).
As someone that never used the built-in plan mode, but did use a lot of spec-driven development, I’m still finding that even with Fable having “plan” docs is still quite helpful.
They’re most useful for broad changes (new features, refactors, etc.) where it’s helpful to avoid breaking changes or unnecessary scope expansion.
The new models are great, but they do more by default, which means I’m finding myself explaining what _not_ to do more often than with previous models (where they’d often end too early).
In my case, the previous plan mode was too ephemeral, and I like having one source of “truth” that sits across context windows without loss/compaction.
As a long time user of Claude Code, I've just naturally stopped using it because the model figures it out, and i prompt it accordingly like 'come up with a plan'. Glad to see your experience matches
hi I’m the author of the post. I think that’s basically the distinction I’m trying to make.
Historically, plan mode served two different roles:
1. making the agent’s instructions precise enough to execute
2. helping the human understand what was about to happen
I think #1 is less necessary as agents get better. #2 is going the other direction, it becomes more important as the model is able to do more on its own because larger chunks of work are happening with increasing complexity.
Where I’ve changed my mind is the interface for #2. I increasingly think an interactive, iterative workflow is closer to how people actually build understanding than being handed a long generated document, especially one they didn’t author themselves.
The human-understanding problem is very real though
Awesome blog. You're a good writer. I enjoyed seeing your article on AI in 2018. Thank you for sharing your expertise.
Also I hope your delivery goes well. My wife (and co-founder) had a challenging delivery and it really put life in to perspective for both of us on a range of issues (how much women's pain is minimized in the health system requiring stronger personal advocacy than I would ever have expected).
As far as plan mode, I still find it essential in keeping agents on track. I build propelcode.app and have a variation on plan mode I still find useful, happy to trade notes on agentic coding if youre interested.
I'm not the person who responded to you but I hope they're able to help her and she's okay.
And thank you for the article - it was a good read. I still see folks in my org playing "throw spaghetti at the wall and see what works" and getting frustrated so plan mode (mostly point 2) has been their guardrails almost as much as for the AI.
I never understood fear and anxiety until I had a kid and was worried about her safety. I am sorry youre facing that and hope she recovers and thrives.
Can't help but think "Doesn't matter if a machine or a human with (even slightly) different background wrote it", maximizing information flow is maximizing common assumptions and "culture" to only have to communicate a small set of current information for the task at hand. Being a team means having built a joint context so to say. This has always been the purpose of design documents and they always were too big or too small. Because you did not write them, but the others. If you only produce code you think they are the past and useless. If you iterate and your team grows, you start seeing the value in always current docs that are containing just what is not in your everyday culture.
All the best for you and your growing family. I had a similar experience recalibrating my values...
I like plan mode personally. I only use claude code for the web, and the questions claude asks me to clarify are usually pretty important - mostly because I was too vague or contradictory in my prompt, or what I was asking for conflicted with something else in the code. I don't know how claude would resolve that without plan mode.
Also for session planning, as in when-can-I-walk-away-from-computer, its nice to know the particular rhythm of initial crunch - ask questions - make plan - do it. Especially with a 5 minute cache timeout.
yeah (at an early stage data-focused startup). I think when our product is a bit more mature we'd probably have a staging tier with full access and then manually promote builds / DB changes to prod. So many things to build lol
we DO daily snapshotting, so the risk is limited... but still spooks me
FYI I had Clod attempt to corrupt a prod db the other day. (Opus 5)
I was experimenting with a rather complicated backfill operation, were I had a validation script I understand and have Clod come up with the backfill script. I was running against a local prod copy, and it proposed running the actual (unfinished) backfill script against prod.
It didn't have access to the secrets and I also caught the command, but a good reminder that this stuff needs guardrails.
yeah that's crazy. its so good 99% of the time but I've seen it have some insane hallucinations before (as late as Fable... cant remember if it was Fable 5 or Fable 5.1).
Hallucination not a big deal when it's on the surface layer. But I can't imagine the damage it could do if it hallucinated while building/validating a "load-bearing" component and then continued down that path
Plan mode was great, but I realized I progressed well beyond it. I found that I was getting these categories repeated errors and oversights from Claude (and frankly it hasn't gotten much better about this). Skills were too generic and got lost to context.
I ended up building out tool an MCP server that I use as a bit of a psuedo harness for Claude. I have a variety of multi-step workflows that are basically micro-skills stacked on top of each other. This helps me make sure that I can get Claude to think in a repeatable and reliable manner.
For coding, I've found that I have a few specific steps that Claude needs to do before I'm comfortable letting it loose:
* It must extensively explore the code base (including certain areas that it misses)
* It must think about what it doesn't know or is making assumptions about
* It MUST scaffold out it's intentions. Essentially, it can write comments, classes, and method stubs - but no actual content. Very much like a spec, but since it's in and alongside other code, it's much easier to identify problems.
* It must spike and validate key assumptions. This, plus the prior step, are the only way I've figured out how to avoid it ending up in a confusion loop. Too often it looks at poor-quality code it's written and thinks it's a long-term solution. By avoiding writing code as much as possible, it knows that it's draft content.
I think over time more and more will be peeled back to just the model and markdown. I have a beautiful factory running with key personas all it is is a few markdown files it is building a mac app fantastically well.
In my experience, around the time the author describes as starting to not need plan mode is when erasing the chat history became an anti-feature. I found the agents were doing better when they had the context already, and with the history I no longer needed to micro-manage persisting various caveats and rejections to the plan artifact. The amount of prompt construction necessary went down overall.
By the time my plan’s done I’m usually between 200 and 350k context. Even if keeping that around gives a performance bump for the implementation (which I haven’t noticed to be the case) it balloons the cost. I would much rather put everything in a plan file and start fresh.
Plus, I usually plan with a more expensive model and guide implementation with a cheaper model (with smaller validation calls back to a more expensive model)
that still essentially works. ask it to write a plan doc to a file. then when you're ready to implement, start a new session with a prompt to review the plan doc and then start building.
> This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For me it’s actually the opposite, and Claude Code’s plan mode isn’t nearly sufficient. Personally I ask Claude to write down a markdown file with its plan, then review the plan using plannotator, and then go back and forth (most of the time it’s actually the comments that are the problem, not the code).
Then start a fresh session, seed it with the plan, tell Claude to find ambiguities / friction points / oversights, resolve those, and then implement it.
Review once again with plannotator, go back and forth, and then send PR.
Maybe not the “vibe coding” that was once imagined, but this does ensure I am fully aware of the code and architecture, the quality, and this also prevents long term degradation.
I do the same, I don't use Claude Code or Codex planning because it is mostly pointless, even with Fable Astra. I just have multiple agents work on a markdown file which I manually perfect, often breaking into multiple different files for large features or PRs. I also create design 'handoff' documents which I feed into Claude Design or Astra along with screenshots and wireframes. By the time an agent does something I'm well prepared.
I've tried doing the incremental, iterative approach with just Code and it's just not as effective unless you're working on something simple or experimental. Or you're shipping to something non-serious or perpetually beta.
Yeah I always balked at the amount bloat in that repo, and I just prefer a more lightweight setup and plannotator’s UI which I can use to interactively review stuff. It solves this one problem better, where superpowers tries to do way too much.
I dived into one or two of the skills there and some of it feels very cargo-culty. In a similar vein to the article, the models have improved markedly now and I’m of convinced that “You are an experienced senior software engineer and an expert reviewer, here are five pages of instructions on how to do a review” style prompts are massively useful any more. I get great review results (as good or better than colleagues using superpowers or even adversarial review skills) just by asking Claude to review a PR and spit out results in order of severity.
Superpowers never really held up in my own testing vs /plan - not just more expensive but markedly worse code organisation, because the plan itself leads to higher cognitive burden for the models - and today it's even worse, because /plan doesn't hold up versus coming up with a few high level slices and having cc work on one per session, for much the same reason.
5.5 is much closer to Fable so i don't even need it. I am pretty sure it's got Fable's DNA in it.
I really need to find a role where I can do more DX...
Going through the superpowers spec-plan-implement-review cycle took one of my test runs (local AI) from 40 minutes to 8 hours, and produced a lower-quality result. I want to use something like that skill, but I'm doubting the efficacy of this one.
Yeah, sure, but you don't need a dedicated plan mode for that at all. you can just do it in auto mode, and say "let's do some planning first", and Claude will (nowadays) be smart enough to understand that it's not supposed to jump straight into the implementation.
So again: You don't need plan mode, auto mode works just fine, there is no difference in the workflows here.
I do something similar but a bit more involved, using a few informal stages. Let's say for example I'm trying to launch a new complex feature for https://coderba.se.
- strategy document
- "sprint" document with technical implementation
- actual implementation
- e2e testing scenarios updates
Every step involves iterating with Claude on it with me in the loop (setting the direction then resolving the "founder questions" as they appear), and importantly a different model for review/code-review, be it Codex (usually, it's great at it) or Antigravity/Gemini (sometimes finds novel things, its precision and recall are abysmal but on the odd occasion it has good accuracy). This iteration on the high-level plan then on the implementation plan is essential to me, and IMHO part of why people are surprised that I tend to get solid results from LLMs. At the very least, it allows me to fill gaps in my own knowledge (primarily front-end development) and be more productive than writing the code myself. I cannot stress enough how nice it is to have a partner in the high-level system design – yes, it often suggests utterly moronic ideas, but the overall experience is still net positive and getting better every quarter.
> (most of the time it’s actually the comments that are the problem, not the code).
I wonder if it's just a consequence of a gigantic training set full of comments completely out-of-date with the code, leading to the model considering this "normal"
I don't usually see Claude leaving comments that are obviously out of date compared to the code. The problem is that the comments are just a dumping grounds for Claude's stream of consciousness, verbosely recording random bits of history and context that are true and at least somewhat relevant, but without cohesively explaining why the code is the way it is. And some sentences in the comments will be beginner-style restating what's obvious from reading the code.
Exactly. Comments are basically its scratchpad for whatever it wants. Ask it to change a number in a TOML config and it'll also add 2 new comment lines above it with some Claudeslop about it being changed and what that accomplishes, as a useless note to itself.
I now make sure to do a big decommenting pass before every PR.
This kind of stuff annoys me too, and I have instructions in agent.md to avoid it, which helps.
But I also am starting to just let go and stop caring. It’s not clear to me that it causes problems down the road, it’s easy to strip out en-masse if needed, and in my experience, agents now are really good at read git blame, the commit log, even prior agent transcripts if available to sleuth out when a change was made and why. So yeah, it’s annoying, but the code agents write for me is increasingly never read by a human, so does it matter?
This. Be honest: who reads the human-written comments? ;)
When I read, I skip most comments, especially the larger "Javadoc" style. My brain sees them colored differently in the editor and it doesn't even take mental effort. Then, when I have a question about the code, I look back up for a relevant comment. That doesn't happen very often.
If Claude writes great code and leaves a garbled Claudese-but-accurate comment ... I can read and comprehend (with like 5x the effort of a human comment) ... that's a small price to pay.
Yes, comments might be my biggest gripe with CC due to what you said. 9 out of 10 of my revisions to Claude's work is deleting or rewriting comments. Such a silly thing for it to fail on.
(I do have detailed instructions for it on how to comment (or not) but it has not fixed this.)
Nah, it’s the typical Claude-isms, and way too much info, info in the wrong places, and putting lots of planning discussions into comments that are completely irrelevant.
It’s always “you explain only what but not why” or “this is way too much prose” or “these comments don’t belong here, they should be inline comments” or “this is completely redundant as it’s already obvious from the code”.
I do something similar. But where I change it up a bit is depending on plan complexity I divvy up parts of the implementation to different subagents with fresh, only relevant context for whatever they're responsible for doing (e.g. a part of that plan). Edit: to clarify, I will also go back and forth with the planning agent making sure edge cases are covered. Sometimes that involves invoking a new subagent without our planning context to validate it without preconceptions.
I personally still find planning a valuable mental exercise; it's not so different from pre-LLMs and whiteboarding or otherwise taking the time to consciously plan a set of work.
I've recently gotten religion on the workflow that is many (relatively) short-lived agent sessions passing planning/handoff docs between themselves. It's better for my own task tracking, better for handling "oh btw I noticed XXXX", and better as a clear review point. Overall it just feels like it takes a lot of the formerly implicit context that was whatever we happened to have talked about and turns it into a much more explicit "this is what you need to know, now go".
Currently looking for a framework for managing this in a more formal way, and I think it's probably beads, but interested to hear from others.
I've been on this kick since I realized the primacy of the initial part of the session context. I created a python app that reads a phased plan and kicks off a new session for each phase. There is a standard prompt and handoff mechanism to determine if we encountered any unforeseen issues that we need to address in chat, but otherwise it will just grind with a clean session with appropriate context for each phase.
Just kicking off a subagent does this automatically...?
My main conversation is usually with an orchestrator that hands off work to various (usually cheaper) subagents to plan / review / etc. It has instructions to find the correct model for each task and not to do too much itself so a multi-phase plan automatically gets a fresh subagent for each phase.
i disable subagents. it also helps that i work in microservice size repos, so a single agent can keep the full relevant context. subagents waste alot of unnecessary tokens doing code exploration that the main agent is going to end up duplicating, i think they make more sense for monorepos.
anecdotally I've tried the same plan on different branches with both approaches (single agent with subagents implement the full plan vs. my little session per phase app) and then had agents judge the code quality and my approach worked better and saves tokens. horses for courses.
Spent the past 1,5 years building a tool that might be relevant, aimed at keeping durable task state between agent sessions. It is an issue tracker persisting state as immutable event logs, allows you to inspect workflows after the fact, lets you inspect diffs inline in the tickets and it is much more lightweight than Jira/Linear. There is no central service to integrate with, as it is Git-backed and lives with your code in your repo.
I have some older projects that use beads (I still run an old version without dolt that's imho pretty good overall) but lately with Fable also have a few newer projects where I just have the agent write docs and keep a worklog with the what/why/decisions etc. (I think I read it here on HN somewhere and figured I'd give that a try.)
The latter seems to work pretty well for now (slightly better than beads) but I'm always looking for ways to improve it. This could be an interesting replacement.
I have found that it works well for tracing decision forks. I tell my agents to use a "human-input-needed" tag when there is a high stakes decision fork, and then I just filter on that tag, immediately highlighting blocked workflows. You can get a long way with clever tagging, filtering + time travel.
I think it works well for tracking intent. I tell my agents to use a "human-input-needed" tag when there is a high stakes decision fork, and then I just filter on that tag, immediately highlighting blocked workflows. I think you can get a long way with clever tagging, filtering + time travel.
This looks really nice and I'm definitely going to give it a go. I been relying on Jira so this will be a breath of fresh air. Beads was great in principle and I haven't given it a look in a while but this looks much cleaner.
Checking this out, the skill it wants me to add to my project seems to reference issues that wouldn't make sense in the context of my project. IE: `QPPR7PE` and `MBB6WY5` ?
You are totally right in that it makes no sense. It seems like the skill.md had regressed as I had updated it via agent proxy. Thanks for pointing it out, the skill has been brushed up.
Got it! I am not up to speed on Claude cloud sessions yet. Interested to hear how that works out for you. Let me know if you figure out what's getting in the way, I'd be happy to look into it.
I have a lot of little projects and I also prefer this way of working with agents. Sometimes I would start to interrogate on a specific portion or ask questions to better understand a concept, and the session would get poisoned and the agent would fixate on that topic for all the remaining turns.
I asked fable to look at my interaction patterns and clearly stated my frustrations and the problems I wanted solved, and it designed a simple process to track things in git and built a couple simple session hook skills. It’s pretty lightweight and I’ve been very happy with it for a couple months.
The power of this mode of work is that after you deconstruct the task into smaller subtasks, it's a lot easier to use cheaper models to implement that task.
I get a long way using models like Opus to make a plan of action and a bunch of tasks, and then using Deepseek to implement that plan of action. Saves a bunch of money and is fast.
- I have a record of work done and work to be done that helps _me_ when I come back to the project after several months. It’s committed and lives with the code.
- when a task inevitably ends up more complicated than I thought, I can in that session break it up
- I initiate sessions from multiple computers, so things stay in sync (through git)
- I also have a “tooling” repo that builds out some views of the work and hosts it for me to see when I’m on my phone.
- The hooks let the agent manage all of the workflow/task management, so there’s very little management overhead for me.
I rejected beads and JIRA. I wanted something more lightweight.
- start: invokes `work brief` skill to put the current tasks into context
- pre-compact: does the same so work state survives compaction
- stop: lints the work-tracking yaml files to make sure that updates mid session don't break
Nice. I found myself doing a lot of repetitive stuff handing off between agents, time to fix that. I think I agree with your approach of keeping it lean.
I also sometimes rewind the conversation after going off on a tangent, sometimes not if I want the agent to have the sorts of things I’m thinking about in context. I probably justify keeping too much crud in context than I should
I use Jobs [0] to manage this—it's an agent-first CLI to track issues and tasks. A single `job orient` command gives the agent the current task in the context of the larger plan. It's a replacement for Plan Mode and issue trackers, and it has allowed me to execute massive plans in parallel with minimal oversight. There's a web UI, but it's a work in progress.
I have roughly the same workflow, also with plannotator - which I like a lot - and haven't used or felt the need to use _plan mode_ for at least 3-4 months.
Then telling Claude to work on a document, the instruction is kept to its core.
Now when bcherny explicitly mentioned that it merely adds a single line - it explains why I don't need it.
What may be concerning about "super plan" mode from the creators (or a skill, for that matter) - is that tuning the amount of effort, and how much deep to dig - may become too hard, as it will interfere with several embedded paragraphs explaining what to do, how to do, where to do, etc'.
What I do look for is even better plannotator ability to track changes, combining historical comments (like Google docs), and git blame of several "generations" before current reviewed doc.
While possible to commit, I'd still loose the comments themselves, and the commit will be at least per changeset, and not connected ro a specific changed section.
Roughly speaking, I'd be happy if plannotator would persist something similar to github PR reviews combined with Google docs comments & suggestions.
Have you recently tried working without generating that plan? What I've been doing is first tell the model I want to plan the implementation, talk about it a few turns, then when I'm happy with the idea and method just tell the model to go ahead with the work.
Note that this is only really necessary for complex work that I don't know yet what the best way to do it is.
I've tried doing it your way as well, but there was just too much fiddling about with writing the plan somewhere, then having another session rebuild their context with whatever info is in the plan. It really didn't result in better output for me.
Currently 9 times out of 10 I just say to the model: xyz is the problem/bug/feature, fix it. Since about Fable and Opus 5, this is more than enough. Opus 5.5 (and previously Fable 5.1) got even better at this. However, this is in a codebase where there are already a few hundred thousand lines of code for the model to look at to see how we generally attack things in our codebase.
Claude Codes plan mode I never use anymore, it was useful a few months ago because the models had a tendency to just start doing work and forget I specifically told them not to. But the UX is just annoying and the models now do adhere when I tell them not to change anything.
I have, and it always does something unexpected and undesirable.
Plan mode ensures I'm spending fewer tokens on the code-test loop, and more on the arch/design, and allows me to keep appraised of what's going on, while planning for future changes better.
Maybe folks who don't need planning, don't have as much concern for the details, and are happy enough with just evaluation of if it works or not.
My workflow is very similar. But I just ask agents to write design in html instead of markdown, due to its richer layout and better interactivity. When the design is about UI or anything related to graphics, this approach is extremely efficient.
I also found that having the design reviewed by multiple agents has very little marginal value. The review agent will always find something to improve, but mostly it’s just nit and not anything super important.
I used to let Claude just upload the html design doc to Claude artifacts for me to review. Recently I switched to codex and started to use my own tool https://github.com/hyperlogue/r3 to complete this workflow.
> plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive
seems like plan mode could turn off some tools, even if it doesn't change the set offered to the model, the ones that they have which would mutate your codebase could just not work with an error message, and plan mode could change permissions in the security approval prompt for "auto"
anyway, isnt the right way to know if plan mode helps or not, to run an experiment? we're all guessing unless we have data
read only agent mode sounds straightforward and useful to me
It’s bizarre to me that Claude Code doesn’t have more guardrails on plan mode. I mostly use Copilot and the plan mode there has teeth. In plan the agent does not have permission to write to the filesystem outside of a temp directory and the plan document itself, and tool calling is heavily restricted.
There’s also a “ask” mode which is read only. Both are enforced by the permissions model, not just a system prompt instruction. I’ve seen the model “forget” and try to start coding - it bounces off a hard permissions failure and that “reminds” it that it’s in plan mode.
My team has been struggling to understand whether or not we should do "spec-driven" development or not. It makes a lot of sense to us to have one developer iterate with the model/harness to generate a markddown document that is a high-level of what will be implemented, and then have the team PR review it before and agent attempts to do the actual implementation work. Do you think this is a good practice?
I specifically work in tooling now, so this probably applies more to that domain than some others, but I find 'very up-front spec-driven development' unappealing for that kind of work.
I'm experimenting just like everyone else, but this is my process right now:
- Quick prototype
- Figure out the language of your app (what terms you want to use for things, what your UI design language will be, etc) and spec that, so you can use words consistently with the agent. You need to be able to describe the things you want well and consistently.
- Keep prototyping. Let the agent write unit tests along the way. Lock down behaviour you like, keep track of those things in a document.
- At some point your idea of the real architecture comes into focus, from actual use cases -- avoids the over-abstracting right away trap.
- Refactoring is cheap with tests, so start refactoring into the architecture you want.
- Your architecture won't necessarily be what would be best for a human, but it will be pretty close.
- Keep relentlessly iterating on small work.
- Things that were expensive before aren't that expensive now -- integrating a library, changing from one library to another, trying out a few architectural refactors, trying out different performance optimizations, etc. That stuff is all 'throw it there and see what sticks' now, so don't be afraid to try stuff which felt big before.
I feel like 'front loading' too much is just the wrong approach. You might feel like you're sitting there 'babysitting the agent'; but that's just what the hard part of the work (hard as in 'zjust slogging through it', not as in 'conceptually complex') looks like now. Your code is much more like clay.
Atleast that's how I'm thinking about it so far, but I'm not working on large sprawling systems that I imagine would need more pre-planning.
I believe spec driven design is a good way to go. But that means keeping all your specs either in the repo (if you want the agent to be able to edit them) or available via for example MCP (if you think only humans should be able to edit them and you got some type of external documentation system keeping specs).
But given that running an agent us cheaper than the cost of waiting for a slot to assemble the team to talk about a change (isn’t it always?), why wait with running the agent?
I propose updating the spec then do the implementation. This will most likely show that a few assumptions were wrong forcing some major or minor updates to the spec. Work through those and then let your team review the spec change together with testing the next iteration of what what’s build.
I have found it doesn't see me any direct time. I either iterate ruthlessly on the spec or on the implementation.
I find that the code is generally in a better place proportionate to the amount of SDD I actually do. But it's just a matter of where and when I want to spend my time.
SDD is good. It's a lot faster and cheaper to have an agent polish a spec, than refactor code. The implementation can then be verified against the spec artifact, and any drift can feed back into process improvements for future specs.
I use plan mode, because my current project benefits from "pair programming."
I have one session define a task, and provide a formal specification plus context in a "cover letter."
The session B, in plan mode, produces the plan back.
Session one reviews the plan and clears it, ratifying portions and often specifying specific changes.
Session one then executes.
What has been striking to me in this approach is that even with two instances of the same model (currently Opus 5.5), there are regularly corrections made. I use "project chat" for session A and Code for session B atm; it is very typical that Code finds and corrects details or oversights in the task spec; it is also typical (though less so with 5.5) that session A (chat) pushes back or clarifies things Code doesn't have the context for.
I have been afraid to open up the potential of negotiation beyond what this is costing as it is. But I am also afraid to simply skip the formalisms, because of the consistent correction that occurs in this back-and-forth.
Each component of the pattern is schematized, generated from a template, and validated, to keep things tight.
Lots of tokens! But I trust this process far more than "just typing" :)
I always append something along the lines of “evaluate”, “investigate” or “report only” to my prompts when I want to see what the agent is gonna do. Because especially with the new models they tend to go easily off rail and do stuff I didn’t ask. To say that they just “get it” is highly dependent on the task, scope and blast radius.
I disagree with the assertion that the model gets it. Here’s a practical example I just tried with Fable 5.1. I gave it this prompt: “Write a Go function that can be used to establish secure communication to a remote system using a certificate. Keep it short, single function, and explain how to use it.” The output forced the use of a private key stored in a file even though that wasn’t specified anywhere as a requirement. The function Claude wrote takes a private key file argument and calls a Golang function that requires a private key file (tls.LoadX509KeyPair) even though Go has crypto.Signer which could support private keys in various other manifestations like HSM or KMS. I argue that a person who “gets it” (or who is reasonably experienced in security) would have opted for not requiring private key material for this to work.
For the record, this isn’t unique to Claude. ChatGPT and Gemini do the same, each with its own quirks. ChatGPT got extra credit for being the only one who allowed the function to also take a CA file for server authentication.
Don’t get me wrong: LLMs are the future (maybe even the present) of software development but I think there’s some way to go before they can be entirely hands-off in some areas. I still find myself having to course correct designs and plan mode helps me with that.
And of course, thank you for your work on Claude. :)
I'm hoping that by "gets it," he meant that if you start a discussion about the design, it doesn't misunderstand and immediately go off to do the work. Some models tend to do this.
I think knowing which models use the standard library and which ones pull in dependencies as highly useful. As a dev I've always favored staying as close to the standard library as possible as it makes refactoring, understanding, and deletion much easier. Not everyone has the same preferences as you, it's nice to understand what matters for others too.
Apologies, I didn’t mean to imply it’s a benchmark, I just wanted to provide a reproducible example of where I see models make decisions that seem to be fine initially but might paint the software architecture into a challenging corner. I don’t expect models to read my mind, but I do see them produce a lot of verbose output, none of which is used to say “here’s a simple response to your ask, but have you also considered...”
Whether or not a distinct "plan mode" is needed, upfront planning remains essential in my experience, even with Fable (albeit not the 5.1 version). I agree that, as the models get better, you can skip planning on increasingly complicated tasks.
But there is still a ceiling above which it is necessary to "preload" the context window before starting to call tools and get into the meat of the work. You want to establish domain language (especially with Claude models which otherwise will invent their own, and it will be inscrutable) and key requirements and assumptions. You want to do a Q&A iteration cycle with the LLM. You definitely should do a sanity check that the LLM actually "understands" what you were trying to achieve, and then make sure that understanding is coherently and plainly stated in the prompt. All of that seems to be necessary still for just about any serious task, if you actually care about the quality of the results and/or don't want to burn hundreds of thousands of tokens on flailing around to get to a good quality result.
So no, you don't "need" plan mode. But you do still need to do all of the things you would do with plan mode.
But in your example you never even asked it to plan so you could check the implementation before writing it you just asked it to write it directly, so this isn't even a comparison to plan mode
Hi, thanks for Claude Code. I use it, and it works well. Have you considered changing it so the text comes down from the top of the screen, in green, like The Matrix?
Plan mode is what makes it easy for me to update my mental model of the codebase, and helps me decide if the mental overhead of all the changes that are needed are even worth doing, which are now the biggest limiting factors (reading the code diffs just doesn’t scale anymore).
> In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet"
Can't help but think if plan mode isn't useful as you say because it's implementation is lacking in claude code, hypothetically speaking.
What I can say is that with other harnesses plan mode helps stabilize my workflows. Actually synthesizing code is only part of the process, lots involved in taking a work-item to production end-to-end and plan mode helps give this flow structure. More than that it's an opportunity to regroup before committing to changes. It slows down the process to a rythim that's sustainable and smooth, which ends up speeding up the process.
So if plan mode in claude was designed to speed up code churn, while oh my pi for instance designed plan mode to be strategic, that might account for the different perceptions here.
And it's not to say you should force yourself to use it, but if you are planning on cutting this mode off the loop just beware of the possible side effects.
Honestly I don't agree with this. A few reasons why I feel like I will always want a plan mode:
1. I want to know whats going to happen, at least at a high level, before changes are actually made.
2. Plan mode helps me flesh out the missing details of my plan before being mid-execution
3. In situations where I have a limited budget for AI usage I will often times use a high powered model like Opus 5.5 or Fable to make a detailed plan, then scale down to a cheaper model for implementation. I feel like this saves cost in the end.
I get plan mode is basically just a small hidden prompt. I get that I can basically just preface my prompts with "make a plan only, don't make actual changes." Maybe this is just a UX trick, but it works well for my brain.
I created a skill "feasibility", which is basically "evaluate feasibility of this idea and propose solution options". So I can check that Claude's idea of how to implement the requirement is close enough to my own before it starts working. It is a lot more lightweight than plan mode, because the skill says not to build a detailed plan, just a high-level summary.
Yeah I think this is a natural consequence of longer task horizons. When I was chaining 4h tasks, I can mostly plan them up front.
Now that I’m frequently designing and delegating day/week scale features, the flow has to change; having the agent go off and build a spike can be a quicker way of us understanding the design space and constraints (especially in a huge codebase). I still have the agent write and update a spec doc as I go, but it’s not waterfall anymore.
At least for my kinesthetic learning mode a rough code PR stack is usually way better than a plan doc anyway, and tokens are cheap enough (vs my time) that going further than just a plan is often cost-effective overall.
The dream of course is (say it with me) loops, but that doesn’t tend to work for me on new features often.
I'm not an Anthropic model user, and the true frontier of the frontiers is beyond my budget. Maybe it's better in the rarefied atmosphere of Astra, Fable and Opus 5.5?
But with GPT 5.6 Sol, I'm still finding that the model makes conceptual mistakes, or gets edge cases wrong, or assumes incorrectly (making an ass out of both user and model). In many cases, I need to at least refine the proposed approach, or amend, correct, or flat out just stop and start over. Not planning and catching these errors, and just letting the agents code their code, would mean I'd have to rollback and redo many times. What a waste!
For a current project, which is ~33k lines of code, I'm also finding that I know the codebase better than the model, and that's vital at the planning stages too. If I wasn't in the planning loop, the model would have reinvented various wheels a few times over. How much spaghetti do you want with your code?
As always, I may simply be doing this wrong. But I'm personally not convinced that the plan is dead, or that I want the plan to be dead. Planning is also good for me -- it keeps me thinking about the code, prompting better, guiding the model better.
If I'm no longer on top of the codebase, then at some point my prompts will devolve to "Do the thing with the thing, that does thing". And I don't want that.
Yer, it’s probably best to face reality and understand that for actual software engineering / complex coding work - Anthropic models are way ahead of OpenAIs…
Agree, and for me I feel like I often have more implicit intentions than I write in a prompt. A plan helps me verify whether an agent gets these right or not. Plus, it highlights tradeoffs I might've not thought about. Removing both feels like lowering a quality bar.
On the other hand, for a low effort hobby project: just do the thing.
It is absolutely true that Opus 5.5 just ‘gets it’ far more often than gpt-5.6-sol, which is more like an idiot savant. It can nearly always do what you ask it to, but that might not be what you want.
It seems unsurprising that if your definition of "plan mode" is as tiny as appending "but don't write code yet", that it would not be that useful for that long. There have to be more sophisticated versions of what "plan mode" means out there.
Also if you are working in a heavily vibe-coded codebase, as Claude Code reportedly is, it's not that surprising if the human doesn't really understand it or have anything useful to add in a collaborative context.
GitHub Copilot for VS Code by default configures "Plan" mode to disable everything except reading, asking the user questions, and writing into the memory scratchpad.
Claude Code's tool system feels quite rudimentary in comparison, but that's what you get I guess when you're only working with models you're going to finetune on being able to handle specifically your own tools anyway.
that's the problem though:
- LLM now output good enough results without a plan.. for coding at least. I'm not saying amazing results..just good enough. it works fine.
- Most people suck at planning anyway
- LLM still don't give you a way to verify and understand to iterate.. you have to ask and then formulate and way so people barely do it, they just trust the vibe
IMO the current successor to plan mode should be the harness knowing when to tell the use "ok here's our overall current state in a simple diagram", auromatically
Reading the replies, it's interesting how idiomatic everything is, eh? We must all opine about our specific setups, panes, prompts, and how they're the true way, or the true alternative to plan mode.
Thanks for the context, and makes sense a lot. That’s my primary reason to use plan mode.
However how about decisions? Do we expect the model to read our minds, just assume the best practice will be followed and that’s what the user want? Plan mode solves those, what is that am I missing?
Fascinating! I literally never use Claude without plan mode and I find it's basically useless without it, constantly wasting tokens going in circles on irrelevant things. Fable or Opus. I feel like neither has a good sense for how to architect things and if I don't use plan mode it usually wastes hours of time chasing it's tail or implementing kludges on kludges to get something working that would be a much simple fix elsewhere, especially when working on a larger codebase.
Hence why the Anthropic employee is telling you to not use it and just aimlessly throw tokens at a wall. It'll eventually get you there, sure, and consume more tokens.
Also this is a way less removed process than I want nothing to do with. The more removed I am from the process the more I hate my job, get burned out and genuinely wish that Anthropic never existed.
Even if it could "just know" or infer my intent. It wouldnt be desirable.
I mean if you want it to plan something without writing code first you literally just need to ask it to plan out the implementation first it's not exactly a difficult concept
So you think telling a model once to "do not write code, just plan" vs having an enum that causes that same text to be added to each message will somehow waste tokens and will surely make anthropic much better off..
The Anthropic employee literally told you plan mode is a "we're still planning!!!" at the end of the prompt. Plan mode is useless because you can literally type "dont write code yet" and get the same effect, not because you should never plan in general.
Sure and I should also have been more specific, I don't always use "plan mode" I often just do planning without explicitly activating it just via prompting, point being that the robot doesn't work for me without some time spent planning out what to do.
I urge you to have a look at tools like tuicr and hunk. It’s a missing component in the chat. I really want a proper chat interface, leave comments on the code and have them replied to. No clue how to implant it tho, it’s a hard UX problem IMO
Recently I started using pi as a harness and its flexibility has really changed the way I think about UI into agents. It feels like I took the red pill and have been liberated from the prescriptive nature of CC or codex. Instead of getting told here’s how you should approach ___, I pay much more attention to my workflows and when I feel friction I make the harness adapt and smooth it out.
For example, there are no shortage of web based uis for pi and they are all cool but I wanted a deep integration between artifacts and how I want to collaborate with the agent on them. So i built my own ui that mirrors what the pi tui sees. It’s chat based but has a deep integration with a GitHub style code review UI so I can leave review and comments whenever I want. Every agent message renders nicely in an annotatable markdown viewer so I don’t need an ask question tool and can more naturally get the agent on the same page as me. I want it to feel like I’m working with a colleague.
I don’t think these features are too unique but having full control over the experience is really nice. Flexibility over model provider, can tailor it to my work’s dev stack, and don’t need to worry about anyone breaking it with a million updates everyday.
Your perfect workflow can be realized in a day or so. You just need to go make it happen.
I also didn't like the chat TUI for back and forth on things, especially on long markdown files where I felt like I needed to copy a line to comment on it. I wonder if LLM companies would rather skip a step and aim for a world where most people one-shot things, which sounds much more impressive.
I've been building https://crit.md to keep that back and forth with agents - GitHub-esque interface and have agents respond to my feedback, iterating until I'm happy.
Admittedly like many others, I use it a lot less for actual plans now, models are indeed getting really good at just getting it.
I wodner what this product space will look like a year from now. Reviews are already dying.
I ask Codex (I am not using Anthropic anymore) to always generate a plan file first, unless explicitly told otherwise.
I also tell it to write deviations and rename plans accordingly once done.
That way I keep the codebases I have to or enjoy to work on in my head and don't become too dependant on any provider or on stochastic parrots in general.
what worked for me is a frontend driving the agent, capturing every user message, every commit, every pull request, putting them in a graph (uses a frontend because I didn't want to load a coding agent with tools + responsibility of book keeping) and the agent get tools to search reason behind code changes and see the high quality user input underneath instead of the sloppy self written justifications
What I do nowadays, for large changes, is have Fable create a HTML explainer for what we're gonna do with code snippets, which is not that fat from plan mode only much more convenient for me and modern models have no issue using that HTML artifact as the source of truth, and then before I get into execution - I interactively build an end to end test that also includes pieces of the plan.
When a test case fails, the relevant part of the plan is surfaced in the error. I find this helps Claude stay on track for longer - I've been able to do 12h most times and even up to 48h unattended (11h of API time) with good results.
Then whenever I do check in, I ask it to update the HTML with current state in an append only fashion (sort of like it's writing a blog), and then based on that, we iterate on the end to end test (I think of it as a "test harness") - update the test cases and error messages.
I've been able to build some truly large projects this way, both greenfield and up to spec (for example, a video game I've always wanted to play), and brownfield while staying within the conventions and design of the codebase, and with very little attention required on my part.
I keep forgetting that CC has access to the "question" tool, even outside plan mode, which is mostly why I still use it sometimes. This isn't (or wasn't until recently) the case for Codex.
I get the idea here, but I find planning very useful as a phase when I'm doing things I am not intimately familiar with (e.g. AppKit programming).
That's when I need to learn about implied patterns, do's and don'ts; not from theory but in the context of my own project. Unfortunately, the model tends to keep implicit knowledge implicit. But I can ask during planning.
The feature did get less useful over time when the model started babbling in newspeak more and more. When it threw a thousand words at me even in concise mode.
So I don't want to let the toolmakers off the hook here. There's a lot to win that would make plan mode much much better without changing plan mode itself.
What, are you suggesting that there were people already successfully doing agentic coding in CLI harnesses with a good number of features before November 2025, the moment HN collectively deciding that it was now "good enough", if using "Opus on Claude Code"? Blasphemy!
Oh I thought it blocked tool calls like editing files...
You don't need to actually change the tool schema or break the prompt cache to do that. In the tool itself you could just check if it's in plan mode and reject the tool call...
I find myself endlessly ctrl+c ing claude now as it flies off doing deep first principles analysis to work out how to find a thing it isn't sure about but I know the answer. Being able to give it that answer without needing to ctrl c would be a massive improvement
I haven't reached for Plan Mode in awhile--maybe, on some blank folder/canvas and I just want that cute "questionnaire" DX to get me going...
But on the whole, CC is smart enough to know when to "rush off and act" and when to "pushback", which is great--and it's no big deal to tell it to pause/stop by adding "what's your thoughts?" or "feel free to pushback" etc to my prompts to make it start a back-and-forth with me (the fact that you say that Plan Mode was really nothing more than a prompt anyways is reassuring).
And yeah, when you're deep in the weeds, you could (and can!) have multiple threads of thought/work going on in the same convo, that stopping and starting a plan in Plan Mode seems like a regression.
I don't really use plan mode, but I do have Fable write a lot of docs, particularly when the work involved is long and multistep, because I may need to do it across multiple sessions.
I am not seeing this and I’m wondering why. Maybe how I plan is different then you? Maybe what we each mean by plan is different? Maybe you trained yourself to write more effective prompts than old you? Maybe the problems you’re solving don’t need planning and the problems I’m solving do? Maybe I’m pickier about solutions?
I think he's very literally referring to /plan mode in Claude Code.
To me, it's about organizing the agent's 'mental model' of your code and your request.
The better the mental model, the less verbose text and planning you need to do. Its mental model is the conversation itself. Skills, rules, whatever is read in, or said by you, all paint a picture that gets put in front of a thinking machine to fulfill your request.
The agent's 'mental model' of your code depends on how well your codebase is structured, how close it is to norms it was trained on (for predictability and better decision-making). And all of this must be organized well into context.
Once you organize your work like this you can focus more on the bigger picture around your work to improve 'flow', reduce churn, and be more productive.
I'm guessing btw, I still use natural language, and planning. I can't keep up with that level of productivity, but I'm fine with what I can do anyways.
I frequently ask claude to write and review org-mode files. I've found it useful to maintain knowledge decisions both for claude and myself. Less so for other members of my team who wont use emacs or (n)vim. But thought I'd share.
To be honest the new behavior is convenient when I just say "X is broken, please fix". But I find myself more and more adding "please do not do anything yet, just tell me what you would do" and please feel free to ask if anything is unclear", especially if is making decisions that are a fork in the road.
I would really like a dry run mode that just disables all external commands from the outside so Claude doesn't proactively go about changing things.
> In Claude Code, all plan mode does is add a little reminder to every user message
I was quite surprised when I learned this (when Claude Code edited a file despite being in plan mode). I had previously assumed "plan mode" was a harness level concern, and restricted what tools could be used. I didn't expect that it was simply an addendum to the prompt.
But I think it gave me some good insight into where the heads are at of Anthropic employees building this. Basically, leaning on pushing everything to the model. That's why alignment is so important: a "sufficiently advanced" model doesn't require any tooling infrastructure around it, and I suppose Claude Code devs are targeting that future. I had previously thought there was more to a harness, but with "auto" mode these days, it seems like there's no desire to build in that direction.
* Scoping discussion - do research and figure out approach (auto mode)
* Planning - take the scoping and convert it to a concrete plan for review (plan mode)
* Implementation - put the plan into action (auto mode)
I find this works really well for my workflow, and it is really easy to trigger each phase because the model has clean boundaries. (This workflow is articulated in my user CLAUDE.md) Plan mode is still useful to me as it forces the model to double-check its plans (I find even with Opus 5.5 it still discovers gaps), and it gives me an opportunity to clear context at a really good spot.
So I would still consider plan mode to be useful. It would make me sad to see it go.
PS- I have a Claude Marketplace directory submission for an MCP server that has been sitting in review hell for six(!) months. I've never received any outcome other than "In Review". I hope I'm not asking too much but would it be possible to put me in touch with someone who might be able to help here. Nobody has ever replied to messages sent to mcp-review@anthropic or the "Anthony at Platform Operations" inquiring about status, and we are getting frustrated
I would have agreed a few months ago before opus 5 but I'm back to plan mode and lots of the old tricks trying to get that insolent model to follow instructions. 5.5 so far is better but certainly doesn't feel under my control (had it go off and read other repos and make changes just yesterday while in "plan mode").
I realize there are already a ton of comments, but I think you're missing the idea of _precision_.
Most people aren't precise when initially describing their problem.
Jumping straight into implementation skips the part where we refine and better define what it is we're trying to do, and think through the implications of those changes.
I suspect "trusting the model" doesn't really work at scale with finite resources.
I m a bit surprised by this as - during the last 2 months - I 've encountered more often models going ahead of what I have asked from them. It's gotten so annoying that now instead of just saying "Let's plan X" I will also say, "Stop once you 've written the plan." because otherwise there is a distinct (~35% anecdotally) probability that the model will go ahead and implementing the plan as well. [This applies to both Anthropic and OpenAI models btw.]
Plan mode is a useful shortcut when I want to have an agent do read-only work without having to worry about giving it appropriate stop conditions.
Plan mode is critical in the beginning, because there's a lot of long tail decisions that haven't been considered, let alone resolved.
But as skills and memory are populated over time, Plan mode isn't as necessary. It becomes simpler to let Claude just build and get something general in place that works, and then refine from there. Auto mode will ask essential questions.
I still use Plan mode for big feature changes, to confirm that I've asked for what I want in the right way. I tend to prompt casually, with only a few specific details. Plan mode helps me see the whole picture before committing. In a few cases, it also helped me decide the feature I asked for was wrong.
except the fact that claude code in plan mode does destructive operations, requiring a hook to tell it not to do stupid things, so its not really a "read only" or "dont code yet" tool, imho.
To be fair, the team also neglected the functionalities of plan mode. It was unusable for since july when they removed the option of running /clear before implementing a plan. Also, there was a period of time during the dark periods of opus 5 and nerfed fable 5, whereby if you didnt explicitly request for code, it went on an on for a few sessions planning and researching in a loop.
I am curious what your take is on skill packages like Obra/Superpowers and grillme and the like. On one hand I find them extremely useful for asking me questions. I wouldn't have thought of myself or that I don't feel I would have necessarily gotten in a structured a format just by having the AI interview me. On the other hand, I am not sure how to tell when they have become obsolete because I'm not sure where to even start with generating evals for such a thing.
Can you fix whatever update has happened that has broken claude code's rendering of a space in the input bar when claude-code is running in a ghostel emacs buffer?
At the moment, instead of moving the cursor along so you can see where the next character will be entered, the cursor remains at the previously entered non-space character until you type another non-space character.
Very distracting. And even claude opus couldn't figure out how to make it behave
> This cut-and-paste workflow felt clunky and made it arduous to work through an idea while keeping track of the current plan.
I mean, it is, and if that were the workflow I’d be very fed up of it by now, but the plan mode in Claude’s VS Code extension has supported select and annotate directly since at the very least early 2026.
oh interesting! I haven’t used vscode for anything meaningful since shortly after leaving GitHub (2023). I generally find it overwhelming but codex added the annotations too and I love it.
When I draft my idea for the implementation of a feature or bug fix, I don't even trust a _human_ to understand what I mean the first time. There are _always_ either errors on my part, or erroneous assumptions on theirs. Everything from "this accounts for X and Y, but not Z which breaks the whole thing" to "this part of the idea directly contradicts what with you said earlier, what do you want to do about it?"
I can't bring myself to trust that an LLM understands what I mean better than any human would, no matter how "good" people claim they are getting.
TFA seems to be advocating for regular old vibecoding. Code now and ask questions later. Which is their choice, and is perhaps even a valid choice in many cases. But at least call it what it is.
Plan mode, for me, is my opportunity to develop and understand my own plan. Sometimes, rarely, Claude demonstrates it misunderstood my intentions by writing a plan to address the wrong problem. But usually it's about me fleshing out the scope and boundary odd the intended changes.
I find that faster than code first, ask questions later. But it takes more time up front.
Coding with plan mode is still vibe coding. No-one outside HN calls it things like "agentic engineering" or whatever. Everyone just says they are vibe coding regardless of how hands on they are.
Whether you use plan mode or not is orthogonal to whether or not you're vibe coding. "Vibe coding" means that you don't read or edit the code; you are running the software and iterating on vibes alone. That's the meaning since the original (and very recent!) coinage and hasn't changed.
If you're reading and editing the code, you're not vibe coding. If you're not reading and editing the code, you're probably vibe coding even if you feel "hands on".
In the before times, and sometimes now even, the way I always coded was by thinking/discussing the work, then hacking up various interesting bits, then throwing it away and doing it the "right" way. Plan mode makes it harder to go back and forth between "hacking" and "thinking" phases.
I'm actively watching understanding slip away from developers, code review getting paired down to no comment checkmarks, and codebases go to bloated messes that nobody can read. Axioms like engineers must understand and take responsibility for the code they ship are getting torn down, and the products coming out are reflecting conway's law, becoming impenetrably obtuse and always "so complex there are no obvious deficiencies" (as opposed to "so simple there are no obvious deficiencies" which used to be the aim).
The one thing plan mode helped is for the humans to get an understanding of the strategy, and be able to poke around and look at the design and architecture. You can achieve this with some self discipline and keeping shorter leashes on agents, but it feels like a losing battle. The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.
> The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.
I agree with this, but the reality is that it's only the result of models empowering devs, and power in good hands amplifies positive results while power in mediocre hands amplifies technical debt.
It's a good time to choose wisely who you work with.
> It's a good time to choose wisely who you work with.
Very true, but this also makes me think what kind of ridiculous obstacle course future hiring process would look like.
In a land where anyone with a pulse can prompt AI to make an app for them - how would future hiring managers and team leads figure out who will drag codebase down with tech debt and who wouldn't?
Probably the same way they have for the last 30 years: poorly, in a dozen different ways, depending on what that particular hiring manager thinks has correlated with successful hiring in the past
I can imagine a code review where you're asked to implement and merge 10 PR in a sample codebase and the codebase is littered with the sorts of mistakes and slop that vibe coders put in when they're careless and you're asked to correct the mess and make it work.
Future hiring will select for people that shipped the most, plain and simple.
The tech debt concerns are much ado about nothing. Use the next model to clean it up, big deal. Code is cheap.
The people that sat around handwringing about tech debt and trying to read every line of LLM code will really struggle to find a job. The profession fundamentally changed, and these people did not catch up.
Of course you shouldn't evaluate people by tokens, you should evaluate them by what they ship. Did it matter to the business? Did it ship fast? Did they build what was needed?
> The people that sat around handwringing about tech debt and trying to read every line of LLM code will really struggle to find a job. The profession fundamentally changed, and these people did not catch up.
I'm inclined to agree - the latest models seem to have crossed a threshold where trying to review by hand becomes akin to a manager insisting on reviewing PRs.
Nothing, the people worrying about this AI tech debt on HN have been worrying about it for years.
The reality is that models just keep getting better and are very good at cleaning up the debt they created.
The engineers I know that were fretting over tech debt in 2024 have just shipped way less value than the engineers I know that just YOLO'd things. The bill never came due. It won't.
Until some catastrophic data loss event or data leak that nobody understands or has any mitigations for. Oops! Guess bankruptcy and prison time are the ultimate uncaught exceptions.
I’m a big proponent of writing simple, understandable code, so I’m playing devil’s advocate here a bit, but: who cares?
A significant (majority?) portion of developers have been shipping JavaScript/node applications for the last decade that contain hundreds of MB to GB of code from god knows where doing god knows what with dependency trees the size of redwoods. It’s not like your average mediocre dev really knew what was going on behind their gluing of frameworks together - at least from what I’ve seen.
If you have remotely competent tech leadership that enforces relatively intelligent patterns (a good one I’ve found is “write everything backend in rust”) you can make AI churn out monstrous amounts of code that… isn’t all that bad? And if you enforce it writing and updating a docs/API.md on every commit/PR you’re probably doing better than 80+% of devs I’ve ever met. Up until a few years ago it wasn’t uncommon to roll up to a new job that was a “legacy” pile of garbage concocted over 20+ years with no comments or API docs and a readme that tells you to ask for help from someone who has been dead for 5 years. At least AI code is full of comments (some of which might even be accurate) and there’s a finite (relatively low!) cost to figuring out “wtf is this doing and how is it doing it”
One is abstraction, one is complexity that _you_ own. Even 100 lines of trad-coded C relies on "hundreds of MB to GB of code from god knows where" in the Linux kernel. The difference is, you can perfectly understand that C, and own it. Then delegate ownership of the rest to Linus. If AI writes 1000 lines of C instead, now there's code no human owns in the world.
I think the main problem with this statement is that different facets of an entity are abstracted at different rates.
AI abstracts effort and cognitive load away from code at a heafty rate, but it doesn’t abstract liability away from code at all.
My business is paid to produce artefacts for which it has liability in the case of error, so we need to do additional work to mitigate and eliminate the liability risk introduced with language models. So far I’ve not found a better way to do that than a plan/act/assert type approach on every feature.
The spec will never be able to capture all the edge cases. The code is what runs. It also doesn't capture all the edge cases (that's why code has bugs) but it does a better job at that than documentation.
Documentation can also become stale easily
I mostly agree with you, but for “Documentation can also become stale easily” … historically I’d 100% agree with you—I often forgot to update separate documentation files, forgot to update comments on related functions, etc—but AI is so much better at automatically catching and fixing this without prompting than I ever was.
That's true, but it's also prone to over documenting. As an example,
for some reason it always tries to include a diagram of the folder structure in the project. I think that's not really useful and it's something that changes so often it's not worth documenting
This, and maintaining the spec is a cat-and-mouse game precisely to cover those edge cases.
Like "if you're using this database, and the engine has these configurations, then do _x_, unless _x'_ and _y'_ are enabled, in which case, do _y_..."
Which, if you're already being THAT specific in your spec, you might as well, idk, write the code yourself?
Because at that point your human language is basically the code and AI is the compiler. A non-deterministic one.
Regardless, your business stakeholders won't understand what's going on anyway (nor should they), so we're back at square zero.
(I wrote Technical and Functional Requirements Documents as a business analyst in college. What's happening now _for most situations_ is more or less the same thing.)
The longer you work in a large engineering organization; the more clear it becomes that no amount of processes, documentation, documentation management systems or training programs can actually transfer the institutional knowledge of the origination to new people other than people-to-people interactions.
Beyond a certain size, the documentation becomes too large to ingest. Below a certain size, it can only contain a fraction of what is needed. If you take any sufficiently engineering project, and give every engineer amnesia; the project will go to shit for a undetermined amount of time, as it takes months or years go build-back the understanding that was lost.
This is clear enough then large companies fire and replace workers randomly to cut costs; a worker that has built up useful knowledge in the origination over a few years is more valuable than three cheap consultants from "low-cost countries" that are fired when the work package is over.
---
AI, looses its memory every time we press "new thread". No spec can bring that back until AIs become able to write and ingest whole books of context without getting confused.
At the same time if the organisation want code fast, it gets code fast.
The revenue for a new feature today is something sure. While the cost associated with supporting such feature will be up to debate in the coming quarters.
As often it is the case, we are moving on a long vs short term trade-off space. And I don't think any experience will generalize
I'm not saying we can get rid of humans (yet). I'm saying that it's not critical to write and review the code yourself anymore, and that the LLM is the new layer of abstraction through which code is written.
I completely agree with everything you’ve said, in the pre-AI world. Now, I’m not so sure.
When I worked at a big tech company on a large codebase (tens of millions of LOC across dozens of repos), it was extremely common to work on something in an area of the codebase I had zero familiarity with, and due to turnover, no one else at the company did either. As you say, docs were frequently missing or out of date.
However, with some effort over a couple hours, I could make a LOT of progress in understanding the history of the code. Every commit and PR were linked to JIRA tickets, most had eng design docs with comments, slack discussions, etc. I could step through the git history and watch the code change, alongside the artifacts of the human discussion and decisions that led to the changes. It wasn’t perfect, but I could make tremendous progress. Now, this was a remote company with pretty strong culture around using JIRA, design docs with review, etc. Probably the biggest gap was meeting transcripts.
Today, an agent can chew through years of history and artifacts on a large codebase in a half hour, documenting as it goes, and have way better understanding of it than I ever will.
It’s true that an agent can’t hold all of that in its head at once without context rot (though this is improving every year), but neither can any human!
If your goal is understanding how to do something, why something was done, why something wasn’t done, etc, and the codebase is large, mature, and extremely well-“artifacted”, I’m not at all convinced that you’re better off asking Bob who has worked on this section of the codebase for a decade than just asking a really good frontier-level agent. Maybe, but it feels like that won’t be true much longer.
> Today, an agent can chew through years of history and artifacts on a large codebase in a half hour, documenting as it goes, and have way better understanding of it than I ever will.
At my last job we worked in a large, but publicly available code base. My experience was that just giving an agent the prompt to look at area X to figure out how it works would use up half or more of the context window. And that was thrown away every time we started again. We had docs, agent generated overviews but the agents still filled their context windows way too quickly to actually really be any use.
I have a hard time finding concrete examples that does not breach contract of dox me. But this issue goes beyond a single large legacy codebase. It can go beyond how the code works into what the code does. You may not even have the code for half of it as it's part of another department.
How should a thing behave may itself be institutional knowledge only Bary that was involved in the design process still remembers. Pointing an agent at a problem in one codebase to write a patch can work quite well. But what do i do when the problem is either in the tool, the database, or the two API layers between? all managed by different teams of office locations?
Exactly. As someone much smarter than me once argued here on HN, the value of a software company is less in the software but much more in the mental map the team has built of the problem and of how the software solves it. Throw the team / mental map away and you'll have a very hard time maintaining the software.
> hundreds of MB to GB of code from god knows where doing god knows what
If you use established libraries then actually the code IS well known to someone (and likely many), even if that's not you. Likely it was built with an actual purpose and with the foresight to not add red herrings to the design.
You can't say any of that for the equivalent amount generated lines. Literally no one knows what it does.
moreover there is not a black and white "code you can read" and "code you cannot read", there is a spectrum between. In this era, IMO if you know which module do what kind of function it should be enough, you don't need to have deep understanding of the file.
What happens when the maintainers lose access to frontier models because of cost, politics or other external factors? What happens if they don't have enough hardware to spin up an open-weights model?
I've seen variants of this play out before AI, so I can tell you: they will inherit a codebase they've never seen before, take forever to ship fixes (forget new features), and they'll either scrap it, completely rewrite it, or, if they're "enterprise" enough, will pay consulting companies literal mountains of cash to make it their problem.
The whole point of writing simple code was to write code that other humans can maintain. If AI is here to stay and becomes economical enough for everyone to use it, then you're right; writing code for other humans is no longer useful. If that doesn't happen though...developers who can/want to still code by hand will be loving life
Then all the more reason to support open weights and the improvements going into them, currently state-sponsored but available to all.
If frontier AI is load-bearing, be thankful you're able to spin up comparable models (face it — 99% of us aren't writing novel software) for a fraction of the cost it would be for a third-party to come in and mop up the mess.
When slop code needs to do something different or new or in a different environment but still stay working on the original things, it’s hard to change, sometimes catastrophically so. When your code has concequences, quality means more people understand it which means more people can maintain it.
Arguments against it are sort of like why have devs on staff at all when you build the thing the first time, or why not outsource everything, or why should I care what my code looks like when the code seems to work?
The cost of tokens is not zero, and the bigger the thing you’re doing the more low quality will cost. When your company is the software, you take on an existential risk based on the software working or not.
Pure AI generated code without human curation is full of bad wordy comments but those comments can mislead, be stale, contain duplicates, and drastically reduce the ability to understand things. I watch teams that still care about code understanding ship good products while those vibe coding in the same org just flounder after the initial burst of features. Some problems only come up after the first 90%, and AI can help you solve them but a big ball of spaghetti is still a big ball of spaghetti.
In many places it's demanded by upper management that devs use AI.. so even if engineers wanted to avoid using it, they would have to meet their quotas.
This is what happens when executives suffer from AI psychosis. They were already impatient, now with AI all they care about is feature velocity.
The faster they can hit that refresh button to see the features, the quicker sales can close the deals for them.
I am reminded of the cartoons where the car kept going faster and faster, the driver kept pushing on the pedal, parts started to fly out, the gauges started to go in circles, smoke started to billow out of everywhere and then boom!
Seems to me that if the AI writes the code, then AI can easily copy the code.
I.e. how hard is it to point an AI at a piece of software and say "AI, copy this"?
Seems like sooner or later copying just becomes a matter of spending enough on tokens.
Seems in that world, all significant software projects get copied. That turns software into a commodity loss leader for other business models or an open source project. Similar to the way Chrome works for Google and the way Firefox works.
Long term, that also just seems like a matter of spending enough tokens. I.e. we have these "red team" AIs that are finding security bugs in very mature software. As AI improves it seems you will be able to just point them at your copy, spend enough tokens and get most of the bugs.
This has been the experience of everyone making decisions in any company without being the one doing the technical work. It's not a novel concept. It's actually the opposite, compared to technical people running companies.
Well they had actual people who understood the code - because they wrote it and had a mental model pretty deep. Now the people owning the code don't understand it.
Won't you be worried if your mechanic didn't understand your car but offloaded it to a robot that made mistakes all the time?
I do try to learn how systems work and remove "abstractions" as much as I can. I do that with coding and my personal life.
Actually baking is a good example, I used to be really bad so I spent time learning. I don't do it every day but now I understand how bread is made.
I bought a 3D printer so I could print parts to fix stuff myself.
I learned to do my own oil changes, I learned how engines work, etc.
My point is that I try to learn more, not less, which is what AI is trying to achieve
It's truly bizarre to read about all these people who just give up on any understanding about what they are working on.
I very regularly use plan mode not to even make a plan of action itself, but to better understand what possible issues might come up when implementing some feature or fixing some bug. And it is quite common for me to fix or rewrite certain findings that AI comes up because its assumptions are not quite right or don't align with overall goal.
And yet so many seem to be perfectly fine leaving all the decisions to AI - even if it's going in the wrong direction. I suppose that's all the people who got into software purely for money or status - never really caring about the actual thing they are working on.
Just like rushing made messes in the before times so too does rushing via LLM. The exact same outcome will happen, but at a far greater velocity and scale than anything we've seen in this industry before. Old Testament, Mr. Mayor, real wrath-of-God type stuff!
"The one thing plan mode helped is for the humans to get an understanding of the strategy"
For me, this phase still happens, but a distinct "plan mode" is unnecessary: I just tell the model, "This is discussion; no code changes yet." and spend hours figuring out what will and will not be done.
I feel the same way but I have successfully refactored some of the early experiments. Our team has settled on targeting a double output from the before times but more ambitious product vision because AI can teach us things we don't know. We actually target 2 days of coding and 3 days of learning with Ai so the increased efficiency allows upskilling rather than just pushing more code
I see this complaint frequently about losing track of what the agents are doing, and I agree you do need to understand your system. But there seems to be this baked in assumption that if you lose track, you now need to manually wade through this massive mess to untangle it and maybe that is impossible. I don't agree.
If you don't understand the codebase, ask the agent to explain it to you. I'm not not kidding. Modern frontier models are fantastic as this - even more so than actually writing the code. It can tell you in words. It can generate architectural diagrams and sequence diagrams. It can write tests and scripts that prove it's assumptions. It can happily refactor so that the system design is aligned with your preferences.
Once you accept this, you can stop worrying so much about it and instead focusing on building the architectures and tools that lets the agents succeed better and faster - so called closed loops or agents prompting agents. Build systems that are more easily verifiable and deterministic so the agent can write very powerful property based tests. Focus more on what and why you are building, how to make sure all external properties are verifiable and leave the internals to the agents. The code is not really for us anymore.
My experience is this is great when the model surfaces something to you. But I'm constantly caught off-guard by things the model didn't volunteer, things I would have quickly stumbled upon if I was working on code the traditional way. The model didn't think it was relevant but I sure do.
In my experience having the agent explain the code doesn’t work very well for real world apps, even the parts written by humans. For example we tried using to generate diagrams, class hierarchies, etc as part of documentation. If you don’t know the code it looks great. If you do, it’s focusing on all the wrong things, missing the mental model, and ignoring lots of important bits. And Claude tends to be extremely verbose to the point of muddling things.
In my experience it's really good at making you feel that you understand things.
Then when you actually dig into the code, there are many things that are not like you'd expect.
When you've experienced that a few times, you stop trusting that the agent gives you the full picture - for good reason.
When I review AI generated code I generally find so many flaws that it makes it hard for me to believe that those who are not reviewing their output are not just fooling themselves. Maybe not all the time, but quite often.
One such recent example was an SSO simulator for a local env. Instead of using a cookie to remember who was logged in, the agent remembered the last log in a variable, assuming the the next requests would come from that login.
This snowballed into our tests, where later agents had created helper tools for working around the SSO simulators statefulness.
Things improve drastically however if you spin up a second session and ask it to adversarially review everything that the first session produces (this goes for everything: not just code, but also design, planning, and explanations).
This works even better if you use models from different families to do so.
I have, on a lark, reached 20 to 30 adversarial sessions a few times. For some tasks, the sessions will just never converge on anything that yet another session won't find fault with, recommending an alternative already ruled out by another session. Even if all the reasoning in between was documented, the new sessions will endlessly claim to find flaws in past reasoning.
We cannot blame just LLM models, it is brains nature to save energy. If agent did tasks consistently good, our brain try to delegate cognitive load to the model to save energy. After consistent use of LLMs anyone can have tendency commit slop just reviewing at high level, this is specially true with busy lifestyle. Also AI generated code do not give dopamine just like solving problems we did before LLMs, we tend to get lazy. Strict discipline is necessary to make good use of LLMs in order to not commit slop and not to make us dumb.
FWIW, our codebase is growing, and the size of each change is also growing, but it's because AI is making us fix all the bugs we'd previously check in because our code long ago surpassed what even our best developers can reason about.
The funny thing is that the AI adopters are in the middle of the bell curve. Our worst devs continue to perform worse than AI yet refuse to use it and our best devs continue to insist AI sucks despite it finding issues in their code and the reviews and designs they've approved.
Sure and sometimes that is what happens. Sometimes it's not. We should be doing a lot of things, but have to triage issues and act pragmatically.
In the case of what I'm currently working on, filing bugs for every issue I found, and then factoring out each fix, and then running each change through the 8 hour ci/cd system, and hoping an unrelated issue doesn't get misattributed to me... No, I'd rather just wrap it up into one coherent refactoring change and be done with it because when I'm done there are several more like it waiting for my attention.
I don't know your exact situation; and there are sometimes genuinely valid reasons for an 8-hour CI/CD system; but ... man, reducing that down to like 20 minutes (which is possible and often common) would pay as many dividends as all the AI stuff that's been added. Man, like ... holy crap @_@
We're uh.. we're a well known and widely used piece of software. There is an insanely large amount of testing that needs to go in to each change across many platforms and configurations and we have a lot of devs hammering that system with changes to test. Even worse now because of AI...
If your company is really interested in doing AI stuff you may be able to sell people on the idea of a "fast agentic CI" where you:
1. Track code coverage from all current test cases in your big 8hr runs.
2. Index the coverage so it goes testcase name -> methods touched.
3. On every commit run an agent which looks at the diffs, tries to predict which tests the code being edited is touching, and run just those tests (or a random subset if there are too many).
Once you have this you can also automatically kick off builds every 8hr that contain all the previously submitted patches. Once that is done, if a test begins to fail, it can automatically bisect the history to find change, and notify the author.
You could also get these benefits with a Bazel like build system which can cache test executions so others don't need to rerun them if the binary is not changed.
Too many comments about AI read like inadvertent admissions of incompetence or dysfunctional projects.
Couple of WTFs that come to mind:
* How many bugs did one have per commit, that commits have to noticeably grow in order to not have those bugs in the first place?
* How does one even do software engineering if the (best) developers can’t reason about the code?
Unfortunately what's happening is people just can't help themselves. Just like an addict reaching for another hit, it's really difficult to make yourself do work when you could just not. I've said it before but LLMs are our cigarettes. It's going to get a lot worse before it gets better.
These comments are so bizarre when we are what, 1-2 years away from AGI?
Like - you really think models won't be able to clean up the tech debt they created!? They are very good at this already. Ask Opus 5.5 to clean up the tech debt from some Opus 4.6 vibe coded app.
Code is cheap now. The most important thing is to ship, ship, ship. If you are handwringing over "tech debt" you have already lost - and you deeply misunderstand how good this technology is getting!
Did I say anything about tech debt? What's the point of "ship, ship, ship" if AI will be able to do it in a couple of years? Don't you see where this is going? We're rapidly losing our ability to think.
The point of "ship, ship, ship" is to cash in (in terms of your career or company) before AGI renders you jobless.
If we get to post-scarcity you never needed the money. But if we end up in some dystopian hellhole where people have no jobs but capitalism still exists, you'll be thankful.
We get it, the future can't come fast enough for you. Code is a commodity and the only important thing is scale. Congratulations on leapfrogging the midwits on the path to the C-suite!
How many people on HN work on mission critical systems? The discussion is not currently constrained to that. I agree, but the exception proves the rule.
The fundamentals of writing correct code have become dramatically cheaper by orders of magnitude. I have been involved in hiring dozens of developers. Fable/Opus 5.5 would be a superior drop-in replacement for perhaps 75% of those hires at roughly 1/100 the cost.
Perhaps for writing generic websites and boilerplate, or early stage prototypes.
I work in embedded systems where it's harder to offload the burden of testing. There is also generally more legacy code, so experienced developers can be costly to replace.
AI makes some things easier, but costs have certainly not fallen by 2 orders of magnitude.
The last couple of weeks, I’ve got rigorous about making the agent refactor old code. We’ve removed, moved, consolidated, reordered loads of cruft. It has made the code much tidier and reduced the chance that the next feature will build on top of cruft.
E.G. over time we’d gained two client-side caches of related server state. This started out as two different parts of the same model, because we couldn’t get all the data we needed from one microservice and had to merge in the client. Over time, more and more features used both caches for different aspects of related processes. At some point one of the microservices changed so as to return all the data in one call. The update to consume that kept both caches, adding code to sync them, because so many parts of the code were using one as a fallback for the other, so they both looked “necessary”. Because they were separate, and “live” sometimes they’d go out of sync after the initial load. Worse: the consumers alternated about which cache was treated as the fallback, making it very hard to see that either might be redundant. Eventually I noticed they were filled by the response to a single call. We all know paying back tech debt never gets prioritised, so I rolled the payback into two feature tasks, and just took longer about them.
My employer expects we use LLMs and provides some budget, but it’s not enough to use even Open4.7 or GLM-5.2 on every task. I do the bulk of my work with Composer 2.5. It’s quite good for “going forward” on smallish tasks and it’s written most of my code this year. It’s possible smarter models would spot these refactorinh opportunities and action them proir to building features or fixing bugs. But I wouldn’t know because I can’t afford it. I’ve never seen even a 4.8 era model spot a refactor and plan to do it prior to a “new build”.
I’m pleased I’ve spotted these trends and started to build the habit of (telling the agent to)“refactor to make the change easier”, but my percieved productivity will go down and I risk the ire of my leaders.
To be fair, it has always been like this. People actively became dependent on the Internet, on StackOverflow, fast Internet connection, cloud services, etc.
At a company I had setup scripts to build our packages, and the CI was running those scripts. Someone more junior (only a few years, not decades) found it strictly superior to remove my scripts and replace them with GitHub actions: people could now know even less about it (as in, no need to know how to copy-paste and adapt the recipe for a new package), but now it depended on GitHub. GitHub is down, nobody can build anything anymore. And it happened once every few weeks, so people would just go have a coffee during the outage.
You know what happened next? That person got promoted for their good work. That was before AI.
With Claude, most tasks do not require discussion, I know exactly what to ask and what approach I want to take but for especially complex functionality with multiple possible options, I will ask it to list the tradeoffs and suggest an approach. It's still kind of 50/50 whether I take its suggestion or not; it's still a bit off in terms of weighing up importance, but it's really good at listing out relevant constraints and tradeoffs. It sometimes misses opportunities but it always sees the tradeoffs and the issues.
I can't speak for Claude, but in Codex, plan mode is very useful for finding serious issues with one's spec. I use it to improve my spec. I keep rerunning plan mode over my updated spec until it stops finding issues. It is only then that I execute the spec.
Lately, I've been using Matt Pocock's “/grill-me” skill more and more, rather than plan mode or a complex, spec-driven set of skills. I made a personal fork of the skill to use Claude’s ask tool, which has a nicer UX.
I used plan mode for two reasons: to review the choices before execution, and to execute with another model (i.e., using the barely documented opusplan feature).
The grill-me skill is much better for reviewing and clarifying choices (and modifying it to use the ask tool makes you go faster). Instead of opusplan, you can explicitly tell Claude to start a subagent with another model to divide the tasks.
Unless I'm going to just do something, I use /wayfinder for almost everything now. It automatically goes into grill mode when appropriate, and it'll suggest just doing the work if it's light enough.
Plan mode isn’t that useful but planning is, I never used it much, but in the agent build mode, I lay down the specs, the architecture, and everything I can think of, and I ask the agent to make the blueprint with the specs that will be used throughout the project, I review them, make the needed edits and revisions, then code and review follow. This is good because you can use the blueprint in any other model or agent, even humans can read it and understand it, sure, sometimes it gets overly verbose but it’s better than nothing.
I'm probably just behind the curve, but I still use plan mode quite a bit in Claude Code. I iterate on the plan repeatedly until it largely agrees with what I want. Once it seems good I ask it to build the plan and open a PR when done.
My only critic of the plan mode is I wish it was easier to see the updates and changes easily in Claude Code as we iterate on the plan. It is wasteful to have to remember what parts I have reviewed and what parts are new (and need another pass). I have thought about fixing this but I also feel the review is the actual thinking (even if ineficient), and so I purposely have not removed it.
I only really used it because the harness was way too trigger happy to start making changes. Even if I just asked a question some times I would come back and it refactored the whole codebase. Now it doesn't seem to do that any more.
I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.
Plan mode is still very useful to me, I just need to read the plan. Even if the model is smarter than me, it’s still working on my project, and thus I need to keep a mental view of the system. My solution has been to ask the assistant for far less than it can do and leave for myself a bigger bunch of the work, so that I’m forced to keep myself up to date on the project. Yeah I’m leaving productivity on the table but… am I really?
I like writing spec files exclusively by hand and just asking the agent to surface questions about it, which I then clarify by editing the spec file further by hand. It keeps the spec file more manageable than having the agent generate the spec file from your conversation.
Using an LLM for any large project shows how useful it is to have data and functions that aren't siloed. That's why CLI apps have had a resurgence: because the LLM can interface with them. Yet the reaction from so many who are deep into LLM development is to build their own little silo: an app.
We really need a better model. One alternative is to have an everything-app: a general purpose tool in which (almost) everything lives. The terminal is one of them. The text editor / word processor is another. (I use Emacs for everything.) In a business context the spreadsheet is probably the best choice.
I still use plan mode, but during plan mode I will spin out sub agents to implement the current draft in a tmp dir and bring back lessons. I feel this keeps me grounded in my original starting point, rather than ending up with a implemention that meanders through the agents own discovery process. It also lets me ask concrete questions about (possible) implementation
472 comments
[ 0.25 ms ] story [ 158 ms ] threadOf course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.
Qwen 3.8 27b is the supervisor
Qwen 3.5 4b are the 6-15 minions it controls
Gemma 4 e4b is the validator for the supervisor.
A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.
What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.
My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.
I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.
This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.
Some people are resource constrained.
- Cold starts impact, context length issues, task lifecycle management
- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).
- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).
This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.
The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.
The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)
The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.
I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.
I also find this "kill Plan mode" push on Twitter odd, because developers have been complaining about AI supposedly killing their jobs, yet they want to take away the main feature that lets them be an active collaborator and participator in the process. Weird.
Maybe it's me (it usually is), but I don't give an LLM small tasks that a human could knock out in a half-day. I give them big tasks, stuff that would take a human a few months to a year - and then I ask it to give me a plan, as a living document, and we spend a good hour or two iterating over the plan.
Then, when its finally at the point where I'm happy with the plan, and I've talked it down from wherever it first wanted to go, or pointed out that we don't need to be all-things-to-all-(wo)men and focus is important, I get it to start going through it.
Likewise, I get it to maintain TODO.md with lists of known bugs, separate lists of future-features, and again, I make a rule that this file must be updated whenever something material changes. I just asked Claude "where are we ?" and got back stuff like:
The ## Still-open detail section lists two items:
- Task #1070: the ported back end doesn't fold offsets into vector loads, so it emits an extra add on 11 files (from 892) at -O3. The output is correct, just longer.
- Task #1080: array sizes must constant-fold. For example, u8 buf[EVSZ * MAXEV] is rejected because size expressions only accept a literal.
This is for 'xc' [1] - a compiler for an Objective-C-like language (but without the excessive []). The language has ARC, blocks and bound-functions/methods named 'block' and 'callback', automatic parsing of DWARF data so you can #use a shared-object, so there's no header files - just read enums/types/functions/methods from the shared object. It's a cross-compiler, runs on mac,windows,linux and creates executables for mac,windows,linux,ios,android,WASM (amongst others). I have a binary running on my iPhone which was written on, and signed on a Linux box - no Apple software used at all. Oh, and it produces code that is very comparable to clang in speed on both arm64 and x86_64.
You can appreciate it's a reasonably large project. It's taken actual months(!) [grin] for me to get working. Months! There's no way I'd approach a problem like this without detailed plans of what I wanted the language to do, where we were going with it
FWIW, "I" wrote blewit.net [2] entirely in xc - both the server back-end (#use <psql> was very useful for binding to Postgres) and the WASM client - which share classes between client and back end, to make it very difficult to get out-of-step between them. No Apache (#use <tls>), no scripting, just a lean-and-mean daemon talking to postgres via valkey (#use <valkey>) - a reddis-alike. Oh yeah, blewit.net has a plan too. Actually it has lots of planning :)
1: https://compile-xc.org/compiler/
2: https://blewit.net/
This has never not worked as expected.
It sounds like the Ant perspective is "you can just ask claude to plan", but at the same time that's a little more tedious than shift+tab
but I think I overfit the interface to how I used to work without agents. when I was at GitHub, I wrote a lot of ADRs, RFCs, and design docs, and I liked that because writing is how I clarify my own thinking. with agents though, I’m often not doing the writing myself. I give the model a rough intent and it fills in a bunch of gaps, and then I get back a long, polished plan containing decisions I didn’t explicitly make.
That’s the part that feels broken to me. The plan can be detailed and technically correct, but still be hard to review because the important bits are buried and feel distant from my own thinking. All the assumptions and tradeoffs and questions may or may not be legitimate, but it’s hard for me to get into flow state and carefully check them.
So I still want the collaboration step. I’m just less convinced that generated prose à la plan mode is the right interface for it.
This tool is a difficult sell. Users _might_ consider an Open Source version, but switching people away from familiar tools is not easy.
I feel like this post is hawking for commercial reasons more than it’s based in reality, especially a reality that takes into account the extremely varied experiences engineers seem to have with agents.
As such I’m not inclined to take it very seriously. More than that, the “X is dead” trope was overdone 10 years ago and I don’t think it’s been long enough to warrant a resurgence.
It's just redundant when you can just collaborate in the "main" mode. There's never been a point to having a separate plan mode.
RE the article: I don't think it's obvious why this process is worth following until you find your time and attention wasted. Conversationally-building is the express train to waste. I'm not sure why you would even be talking to claude if you don't understand what you want to build.
You can use Docker’s sbx or similar VM/containers for that.
If you do use it you might like https://plannotator.ai/
I typically converse with the default model to point the plan in the right direction, then have it iterate with a smarter reviewer to find flaws until the plan is converged.
/s
This is why open harnesses and open models are superior, it's only a matter of time until someone decides that something you use isn't worth maintaining anymore. The reply from the Anthropic employee up above doesn't give me much confidence that Claude users will be able to continue using plan mode forever.
The flexibility and lightweight nature of open harnesses is mainly useful for open models. These models are considerably far behind the frontier and need a lot more steering. Complex harnesses confuse them. I use open harnesses for my OSS model rig because the model needs it.
https://x.com/bcherny/status/2035375125382451648
Also, since sometimes an AI can be a busy beaver, I added instructions to not edit any checked-in files if the prompt contains a question.
I can’t imagine how high that number would go if I got rid of the collaborative style and also the “please don’t run scary commands on random directories without permission” mode.
It’s like this: If you know the right words to say to the LLM, you’ll likely get back the “right words” also. And the more right words up front, the less steering you need to do.
(Having 20yrs of experience writing these user stories and designs help, my “unfair” advantage)
I want it to go and read stuff but also invoke some commands and generate reports and discuss on the results. Latest GPTs in codex plan mode always go.. write a plan. Shocking, i know, but thats not what i want on every turn when im in that mode.
In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.
I still use plan mode in Astra to come up with a plan that I then feed into Fable. I feel like OpenAI models still do better big picture investigation and planning, while Claude is the better software engineer, if that makes any sense.
Of course this could well come down to my own biases and the specific things I’m working on.
I greatly prefer this, since it lets me iterate on the plan with Claude for a while without it repeatedly asking if I’m ready to implement the plan.
Once I’m satisfied, I usually start a fresh session and tell it to implement the plan.
For smaller plans, you don’t need the file. Just ask it to come up with a plan. I don’t recall the last time it just started implementing if I only asked for a plan.
“I want to ...” / “Let’s ...” -> Plan
“Do X” -> Actually Act.
But then again I also have it configured to only ever answer questions instead of inferring them to be instructions (which I’ve seen others do differently).
They’re most useful for broad changes (new features, refactors, etc.) where it’s helpful to avoid breaking changes or unnecessary scope expansion.
The new models are great, but they do more by default, which means I’m finding myself explaining what _not_ to do more often than with previous models (where they’d often end too early).
In my case, the previous plan mode was too ephemeral, and I like having one source of “truth” that sits across context windows without loss/compaction.
Historically, plan mode served two different roles:
1. making the agent’s instructions precise enough to execute 2. helping the human understand what was about to happen
I think #1 is less necessary as agents get better. #2 is going the other direction, it becomes more important as the model is able to do more on its own because larger chunks of work are happening with increasing complexity.
Where I’ve changed my mind is the interface for #2. I increasingly think an interactive, iterative workflow is closer to how people actually build understanding than being handed a long generated document, especially one they didn’t author themselves.
The human-understanding problem is very real though
Also I hope your delivery goes well. My wife (and co-founder) had a challenging delivery and it really put life in to perspective for both of us on a range of issues (how much women's pain is minimized in the health system requiring stronger personal advocacy than I would ever have expected).
As far as plan mode, I still find it essential in keeping agents on track. I build propelcode.app and have a variation on plan mode I still find useful, happy to trade notes on agentic coding if youre interested.
Propelcode looks awesome! So cool to see different people and perspectives shaping this space.
And thank you for the article - it was a good read. I still see folks in my org playing "throw spaghetti at the wall and see what works" and getting frustrated so plan mode (mostly point 2) has been their guardrails almost as much as for the AI.
Can't help but think "Doesn't matter if a machine or a human with (even slightly) different background wrote it", maximizing information flow is maximizing common assumptions and "culture" to only have to communicate a small set of current information for the task at hand. Being a team means having built a joint context so to say. This has always been the purpose of design documents and they always were too big or too small. Because you did not write them, but the others. If you only produce code you think they are the past and useless. If you iterate and your team grows, you start seeing the value in always current docs that are containing just what is not in your everyday culture.
All the best for you and your growing family. I had a similar experience recalibrating my values...
Also for session planning, as in when-can-I-walk-away-from-computer, its nice to know the particular rhythm of initial crunch - ask questions - make plan - do it. Especially with a 5 minute cache timeout.
1. I dont want to have to accept every time Claude touches our DB
2. I'm scared out of my mind it might do something bad to the DB
Plan mode gives me enough confidence that it wont do (2) --> allowing me to give it enough permissions to do (1)
we DO daily snapshotting, so the risk is limited... but still spooks me
What could possibly go wrong with building a data-focused company on a foundation of violating the most fundamental precepts of data management
I was experimenting with a rather complicated backfill operation, were I had a validation script I understand and have Clod come up with the backfill script. I was running against a local prod copy, and it proposed running the actual (unfinished) backfill script against prod.
It didn't have access to the secrets and I also caught the command, but a good reminder that this stuff needs guardrails.
Hallucination not a big deal when it's on the surface layer. But I can't imagine the damage it could do if it hallucinated while building/validating a "load-bearing" component and then continued down that path
For my small-scale sqlite dB, it gets read access, and I encourage it to test modifications by copying it somewhere and writing into that.
Scale-dependent, but I hope to not have to work at a scale where it gets write access to the production DB. That just seems like asking for trouble.
I ended up building out tool an MCP server that I use as a bit of a psuedo harness for Claude. I have a variety of multi-step workflows that are basically micro-skills stacked on top of each other. This helps me make sure that I can get Claude to think in a repeatable and reliable manner.
For coding, I've found that I have a few specific steps that Claude needs to do before I'm comfortable letting it loose:
* It must extensively explore the code base (including certain areas that it misses)
* It must think about what it doesn't know or is making assumptions about
* It MUST scaffold out it's intentions. Essentially, it can write comments, classes, and method stubs - but no actual content. Very much like a spec, but since it's in and alongside other code, it's much easier to identify problems.
* It must spike and validate key assumptions. This, plus the prior step, are the only way I've figured out how to avoid it ending up in a confusion loop. Too often it looks at poor-quality code it's written and thinks it's a long-term solution. By avoiding writing code as much as possible, it knows that it's draft content.
* Only, then can I review it and send it it.
Said MCP server (missing the actual ops): https://github.com/clops-mcp/clops-mcp
https://innerloop.test/breadcrumb (for reference)
Is this what happens when you vibe code long enough?
https://innerloop.works/breadcrumb
What a rookie mistake!
I’ve heard “earns its keep” in only two contexts in my life - the intro to the song “Regulate” and terrible Claude docs
I miss that. It worked really well, and it kept the context clean.
Plus, I usually plan with a more expensive model and guide implementation with a cheaper model (with smaller validation calls back to a more expensive model)
For me it’s actually the opposite, and Claude Code’s plan mode isn’t nearly sufficient. Personally I ask Claude to write down a markdown file with its plan, then review the plan using plannotator, and then go back and forth (most of the time it’s actually the comments that are the problem, not the code).
Then start a fresh session, seed it with the plan, tell Claude to find ambiguities / friction points / oversights, resolve those, and then implement it.
Review once again with plannotator, go back and forth, and then send PR.
Maybe not the “vibe coding” that was once imagined, but this does ensure I am fully aware of the code and architecture, the quality, and this also prevents long term degradation.
I've tried doing the incremental, iterative approach with just Code and it's just not as effective unless you're working on something simple or experimental. Or you're shipping to something non-serious or perpetually beta.
https://plannotator.ai/
5.5 is much closer to Fable so i don't even need it. I am pretty sure it's got Fable's DNA in it.
I really need to find a role where I can do more DX...
So again: You don't need plan mode, auto mode works just fine, there is no difference in the workflows here.
- strategy document
- "sprint" document with technical implementation
- actual implementation
- e2e testing scenarios updates
Every step involves iterating with Claude on it with me in the loop (setting the direction then resolving the "founder questions" as they appear), and importantly a different model for review/code-review, be it Codex (usually, it's great at it) or Antigravity/Gemini (sometimes finds novel things, its precision and recall are abysmal but on the odd occasion it has good accuracy). This iteration on the high-level plan then on the implementation plan is essential to me, and IMHO part of why people are surprised that I tend to get solid results from LLMs. At the very least, it allows me to fill gaps in my own knowledge (primarily front-end development) and be more productive than writing the code myself. I cannot stress enough how nice it is to have a partner in the high-level system design – yes, it often suggests utterly moronic ideas, but the overall experience is still net positive and getting better every quarter.
I wonder if it's just a consequence of a gigantic training set full of comments completely out-of-date with the code, leading to the model considering this "normal"
I now make sure to do a big decommenting pass before every PR.
But I also am starting to just let go and stop caring. It’s not clear to me that it causes problems down the road, it’s easy to strip out en-masse if needed, and in my experience, agents now are really good at read git blame, the commit log, even prior agent transcripts if available to sleuth out when a change was made and why. So yeah, it’s annoying, but the code agents write for me is increasingly never read by a human, so does it matter?
When I read, I skip most comments, especially the larger "Javadoc" style. My brain sees them colored differently in the editor and it doesn't even take mental effort. Then, when I have a question about the code, I look back up for a relevant comment. That doesn't happen very often.
If Claude writes great code and leaves a garbled Claudese-but-accurate comment ... I can read and comprehend (with like 5x the effort of a human comment) ... that's a small price to pay.
(I do have detailed instructions for it on how to comment (or not) but it has not fixed this.)
It’s always “you explain only what but not why” or “this is way too much prose” or “these comments don’t belong here, they should be inline comments” or “this is completely redundant as it’s already obvious from the code”.
I personally still find planning a valuable mental exercise; it's not so different from pre-LLMs and whiteboarding or otherwise taking the time to consciously plan a set of work.
Currently looking for a framework for managing this in a more formal way, and I think it's probably beads, but interested to hear from others.
My main conversation is usually with an orchestrator that hands off work to various (usually cheaper) subagents to plan / review / etc. It has instructions to find the correct model for each task and not to do too much itself so a multi-phase plan automatically gets a fresh subagent for each phase.
anecdotally I've tried the same plan on different branches with both approaches (single agent with subagents implement the full plan vs. my little session per phase app) and then had agents judge the code quality and my approach worked better and saves tokens. horses for courses.
https://ljtn.github.io/epiq/
Might be worth a look if you’re evaluating alternatives to Beads.
Great work man.
I have some older projects that use beads (I still run an old version without dolt that's imho pretty good overall) but lately with Fable also have a few newer projects where I just have the agent write docs and keep a worklog with the what/why/decisions etc. (I think I read it here on HN somewhere and figured I'd give that a try.)
The latter seems to work pretty well for now (slightly better than beads) but I'm always looking for ways to improve it. This could be an interesting replacement.
“ayy lmao”
https://aventasks.dev/
I asked fable to look at my interaction patterns and clearly stated my frustrations and the problems I wanted solved, and it designed a simple process to track things in git and built a couple simple session hook skills. It’s pretty lightweight and I’ve been very happy with it for a couple months.
I get a long way using models like Opus to make a plan of action and a bunch of tasks, and then using Deepseek to implement that plan of action. Saves a bunch of money and is fast.
- I have a record of work done and work to be done that helps _me_ when I come back to the project after several months. It’s committed and lives with the code.
- when a task inevitably ends up more complicated than I thought, I can in that session break it up
- I initiate sessions from multiple computers, so things stay in sync (through git)
- I also have a “tooling” repo that builds out some views of the work and hosts it for me to see when I’m on my phone.
- The hooks let the agent manage all of the workflow/task management, so there’s very little management overhead for me.
I rejected beads and JIRA. I wanted something more lightweight.
I also sometimes rewind the conversation after going off on a tangent, sometimes not if I want the agent to have the sorts of things I’m thinking about in context. I probably justify keeping too much crud in context than I should
[0]: https://github.com/bensyverson/jobs
How do I use plannotator to review an arbitrary markdown file? It always opens the Claude Code plan file for me.
Then telling Claude to work on a document, the instruction is kept to its core.
Now when bcherny explicitly mentioned that it merely adds a single line - it explains why I don't need it.
What may be concerning about "super plan" mode from the creators (or a skill, for that matter) - is that tuning the amount of effort, and how much deep to dig - may become too hard, as it will interfere with several embedded paragraphs explaining what to do, how to do, where to do, etc'.
What I do look for is even better plannotator ability to track changes, combining historical comments (like Google docs), and git blame of several "generations" before current reviewed doc.
Roughly speaking, I'd be happy if plannotator would persist something similar to github PR reviews combined with Google docs comments & suggestions.
Note that this is only really necessary for complex work that I don't know yet what the best way to do it is.
I've tried doing it your way as well, but there was just too much fiddling about with writing the plan somewhere, then having another session rebuild their context with whatever info is in the plan. It really didn't result in better output for me.
Currently 9 times out of 10 I just say to the model: xyz is the problem/bug/feature, fix it. Since about Fable and Opus 5, this is more than enough. Opus 5.5 (and previously Fable 5.1) got even better at this. However, this is in a codebase where there are already a few hundred thousand lines of code for the model to look at to see how we generally attack things in our codebase.
Claude Codes plan mode I never use anymore, it was useful a few months ago because the models had a tendency to just start doing work and forget I specifically told them not to. But the UX is just annoying and the models now do adhere when I tell them not to change anything.
Plan mode ensures I'm spending fewer tokens on the code-test loop, and more on the arch/design, and allows me to keep appraised of what's going on, while planning for future changes better.
Maybe folks who don't need planning, don't have as much concern for the details, and are happy enough with just evaluation of if it works or not.
I also found that having the design reviewed by multiple agents has very little marginal value. The review agent will always find something to improve, but mostly it’s just nit and not anything super important.
I used to let Claude just upload the html design doc to Claude artifacts for me to review. Recently I switched to codex and started to use my own tool https://github.com/hyperlogue/r3 to complete this workflow.
seems like plan mode could turn off some tools, even if it doesn't change the set offered to the model, the ones that they have which would mutate your codebase could just not work with an error message, and plan mode could change permissions in the security approval prompt for "auto"
anyway, isnt the right way to know if plan mode helps or not, to run an experiment? we're all guessing unless we have data
read only agent mode sounds straightforward and useful to me
There’s also a “ask” mode which is read only. Both are enforced by the permissions model, not just a system prompt instruction. I’ve seen the model “forget” and try to start coding - it bounces off a hard permissions failure and that “reminds” it that it’s in plan mode.
I'm experimenting just like everyone else, but this is my process right now:
- Quick prototype
- Figure out the language of your app (what terms you want to use for things, what your UI design language will be, etc) and spec that, so you can use words consistently with the agent. You need to be able to describe the things you want well and consistently.
- Keep prototyping. Let the agent write unit tests along the way. Lock down behaviour you like, keep track of those things in a document.
- At some point your idea of the real architecture comes into focus, from actual use cases -- avoids the over-abstracting right away trap.
- Refactoring is cheap with tests, so start refactoring into the architecture you want.
- Your architecture won't necessarily be what would be best for a human, but it will be pretty close.
- Keep relentlessly iterating on small work.
- Things that were expensive before aren't that expensive now -- integrating a library, changing from one library to another, trying out a few architectural refactors, trying out different performance optimizations, etc. That stuff is all 'throw it there and see what sticks' now, so don't be afraid to try stuff which felt big before.
I feel like 'front loading' too much is just the wrong approach. You might feel like you're sitting there 'babysitting the agent'; but that's just what the hard part of the work (hard as in 'zjust slogging through it', not as in 'conceptually complex') looks like now. Your code is much more like clay.
Atleast that's how I'm thinking about it so far, but I'm not working on large sprawling systems that I imagine would need more pre-planning.
For who? The more control you hand over to the AI and let it think for you with no supervision, the better it is for Anthropic
But given that running an agent us cheaper than the cost of waiting for a slot to assemble the team to talk about a change (isn’t it always?), why wait with running the agent?
I propose updating the spec then do the implementation. This will most likely show that a few assumptions were wrong forcing some major or minor updates to the spec. Work through those and then let your team review the spec change together with testing the next iteration of what what’s build.
I find that the code is generally in a better place proportionate to the amount of SDD I actually do. But it's just a matter of where and when I want to spend my time.
I have one session define a task, and provide a formal specification plus context in a "cover letter."
The session B, in plan mode, produces the plan back.
Session one reviews the plan and clears it, ratifying portions and often specifying specific changes.
Session one then executes.
What has been striking to me in this approach is that even with two instances of the same model (currently Opus 5.5), there are regularly corrections made. I use "project chat" for session A and Code for session B atm; it is very typical that Code finds and corrects details or oversights in the task spec; it is also typical (though less so with 5.5) that session A (chat) pushes back or clarifies things Code doesn't have the context for.
I have been afraid to open up the potential of negotiation beyond what this is costing as it is. But I am also afraid to simply skip the formalisms, because of the consistent correction that occurs in this back-and-forth.
Each component of the pattern is schematized, generated from a template, and validated, to keep things tight.
Lots of tokens! But I trust this process far more than "just typing" :)
For the record, this isn’t unique to Claude. ChatGPT and Gemini do the same, each with its own quirks. ChatGPT got extra credit for being the only one who allowed the function to also take a CA file for server authentication.
Don’t get me wrong: LLMs are the future (maybe even the present) of software development but I think there’s some way to go before they can be entirely hands-off in some areas. I still find myself having to course correct designs and plan mode helps me with that.
And of course, thank you for your work on Claude. :)
But there is still a ceiling above which it is necessary to "preload" the context window before starting to call tools and get into the meat of the work. You want to establish domain language (especially with Claude models which otherwise will invent their own, and it will be inscrutable) and key requirements and assumptions. You want to do a Q&A iteration cycle with the LLM. You definitely should do a sanity check that the LLM actually "understands" what you were trying to achieve, and then make sure that understanding is coherently and plainly stated in the prompt. All of that seems to be necessary still for just about any serious task, if you actually care about the quality of the results and/or don't want to burn hundreds of thousands of tokens on flailing around to get to a good quality result.
So no, you don't "need" plan mode. But you do still need to do all of the things you would do with plan mode.
> In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet"
Can't help but think if plan mode isn't useful as you say because it's implementation is lacking in claude code, hypothetically speaking.
What I can say is that with other harnesses plan mode helps stabilize my workflows. Actually synthesizing code is only part of the process, lots involved in taking a work-item to production end-to-end and plan mode helps give this flow structure. More than that it's an opportunity to regroup before committing to changes. It slows down the process to a rythim that's sustainable and smooth, which ends up speeding up the process.
So if plan mode in claude was designed to speed up code churn, while oh my pi for instance designed plan mode to be strategic, that might account for the different perceptions here.
And it's not to say you should force yourself to use it, but if you are planning on cutting this mode off the loop just beware of the possible side effects.
1. I want to know whats going to happen, at least at a high level, before changes are actually made. 2. Plan mode helps me flesh out the missing details of my plan before being mid-execution 3. In situations where I have a limited budget for AI usage I will often times use a high powered model like Opus 5.5 or Fable to make a detailed plan, then scale down to a cheaper model for implementation. I feel like this saves cost in the end.
I get plan mode is basically just a small hidden prompt. I get that I can basically just preface my prompts with "make a plan only, don't make actual changes." Maybe this is just a UX trick, but it works well for my brain.
I just start my day writing about 20 queues /goal prompts and then check the work at the end of the day. It’s almost always right!
Now that I’m frequently designing and delegating day/week scale features, the flow has to change; having the agent go off and build a spike can be a quicker way of us understanding the design space and constraints (especially in a huge codebase). I still have the agent write and update a spec doc as I go, but it’s not waterfall anymore.
At least for my kinesthetic learning mode a rough code PR stack is usually way better than a plan doc anyway, and tokens are cheap enough (vs my time) that going further than just a plan is often cost-effective overall.
The dream of course is (say it with me) loops, but that doesn’t tend to work for me on new features often.
But with GPT 5.6 Sol, I'm still finding that the model makes conceptual mistakes, or gets edge cases wrong, or assumes incorrectly (making an ass out of both user and model). In many cases, I need to at least refine the proposed approach, or amend, correct, or flat out just stop and start over. Not planning and catching these errors, and just letting the agents code their code, would mean I'd have to rollback and redo many times. What a waste!
For a current project, which is ~33k lines of code, I'm also finding that I know the codebase better than the model, and that's vital at the planning stages too. If I wasn't in the planning loop, the model would have reinvented various wheels a few times over. How much spaghetti do you want with your code?
As always, I may simply be doing this wrong. But I'm personally not convinced that the plan is dead, or that I want the plan to be dead. Planning is also good for me -- it keeps me thinking about the code, prompting better, guiding the model better.
If I'm no longer on top of the codebase, then at some point my prompts will devolve to "Do the thing with the thing, that does thing". And I don't want that.
On the other hand, for a low effort hobby project: just do the thing.
Also if you are working in a heavily vibe-coded codebase, as Claude Code reportedly is, it's not that surprising if the human doesn't really understand it or have anything useful to add in a collaborative context.
- Most people suck at planning anyway
- LLM still don't give you a way to verify and understand to iterate.. you have to ask and then formulate and way so people barely do it, they just trust the vibe
IMO the current successor to plan mode should be the harness knowing when to tell the use "ok here's our overall current state in a simple diagram", auromatically
However how about decisions? Do we expect the model to read our minds, just assume the best practice will be followed and that’s what the user want? Plan mode solves those, what is that am I missing?
It could still make the tools into no-ops or disabled if it actually tries to use them, without changing the context history at all.
Also this is a way less removed process than I want nothing to do with. The more removed I am from the process the more I hate my job, get burned out and genuinely wish that Anthropic never existed.
Even if it could "just know" or infer my intent. It wouldnt be desirable.
The Anthropic employee literally told you plan mode is a "we're still planning!!!" at the end of the prompt. Plan mode is useless because you can literally type "dont write code yet" and get the same effect, not because you should never plan in general.
For example, there are no shortage of web based uis for pi and they are all cool but I wanted a deep integration between artifacts and how I want to collaborate with the agent on them. So i built my own ui that mirrors what the pi tui sees. It’s chat based but has a deep integration with a GitHub style code review UI so I can leave review and comments whenever I want. Every agent message renders nicely in an annotatable markdown viewer so I don’t need an ask question tool and can more naturally get the agent on the same page as me. I want it to feel like I’m working with a colleague.
I don’t think these features are too unique but having full control over the experience is really nice. Flexibility over model provider, can tailor it to my work’s dev stack, and don’t need to worry about anyone breaking it with a million updates everyday.
Your perfect workflow can be realized in a day or so. You just need to go make it happen.
I've been building https://crit.md to keep that back and forth with agents - GitHub-esque interface and have agents respond to my feedback, iterating until I'm happy.
Admittedly like many others, I use it a lot less for actual plans now, models are indeed getting really good at just getting it.
I wodner what this product space will look like a year from now. Reviews are already dying.
So basically you don't know what the fuck you're talking about. Not everyone is employed to burn money.
I also tell it to write deviations and rename plans accordingly once done.
That way I keep the codebases I have to or enjoy to work on in my head and don't become too dependant on any provider or on stochastic parrots in general.
When a test case fails, the relevant part of the plan is surfaced in the error. I find this helps Claude stay on track for longer - I've been able to do 12h most times and even up to 48h unattended (11h of API time) with good results.
Then whenever I do check in, I ask it to update the HTML with current state in an append only fashion (sort of like it's writing a blog), and then based on that, we iterate on the end to end test (I think of it as a "test harness") - update the test cases and error messages.
I've been able to build some truly large projects this way, both greenfield and up to spec (for example, a video game I've always wanted to play), and brownfield while staying within the conventions and design of the codebase, and with very little attention required on my part.
That's when I need to learn about implied patterns, do's and don'ts; not from theory but in the context of my own project. Unfortunately, the model tends to keep implicit knowledge implicit. But I can ask during planning.
The feature did get less useful over time when the model started babbling in newspeak more and more. When it threw a thousand words at me even in concise mode.
So I don't want to let the toolmakers off the hook here. There's a lot to win that would make plan mode much much better without changing plan mode itself.
Aider, Cline and many other agents had plan mode before Claude Code existed.
You don't need to actually change the tool schema or break the prompt cache to do that. In the tool itself you could just check if it's in plan mode and reject the tool call...
I find myself endlessly ctrl+c ing claude now as it flies off doing deep first principles analysis to work out how to find a thing it isn't sure about but I know the answer. Being able to give it that answer without needing to ctrl c would be a massive improvement
I haven't reached for Plan Mode in awhile--maybe, on some blank folder/canvas and I just want that cute "questionnaire" DX to get me going...
But on the whole, CC is smart enough to know when to "rush off and act" and when to "pushback", which is great--and it's no big deal to tell it to pause/stop by adding "what's your thoughts?" or "feel free to pushback" etc to my prompts to make it start a back-and-forth with me (the fact that you say that Plan Mode was really nothing more than a prompt anyways is reassuring).
And yeah, when you're deep in the weeds, you could (and can!) have multiple threads of thought/work going on in the same convo, that stopping and starting a plan in Plan Mode seems like a regression.
(happily using Claude Code Opus 5.5 on High rn)
# YOLO mode
To me, it's about organizing the agent's 'mental model' of your code and your request.
The better the mental model, the less verbose text and planning you need to do. Its mental model is the conversation itself. Skills, rules, whatever is read in, or said by you, all paint a picture that gets put in front of a thinking machine to fulfill your request.
The agent's 'mental model' of your code depends on how well your codebase is structured, how close it is to norms it was trained on (for predictability and better decision-making). And all of this must be organized well into context.
Once you organize your work like this you can focus more on the bigger picture around your work to improve 'flow', reduce churn, and be more productive.
I'm guessing btw, I still use natural language, and planning. I can't keep up with that level of productivity, but I'm fine with what I can do anyways.
I would really like a dry run mode that just disables all external commands from the outside so Claude doesn't proactively go about changing things.
I was quite surprised when I learned this (when Claude Code edited a file despite being in plan mode). I had previously assumed "plan mode" was a harness level concern, and restricted what tools could be used. I didn't expect that it was simply an addendum to the prompt.
But I think it gave me some good insight into where the heads are at of Anthropic employees building this. Basically, leaning on pushing everything to the model. That's why alignment is so important: a "sufficiently advanced" model doesn't require any tooling infrastructure around it, and I suppose Claude Code devs are targeting that future. I had previously thought there was more to a harness, but with "auto" mode these days, it seems like there's no desire to build in that direction.
* Scoping discussion - do research and figure out approach (auto mode)
* Planning - take the scoping and convert it to a concrete plan for review (plan mode)
* Implementation - put the plan into action (auto mode)
I find this works really well for my workflow, and it is really easy to trigger each phase because the model has clean boundaries. (This workflow is articulated in my user CLAUDE.md) Plan mode is still useful to me as it forces the model to double-check its plans (I find even with Opus 5.5 it still discovers gaps), and it gives me an opportunity to clear context at a really good spot.
So I would still consider plan mode to be useful. It would make me sad to see it go.
PS- I have a Claude Marketplace directory submission for an MCP server that has been sitting in review hell for six(!) months. I've never received any outcome other than "In Review". I hope I'm not asking too much but would it be possible to put me in touch with someone who might be able to help here. Nobody has ever replied to messages sent to mcp-review@anthropic or the "Anthony at Platform Operations" inquiring about status, and we are getting frustrated
Most people aren't precise when initially describing their problem.
Jumping straight into implementation skips the part where we refine and better define what it is we're trying to do, and think through the implications of those changes.
I suspect "trusting the model" doesn't really work at scale with finite resources.
Plan mode is a useful shortcut when I want to have an agent do read-only work without having to worry about giving it appropriate stop conditions.
But as skills and memory are populated over time, Plan mode isn't as necessary. It becomes simpler to let Claude just build and get something general in place that works, and then refine from there. Auto mode will ask essential questions.
I still use Plan mode for big feature changes, to confirm that I've asked for what I want in the right way. I tend to prompt casually, with only a few specific details. Plan mode helps me see the whole picture before committing. In a few cases, it also helped me decide the feature I asked for was wrong.
In case anyone else is interested, the skill is public: https://github.com/Mudlet/Mudlet/blob/development/docs/demo-...
Can you fix whatever update has happened that has broken claude code's rendering of a space in the input bar when claude-code is running in a ghostel emacs buffer?
At the moment, instead of moving the cursor along so you can see where the next character will be entered, the cursor remains at the previously entered non-space character until you type another non-space character.
Very distracting. And even claude opus couldn't figure out how to make it behave
I mean, it is, and if that were the workflow I’d be very fed up of it by now, but the plan mode in Claude’s VS Code extension has supported select and annotate directly since at the very least early 2026.
I can't bring myself to trust that an LLM understands what I mean better than any human would, no matter how "good" people claim they are getting.
TFA seems to be advocating for regular old vibecoding. Code now and ask questions later. Which is their choice, and is perhaps even a valid choice in many cases. But at least call it what it is.
I find that faster than code first, ask questions later. But it takes more time up front.
If you're reading and editing the code, you're not vibe coding. If you're not reading and editing the code, you're probably vibe coding even if you feel "hands on".
The one thing plan mode helped is for the humans to get an understanding of the strategy, and be able to poke around and look at the design and architecture. You can achieve this with some self discipline and keeping shorter leashes on agents, but it feels like a losing battle. The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.
I agree with this, but the reality is that it's only the result of models empowering devs, and power in good hands amplifies positive results while power in mediocre hands amplifies technical debt.
It's a good time to choose wisely who you work with.
Very true, but this also makes me think what kind of ridiculous obstacle course future hiring process would look like.
In a land where anyone with a pulse can prompt AI to make an app for them - how would future hiring managers and team leads figure out who will drag codebase down with tech debt and who wouldn't?
The tech debt concerns are much ado about nothing. Use the next model to clean it up, big deal. Code is cheap.
The people that sat around handwringing about tech debt and trying to read every line of LLM code will really struggle to find a job. The profession fundamentally changed, and these people did not catch up.
I think there has to be a better rubric by which person would be evaluated, not just amount of tokens they managed to waste.
I'm inclined to agree - the latest models seem to have crossed a threshold where trying to review by hand becomes akin to a manager insisting on reviewing PRs.
The reality is that models just keep getting better and are very good at cleaning up the debt they created.
The engineers I know that were fretting over tech debt in 2024 have just shipped way less value than the engineers I know that just YOLO'd things. The bill never came due. It won't.
> The "tech debt" bill never came due.
Companies paying $200k a day for coding models to churn on what the coding models are messing up is one thing.
I’m not even worrying about tech debt, I’m talking full on defects, production incidents, security holes, and reputational damage.
A significant (majority?) portion of developers have been shipping JavaScript/node applications for the last decade that contain hundreds of MB to GB of code from god knows where doing god knows what with dependency trees the size of redwoods. It’s not like your average mediocre dev really knew what was going on behind their gluing of frameworks together - at least from what I’ve seen.
If you have remotely competent tech leadership that enforces relatively intelligent patterns (a good one I’ve found is “write everything backend in rust”) you can make AI churn out monstrous amounts of code that… isn’t all that bad? And if you enforce it writing and updating a docs/API.md on every commit/PR you’re probably doing better than 80+% of devs I’ve ever met. Up until a few years ago it wasn’t uncommon to roll up to a new job that was a “legacy” pile of garbage concocted over 20+ years with no comments or API docs and a readme that tells you to ask for help from someone who has been dead for 5 years. At least AI code is full of comments (some of which might even be accurate) and there’s a finite (relatively low!) cost to figuring out “wtf is this doing and how is it doing it”
Code is just another abstraction.
You can give the AI the spec and own the spec instead of the code.
Eventually, the spec too will be an abstraction.
AI abstracts effort and cognitive load away from code at a heafty rate, but it doesn’t abstract liability away from code at all.
My business is paid to produce artefacts for which it has liability in the case of error, so we need to do additional work to mitigate and eliminate the liability risk introduced with language models. So far I’ve not found a better way to do that than a plan/act/assert type approach on every feature.
Like "if you're using this database, and the engine has these configurations, then do _x_, unless _x'_ and _y'_ are enabled, in which case, do _y_..."
Which, if you're already being THAT specific in your spec, you might as well, idk, write the code yourself?
Because at that point your human language is basically the code and AI is the compiler. A non-deterministic one.
Regardless, your business stakeholders won't understand what's going on anyway (nor should they), so we're back at square zero.
(I wrote Technical and Functional Requirements Documents as a business analyst in college. What's happening now _for most situations_ is more or less the same thing.)
Beyond a certain size, the documentation becomes too large to ingest. Below a certain size, it can only contain a fraction of what is needed. If you take any sufficiently engineering project, and give every engineer amnesia; the project will go to shit for a undetermined amount of time, as it takes months or years go build-back the understanding that was lost.
This is clear enough then large companies fire and replace workers randomly to cut costs; a worker that has built up useful knowledge in the origination over a few years is more valuable than three cheap consultants from "low-cost countries" that are fired when the work package is over.
---
AI, looses its memory every time we press "new thread". No spec can bring that back until AIs become able to write and ingest whole books of context without getting confused.
The revenue for a new feature today is something sure. While the cost associated with supporting such feature will be up to debate in the coming quarters.
As often it is the case, we are moving on a long vs short term trade-off space. And I don't think any experience will generalize
When I worked at a big tech company on a large codebase (tens of millions of LOC across dozens of repos), it was extremely common to work on something in an area of the codebase I had zero familiarity with, and due to turnover, no one else at the company did either. As you say, docs were frequently missing or out of date.
However, with some effort over a couple hours, I could make a LOT of progress in understanding the history of the code. Every commit and PR were linked to JIRA tickets, most had eng design docs with comments, slack discussions, etc. I could step through the git history and watch the code change, alongside the artifacts of the human discussion and decisions that led to the changes. It wasn’t perfect, but I could make tremendous progress. Now, this was a remote company with pretty strong culture around using JIRA, design docs with review, etc. Probably the biggest gap was meeting transcripts.
Today, an agent can chew through years of history and artifacts on a large codebase in a half hour, documenting as it goes, and have way better understanding of it than I ever will.
It’s true that an agent can’t hold all of that in its head at once without context rot (though this is improving every year), but neither can any human!
If your goal is understanding how to do something, why something was done, why something wasn’t done, etc, and the codebase is large, mature, and extremely well-“artifacted”, I’m not at all convinced that you’re better off asking Bob who has worked on this section of the codebase for a decade than just asking a really good frontier-level agent. Maybe, but it feels like that won’t be true much longer.
At my last job we worked in a large, but publicly available code base. My experience was that just giving an agent the prompt to look at area X to figure out how it works would use up half or more of the context window. And that was thrown away every time we started again. We had docs, agent generated overviews but the agents still filled their context windows way too quickly to actually really be any use.
How should a thing behave may itself be institutional knowledge only Bary that was involved in the design process still remembers. Pointing an agent at a problem in one codebase to write a patch can work quite well. But what do i do when the problem is either in the tool, the database, or the two API layers between? all managed by different teams of office locations?
Bary would know.
and there is a community of developers who care about the quality of those libraries and take the weight of their responsibility seriously
If you use established libraries then actually the code IS well known to someone (and likely many), even if that's not you. Likely it was built with an actual purpose and with the foresight to not add red herrings to the design.
You can't say any of that for the equivalent amount generated lines. Literally no one knows what it does.
What happens when the maintainers lose access to frontier models because of cost, politics or other external factors? What happens if they don't have enough hardware to spin up an open-weights model?
I've seen variants of this play out before AI, so I can tell you: they will inherit a codebase they've never seen before, take forever to ship fixes (forget new features), and they'll either scrap it, completely rewrite it, or, if they're "enterprise" enough, will pay consulting companies literal mountains of cash to make it their problem.
The whole point of writing simple code was to write code that other humans can maintain. If AI is here to stay and becomes economical enough for everyone to use it, then you're right; writing code for other humans is no longer useful. If that doesn't happen though...developers who can/want to still code by hand will be loving life
If frontier AI is load-bearing, be thankful you're able to spin up comparable models (face it — 99% of us aren't writing novel software) for a fraction of the cost it would be for a third-party to come in and mop up the mess.
Arguments against it are sort of like why have devs on staff at all when you build the thing the first time, or why not outsource everything, or why should I care what my code looks like when the code seems to work?
The cost of tokens is not zero, and the bigger the thing you’re doing the more low quality will cost. When your company is the software, you take on an existential risk based on the software working or not.
Pure AI generated code without human curation is full of bad wordy comments but those comments can mislead, be stale, contain duplicates, and drastically reduce the ability to understand things. I watch teams that still care about code understanding ship good products while those vibe coding in the same org just flounder after the initial burst of features. Some problems only come up after the first 90%, and AI can help you solve them but a big ball of spaghetti is still a big ball of spaghetti.
This is what happens when executives suffer from AI psychosis. They were already impatient, now with AI all they care about is feature velocity.
The faster they can hit that refresh button to see the features, the quicker sales can close the deals for them.
AI has basically sold them to wet dream.
I guess we just wait for the boom.
I.e. how hard is it to point an AI at a piece of software and say "AI, copy this"?
Seems like sooner or later copying just becomes a matter of spending enough on tokens.
Seems in that world, all significant software projects get copied. That turns software into a commodity loss leader for other business models or an open source project. Similar to the way Chrome works for Google and the way Firefox works.
Won't you be worried if your mechanic didn't understand your car but offloaded it to a robot that made mistakes all the time?
I see it kind of like baking/cooking. Do you bake your bread from scratch? Do you grow your own wheat and grist your own flour?
I think over reliance on it or not even trying to understand what is happening is a big problem to be sure, but it's certainly not a new problem.
Actually baking is a good example, I used to be really bad so I spent time learning. I don't do it every day but now I understand how bread is made. I bought a 3D printer so I could print parts to fix stuff myself. I learned to do my own oil changes, I learned how engines work, etc.
My point is that I try to learn more, not less, which is what AI is trying to achieve
And I would understand how to bake bread from scratch.
Now in the entire chain, we will get to a point where no one knows anything.
I very regularly use plan mode not to even make a plan of action itself, but to better understand what possible issues might come up when implementing some feature or fixing some bug. And it is quite common for me to fix or rewrite certain findings that AI comes up because its assumptions are not quite right or don't align with overall goal.
And yet so many seem to be perfectly fine leaving all the decisions to AI - even if it's going in the wrong direction. I suppose that's all the people who got into software purely for money or status - never really caring about the actual thing they are working on.
For me, this phase still happens, but a distinct "plan mode" is unnecessary: I just tell the model, "This is discussion; no code changes yet." and spend hours figuring out what will and will not be done.
vs
Press "Tab"
If you don't understand the codebase, ask the agent to explain it to you. I'm not not kidding. Modern frontier models are fantastic as this - even more so than actually writing the code. It can tell you in words. It can generate architectural diagrams and sequence diagrams. It can write tests and scripts that prove it's assumptions. It can happily refactor so that the system design is aligned with your preferences.
Once you accept this, you can stop worrying so much about it and instead focusing on building the architectures and tools that lets the agents succeed better and faster - so called closed loops or agents prompting agents. Build systems that are more easily verifiable and deterministic so the agent can write very powerful property based tests. Focus more on what and why you are building, how to make sure all external properties are verifiable and leave the internals to the agents. The code is not really for us anymore.
Then when you actually dig into the code, there are many things that are not like you'd expect.
When you've experienced that a few times, you stop trusting that the agent gives you the full picture - for good reason.
When I review AI generated code I generally find so many flaws that it makes it hard for me to believe that those who are not reviewing their output are not just fooling themselves. Maybe not all the time, but quite often.
One such recent example was an SSO simulator for a local env. Instead of using a cookie to remember who was logged in, the agent remembered the last log in a variable, assuming the the next requests would come from that login.
This snowballed into our tests, where later agents had created helper tools for working around the SSO simulators statefulness.
Things improve drastically however if you spin up a second session and ask it to adversarially review everything that the first session produces (this goes for everything: not just code, but also design, planning, and explanations).
This works even better if you use models from different families to do so.
The funny thing is that the AI adopters are in the middle of the bell curve. Our worst devs continue to perform worse than AI yet refuse to use it and our best devs continue to insist AI sucks despite it finding issues in their code and the reviews and designs they've approved.
Shouldn’t those “fixing bugs we gained in the past” be their own MR that can be read, reasoned about and have evaluated test coverage?
In the case of what I'm currently working on, filing bugs for every issue I found, and then factoring out each fix, and then running each change through the 8 hour ci/cd system, and hoping an unrelated issue doesn't get misattributed to me... No, I'd rather just wrap it up into one coherent refactoring change and be done with it because when I'm done there are several more like it waiting for my attention.
1. Track code coverage from all current test cases in your big 8hr runs.
2. Index the coverage so it goes testcase name -> methods touched.
3. On every commit run an agent which looks at the diffs, tries to predict which tests the code being edited is touching, and run just those tests (or a random subset if there are too many).
Once you have this you can also automatically kick off builds every 8hr that contain all the previously submitted patches. Once that is done, if a test begins to fail, it can automatically bisect the history to find change, and notify the author.
You could also get these benefits with a Bazel like build system which can cache test executions so others don't need to rerun them if the binary is not changed.
Couple of WTFs that come to mind: * How many bugs did one have per commit, that commits have to noticeably grow in order to not have those bugs in the first place? * How does one even do software engineering if the (best) developers can’t reason about the code?
Software engineering is possible but largely a myth in practice.
The quote is "so simple that there are obviously no deficiencies"
Like - you really think models won't be able to clean up the tech debt they created!? They are very good at this already. Ask Opus 5.5 to clean up the tech debt from some Opus 4.6 vibe coded app.
Code is cheap now. The most important thing is to ship, ship, ship. If you are handwringing over "tech debt" you have already lost - and you deeply misunderstand how good this technology is getting!
If we get to post-scarcity you never needed the money. But if we end up in some dystopian hellhole where people have no jobs but capitalism still exists, you'll be thankful.
I work in embedded systems where it's harder to offload the burden of testing. There is also generally more legacy code, so experienced developers can be costly to replace.
AI makes some things easier, but costs have certainly not fallen by 2 orders of magnitude.
E.G. over time we’d gained two client-side caches of related server state. This started out as two different parts of the same model, because we couldn’t get all the data we needed from one microservice and had to merge in the client. Over time, more and more features used both caches for different aspects of related processes. At some point one of the microservices changed so as to return all the data in one call. The update to consume that kept both caches, adding code to sync them, because so many parts of the code were using one as a fallback for the other, so they both looked “necessary”. Because they were separate, and “live” sometimes they’d go out of sync after the initial load. Worse: the consumers alternated about which cache was treated as the fallback, making it very hard to see that either might be redundant. Eventually I noticed they were filled by the response to a single call. We all know paying back tech debt never gets prioritised, so I rolled the payback into two feature tasks, and just took longer about them.
My employer expects we use LLMs and provides some budget, but it’s not enough to use even Open4.7 or GLM-5.2 on every task. I do the bulk of my work with Composer 2.5. It’s quite good for “going forward” on smallish tasks and it’s written most of my code this year. It’s possible smarter models would spot these refactorinh opportunities and action them proir to building features or fixing bugs. But I wouldn’t know because I can’t afford it. I’ve never seen even a 4.8 era model spot a refactor and plan to do it prior to a “new build”.
I’m pleased I’ve spotted these trends and started to build the habit of (telling the agent to)“refactor to make the change easier”, but my percieved productivity will go down and I risk the ire of my leaders.
At a company I had setup scripts to build our packages, and the CI was running those scripts. Someone more junior (only a few years, not decades) found it strictly superior to remove my scripts and replace them with GitHub actions: people could now know even less about it (as in, no need to know how to copy-paste and adapt the recipe for a new package), but now it depended on GitHub. GitHub is down, nobody can build anything anymore. And it happened once every few weeks, so people would just go have a coffee during the outage.
You know what happened next? That person got promoted for their good work. That was before AI.
Future AI agents will refactor tech debt at lower interest rates.
I used plan mode for two reasons: to review the choices before execution, and to execute with another model (i.e., using the barely documented opusplan feature).
The grill-me skill is much better for reviewing and clarifying choices (and modifying it to use the ask tool makes you go faster). Instead of opusplan, you can explicitly tell Claude to start a subagent with another model to divide the tasks.
My only critic of the plan mode is I wish it was easier to see the updates and changes easily in Claude Code as we iterate on the plan. It is wasteful to have to remember what parts I have reviewed and what parts are new (and need another pass). I have thought about fixing this but I also feel the review is the actual thinking (even if ineficient), and so I purposely have not removed it.
I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.
We really need a better model. One alternative is to have an everything-app: a general purpose tool in which (almost) everything lives. The terminal is one of them. The text editor / word processor is another. (I use Emacs for everything.) In a business context the spreadsheet is probably the best choice.
“discuss your plan with me before implementing anything”
theres your plan mode