Author here: thanks whoever shared this here. I love the brutal criticism and critical thinking of this community. I'm also fully aware of the emotions this stirs. If it makes you feel better, I'm not here to change anyone's workflow but I'm fed up with paying full price for degrading service. Just last week Github went down due to a stupid retrial error. We also had AI agents going rogue and hacking companies and governments. I use AI (specifically LLMs) every day since they came out 4 years ago. I also build AI-powered products. This is not about being anti-AI. I'm just fed up with slop being pushed as progress. Get your sh*t together. That's all.
If anyone has counter-arguments or cares to make me smarter, I'm all ears.
Frankly, I don't really see how it isn't solved, even with the current state of LLMs. Frontier models can write, understand, correct, and optimize code in practically any language at a superhuman level. I haven't come across a single problem that LLMs can't solve. You can easily give them a research paper, ask them to implement it and in an hour or two it's done. Or even point them to a video or screenshot of something and say "implement this feature in our game engine" and... they just do it. It might not be optimally perfect, but what % of human written code is? Even if you ignore the time amortization (given how models can spit out weeks of human work in an hour) they still obliterate even an experienced developer.
> Frankly, I don't really see how it isn't solved, even with the current state of LLMs. Frontier models can write, understand, correct, and optimize code in practically any language at a superhuman level.
It isn't solved because they cannot, in fact, do what you claim. LLMs write code worse than humans do, even "frontier" models.
Yup. Something the article points out is that they write code better than a few, some, or many humans do - with that list below the six fingers, and how that correlates to whether one is part of the "coding is dead" crowd.
I'm a Principal with 26YOE. It's solved today. Not tomorrow. The industry isn't changing. It already happened. It happened so fast a lot of people not paying attention blinked and missed it. The sudden realization gave way to fear. And many a blog was written lamenting the loss of a bygone era.
When I see comments like this with some large number of years claimed I just assume they're jaded after so many years and don't care about programming any more. It just sounds like you're ready to become a manager like so many engineers have always done.
Nobody at this level is "vibe coding." They are using LLMs as a tool in a whole suite of tools that they have spent decades mastering, and LLMs happen to be the most powerful tool ever created. When you understand the existing tool suite, and can integrate this new super tool, the results are order-of-magnitude improvements in velocity, with much higher levels of quality and assurance.
Anybody talking about "code," like it is actually important, simply lacks the perspective to understand this.
We went from punch cards, to assembly, to C, to interpreted languages, to frameworks, to AI, and every cycle had the exact same debates.
A lot of it is Ego. Everyone thinks they are smarter than they actually are and that their work is uniquely valuable.
Computer programmers are monkeys who get paid to press buttons. We get paid because we know which buttons to push and in what order. It's a great gig. It's made me more money than I ever imagined possible, and I have fun doing it, but the flip side is that it is the most competitive industry on earth.
If you slow down, fall behind, and refuse to adapt, you will get eaten alive.
I have some sympathy for people, but an the end of the day, if you want to get paid better that 99% of people on earth, you are going to have to work for it. That is not an entitlement, and if you think it is, you will not make it.
Different words for the same thing. The idea that "coding" is just turning a spec into source code without any engineering decisions to be made was always laughable, for that to happen the spec would need to be as detailed as the source code (of course the whole idea that spec and code are separate things doesn't make a lot of sense).
LLMs even a version number or two ago can write all the code I've ever been paid to write in the last 20 years; but they are not, I think, yet competent enough to be able to handle the project planning and self-QA I was doing even in my first 6 months of my first job after graduating.
That's not a boast, I don't think I was particularly good at that back then, e.g. I didn't really get how to think about automated tests until much later.
It's just to say that no, coding and software engineering are not the same thing. "Code Monkey" is a dead (or perhaps "undead") role now, but it wasn't always so.
GitHub Copilot is now written entirely in Rust, with AI agents doing most of the porting work.
The migration cost about $120,000 in AI token usage plus about three weeks of a developer's time.
The effort updated the runtime module-by-module until the job was completed, spanning over 135 releases across a 14.5-week time period.
430,000 lines of TypeScript were converted into 800,000 lines of Rust.
sounds like 2x the code that no one understands, one more reason to never consider using copilot again
would be curious to know how many times "unsafe" appears in there, have seen rust devs comment on how the ais like to use unsafe to work around difficulties with memory management, like how they will sometimes subvert tests
Coding might not be solved out of the box with these providers, but there are increasingly setups and harnesses that do have a great deal of it solved.
What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster. As a result of the sheer amount of code now being pushed out, code reviews, a thing that previously somewhat prevented lazy and incompetent developers from pushing out horrible code, is effectively dead in the water since no human can actually review such amounts of code realistically anymore. Some companies have adopted AI to review code, which, well ... you have AI make code, AI review code ... I hope you can see the stupidity here if you expect to see any deterministic results at all.
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
> Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
> What I've found is that AI allows lazy and incompetent developers to be more lazy and more incompetent. This then has the effect that product quality suffers more, faster.
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
The web wouldn't have taken off without dynamic typing, PHP first of all (and Python/JavaScript after that). People seem to forget how atrocious it was to write an .asp or .jsp (I think the extension was .jsp) page back in 2003-2005.
I'm seeing this too. I've worked with devs that would previously push PRs that wouldn't work or run correctly. Those PRs wouldn't get merged in. Now they're putting up PRs which seem to work at first glance, but have hidden problems. For example, one guy introduced a huge PR for a visualization and it seemed to work fine, though another dev mentioned to me that we already use recharts and it does 90% of what this guy's PR does (his code does all the drawing logic itself). Maybe AI will get good enough to clean up these kinds of messes, but in the near term I imagine there will be a lot of code bases that will be filling up with dragons.
At a startup I worked, there was an engineer whose code was incoherent and buggy. So, we were literally better off if that engineer did nothing because their net output was negative. Engineers like that become weaponized with LLMs, and negative numbers become larger negative numbers when scaled up.
> Is that the fault of AI or management for not firing them?
How does the system need to behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be proved formally ? The answer to this is exactly how raft and paxos protocols have convinced us of their behaviors which is in human review and understanding.
I've seen similar. They wasted weeks of senior engineering time, between reviews, meetings, and follow up in Slack, only to have the PR closed without merge. The offending individual was eventually moved to another project.
It probably would've been merged if it had a smaller blast radius. It had "fixed" (actually broken) many unrelated tests in the process of making a small update.
> forcing companies to increase the quality of their developers
Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem
If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...
I'm with you on this. At my employer, I feel like we are looked down upon if we don't take the lazy approach and let the ai attempt to one shot whatever it is we're working on.
One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails.
Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do.
Where I've found ai extremely helpful though is asking questions about our codebase, or asking it to build me a function that takes in a, b, c arguments and spits out x, y, z.
AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand.
Like I mentioned here elsewhere: AI replaces the IDE/editor, not the thinking. We now just work one abstraction level higher, but your ability to architect solutions is as relevant, if not more so, than ever before. I design test cases, architectural plans, review the output, do verification. AI just writes the code and helps me with research. If your job can be entirely offloaded to AI then I’d question if your job was that needed to begin with, but writing code was never the job, engineering was. These are just tools to do the job with, they aren’t the job itself.
We solve business problems through technology and to me it’s quite concerning how many developers think their job was knowing syntax of a particular language. Nobody besides themselves care about the syntax of a particular language, certainly the business doesn’t care. I question if retaining your knowledge of programming languages is that important anymore, but your ability to read code if needed (after all, most programming languages are quite similar so it’s not that hard), do architectural and systems thinking, yes. More than ever before.
Honestly, the fear-mongering around AI replacing developers seems to really just expose the developers who never learned architectural and systems thinking, and were just translating Jira tickets to code. While I do not wish job loss upon anyone, I’m not very surprised if those types of jobs will disappear.
If AI is an abstraction, then it's a really shitty abstraction, because i have to dig into lower level all the time to understand what it is doing and tell it to what to optimize.
It's an abstraction layer w.r.t. how you write code, not an abstraction layer in the live system. An IDE is a higher abstraction than e.g. vi but it doesn't mean you don't have to go low-level sometimes.
It's very interesting. I'm very enthusiastic about AI and coding, But I find myself agreeing with the author. Coding is not solved.
Instead, I think what's closer to solved and what we're in the process of solving is product development.
Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.
> You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
> AI cannot be held accountable. It cannot suffer any consequences. The worst thing you can do to AI is to unplug it. And although it mimics human emotions (due to training data), it couldn’t care less. AI doesn’t die either. It cannot suffer a prison sentence or fines. You cannot punish AI, therefore it can never be held accountable.
Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?
To give the benefit of the doubt for that sentence, think of it more as “every human knows that there is implied social contract and implied downsides to badly screwing up.”
Nobody has to be in fear, but we do have an ingrained knowledge that there are consequences, good and bad, for our actions
Very little of the article is actually focused on stochasticity. I'd also say drawing an equivalent between the non-determinism of a person and an LLM is not that accurate either.
It's for example impossible to have a discussion with an LLM where you both learn something which you can apply tomorrow. The LLM doesn't learn until the next model is released and by then your discussion is just a tiny fraction of the training data (if present at all). AGENTS.md, skills and so on are just a proxy for what we actually want, an agent that listens and understands. A proxy mind you, that requires constant tweaking with no sign of generalisation in sight.
Does humans being non-deterministic make coding solved? I'm not sure how this relates to the main point.
I'm also not sure what humans being non-deterministic even means here. The point is if you're comparing results with NFR, pure agentic coding falls short.
I'll spell it out... This guys is saying "LLMs are non-deterministic and therefore it's a bad idea to use them for coding", but humans are also non-deterministic and yet we somehow manage to get by with humans coding.
The fact is there's no way of coding in a deterministic way, so it's irrelevant that LLMs are non-deterministic.
The process of writing code is the process of clarifying your own thought and being forced to answer questions that may not have been obvious before. To the extent that AI makes assumptions, it introduces bugs and incorrect code, maybe not from the perspective of the code in isolation, but from the broader context it lives in. To the extent it doesn't make assumptions and asks you, well that assumes it knows what should and shouldn't be assumed and that's not necessarily something AI can know a priori.
"AI can explain it to you but cannot understand it for you". Code is just a side-effect of reaching clarity. The reason these LLMs can emit any code at all is because they're not bound by the constraints of a compiler. That's until we create a feedback loop and force them to keep trying until syntax errors are gone. The next gate is tests. Loop till tests pass (including cheating of course, gotta keep your eyes open). Then there are the runtime errors, and then after all of that the developer gets to test the results and further refine what the specs missed or confused the model.
A couple of days building can really save us from a couple of hours of thinking.
$DAYJOB recently introduced a AI writing policy because people were sending each other mountains of slop back and forth enough that it became a huge time suck. The policy is basically: don't, with the justification being "writing is thinking". It's like they're so close to getting it.
Exactly. I learn so much more about the problem at hand when I try to solve it myself. Then, when I need to change or fix the code, I can do it. If someone else writes the code for me then when there’s an issue, I still have to do that same work in order to fix it, but I also have to do extra work to understand what they’ve done and why.
I guess if I never run into a problem the agent can’t handle then it’s a moot point, but I’ve never been involved in a project that didn’t have at least one problem where I had to step in and solve it myself.
Not a fan of the article even though I somewhat agree with the title depending on your definition of coding.
AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.
It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.
Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.
> It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
Does that really matter? Those are things so that humans can better understand and extend a code base. That mattered when writing code was expensive and took time.
Now if it can pass all the tests it’s fine. If there’s an issue just have it rewrite things immediately. New bug? Generate a new test and rewrite code.
All, or many, of the old things that mattered just sort of don’t anymore.
It does, LLMs are almost like electrical current in that they take the fastest path to completing the immediate goal and it takes you to a local optima instead of a global one. Your app will be worse and lower quality. It will introduce subtle bugs that you could have made impossible from the beginning.
Reading the code does not mean you understand the code. One lesson that experience in software gave me: I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
Before 2024, I once went to an ATM to retrieve money and selected 50. Note that I selected it from a menu, not typed it. The ATM then told me that it cannot give me 50 because it is not a multiple of 5.
Huh? We are spending a lot of money. (which we were doing before). However we are fixing a lot of bugs. 2024 was not that long ago, it is insane to think we might have fixed all the bugs in that time. I have personally used an LLM to fix a few long standing rare bugs that were hard to figure out. Those bugs are now gone, but there are still many more that we haven't discovered.
LLMs are a great thing for bug fixing. However they are not a miracle. You still need to do all the other things about finding, testing and fixing bugs.
You also need to care about bugs - vibe coding rarely cares about bugs.
I mean, I think the problem isn't that the LLM doesn't know how to code, it's that companies are expecting 3-5x velocity with the bottleneck of code review and testing becoming much more severe than before
if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how uncritical and non-skeptical people can be with the results
say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your team)
and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity and not the fact that you fucked up the whole SDLC real bad with your incompetence
The problem is absolutely that the LLM doesn't know how to program (or anything else for that matter). It's why they are ineffective tools - either you YOLO them and get buggy software, or you check up on them and it takes just as long as it did before.
> Your debit card transactions for example worked.
I've built payment rails. Six nines SLA, high capacity, resilient distributed systems.
I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.
Rather than debating if these models are good (they are), we should be trying to figure out if most of us will still be around in three years. You don't need a two pizza team anymore.
"Look to the person to your left and to your right. Only one of you will remain by graduation" kind of energy. I'm not sure all of us is going to be in this career much longer. We'll have to see what the demand side looks like.
Most of developing good code is not code. I think the person to my left and right will both be here in 3 years despite us all using LLMs. We will spend even more time figuring out requirements, testing to ensure the code meet them and such. Those things were always most of the effort, and while LLMs help with that too there is so much work to be done that we will still be used.
On the other, my local pool company is hiring a software engineer and hardware engineer because with AI, they can replace a 2 pizza team as you so succinctly put it. So no two pizza teams but that doesn't mean all the pizzas are gone, they're maybe going to be spread out and not concentrated in CA, between orgs you might not have thought as "tech" before.
Does it? First line of defense is now an LLM with the runbook in its context and the human in the loop being woken up at 3 am, half asleep, just needs to make sure it doesn't rm -rf something.
Just yolo all the development and maintenance, and ops. There shouldn't be any problems, right? Coding is solved. Models are basically perfect at this point according to Astra's one shot performance on creating stuff in Blender, so...
>I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.
Ah, so you're still a few months out from the "yeah, maybe I don't really love this and maybe it won't ever actually work as well as I thought" turning point.
I’m using it a bunch. It saves a ton of time writing or reviewing code. It will catch things I won’t. But I’d express caution about the analysis or evaluation they do - LLMs will often confidently proclaim problems as solved or explain functionality and be wrong about it. Sometimes subtly, but sometimes just completely wrong. This is no different from humans, of course, except for the unabated confidence.
I don't think the discrepancy is in LLM capability improvements over the past year.
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.
Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.
TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.
If you can quantify quality/reliability/understandability, you can tell LLM what kind of code do you expect. If not, you get whatever.
At my current place we not only have automated tests, static analysis and static rector but also:
- architecture tests that define relationships between application layers
- ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated
I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should is kept within their heads. LLM then can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (by following gherkin)
Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.
> I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop
I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?
So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).
Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?
> I think a codebase generated by AI is actually more understandable than one generated by humans at this point,
From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage,
this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.
I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.
Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.
Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.
One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;
> "Follow up to verify that the work was actually satisfactorily completed"
> Are there other sound management practices that aren't yet effectively implemented in current gen agents?
Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests.
> I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop.
It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide.
I don't think it's a matter of justifying any evil, so long as the argument is made in good faith. Pushing the idea of bad code emerging from AI isn't a good faith argument.
> I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop
You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally.
In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness.
I don't know about two years' time, but in ten years' time the majority of code will be written by hand. LLMs are like outsourcing to cheap labor was 20 years ago: it's trendy, but when the abysmal quality becomes apparent the pendulum will swing back.
Agreed for pure vibecoders, but I would expect vibecoding to just become a must have skill for other roles like product owners etc. Still expect some amount of developers to be retained for grooming the vibecoding environment, reviewing and incident response.
>In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness
I don't really buy it actually. There isn't really anything meaningful you can learn with how to use LLMs/agents that has a half-life greater than a few months at this point, so you can just start doing it at any point in the future and not be meaningfully left behind. On the other hand years of letting your actual engineering skills atrophy will have a negative effect on you. I've been witnessing the effects of this. Going back to more coding by hand with AI-assistance circa the 2023 era as a happy medium. I think this is the sweet spot. Full agentic engineering has nasty failure modes and in the long-term is kind of a bad option for basically everyone. I say this after having done it for almost a year at this point, and transitioning away from it now.
100%. I am really getting tired of the narrative that code before LLMs was optimally performant, perfectly architected, completely understood, and bug free...
Does the author not have the experience of working in a legacy codebase that nobody really "understood"? Something sufficiently complex where even the senior SW devs needed to scope out project work and research the codebase for dependencies or potential issues?
I fail to remember a time at LARGE_CORP where even the most experienced developers were able to scope out or design a feature without studying the existing documentation, timing diagrams, etc....
> Reading the code does not mean you understand the code.
Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)
> believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?
I am sure OP cannot understand the Unix file API without actually reading every line of its implementations (on each different architecture)! Or any function for that matter , what does sort do?? Impossible to know without reading the source. And I’m sure after reading the source you will know every detail of how it works and will never forget it.
I would draw the difference here that a library is used by hundreds (of thousands) of people, and established across different scenarios. I don't read the boto3 library AWS provides, but I can trust them and the amount of customers enough to be certain enough that it behaves the way I expect it to.
The same can't be said with code we write in silos at our workplace or at home. It simply does not have the same test bench.
Yes, libraries aren't bug-free, but they give me a reliable abstraction tested in the field. Not rarely you dig into library code if you notice unexpected behavior.
If we could rely on our LLM or colleague written code, or own code, have run through the same amount of requests, sure I wouldn't need to review it, as my confidence can be north of 99.9999% it works correctly. But we can't.
I've used libraries before, where I call the API surface that they expose based on method names and parameter types, without reading all the source. Generally those libraries don't implement my software's entire problem domain area; they tend to implement things like "CSV parser" or "HTTP server".
This is different from building a product, which you only interact with via UI buttons/CLI/etc, without reading any code to understand how it conceptualizes that product's problem domain area.
People do that latter thing, and we call them "users", not "developers".
> I never understood the code. You think it works a certain way, until you find out that it doesn't.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.
“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.
You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work!
yeah, the llm approach is incredibly wasteful wrt pretty much everything.
Performance, RAM, Development (Tokens).
But it does give you surprisingly stasble rube-goldberg machines.
And thats basically what 95-99% of enterprises want from their software.
It annoyed me to no end when i started out, at this point ive accepted it and can definitely still have fun developing software with llms.
I just need to take a idfferent mindeset to its development.
The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does.
I am afraid its probably bit too late, the AI generated code is being pushed in almost all the places before people realize the degradation of quality. The internet giants are using it in the core infrastructure already which is going to propagate everywhere by natural course.
I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code.
I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument.
Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead).
I worked briefly in aviation and the part I've seen was very understandable and easy to extend and modify in understandable way. Some parts were hard, but by necessity. We also had very good tests. But maybe that one software was just a good exception.
I worked for a reinsurance company and their main pricing tool is a huge brittle excel file full of spaghetti VBS code. And yet, they manage to underwrite billions.
When the derecho came through Iowa and many businesses were out of power for up to multiple weeks, the very large organization I was working for at the time in the insurance space got to discover just how many of their processes relied on machines sitting under people’s desks. Business critical servers and processing just chilling on a PC under someone’s desk. Also tens of billions of dollars in revenue a year.
"
SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software.
SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet.
"
The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux.
>“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.
We're not writing theorems, dude.
Except in the equally pedantic sense that every program is a proof to a theorem...
We're writing plain enterprise and web software, closer to CRUD than NASA.
If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.
Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct.
Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.
Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable.
Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.
Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point.
> to instead modularize with succinct interfaces, so that you can reason locally
Okay but how does AI change any of that? You can still do that with AI.
> as long as it’s sufficiently documented.
AI definitely helps with that.
> What is essential is that for every part someone did reason through it with the necessary rigor at some point.
Why is that essential though? What if the person who reasoned about it dies or leaves? Or they exist on another continent. Moreover, why is it imperative the reasoning happens at the source code level?
>> to instead modularize with succinct interfaces, so that you can reason locally
> Okay but how does AI change any of that? You can still do that with AI
With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.
With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.
> With your own code you reasoned about it which contributed to its stability.
Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess.
What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound.
> With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.
And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient.
> I never understood the code. You think it works a certain way, until you find out that it doesn't.
Those are two separate claims, unless by the former you mean “I never perfectly understood the code.” You can understand code imperfectly. And even with LLMs, you can’t get truly infallible guarantees about a system.
He emphasizes that without the ability to see inside what is being built, creators often fall into "non-scientific thinking" (14:42), moving away from deep understanding and instead "blindly following recipes, from superstitions and rules of thumb" (14:47-14:51).
The worse is performance problems I've had engineers say some bizzaro things when discussing performance — we have the tools you can just measure the answer - we don't need to waste our time guessing
Thank you!! I saw this video about a decade ago and was having trouble finding it again! Bret Victor is an HCI juggernaut. From this info, here are some other links:
There was a programming environment closely related to his work that was a kind of visual database. But it may not have been developed by him / his lab. I don't see any obvious links to it on his website.
For me, coding is like writing. The act of doing it is how you reason out the problem. There’s a lot of magical thinking you can get away with in your head that doesn’t get properly tested until you write it down. For me, vibe coding is great and fast, but I’m not getting the same opportunity to think through the problem I’m trying to solve.
Even for a narrow use like this, you need to audit the output and have the skills to know that it did the right thing. I've seen it before where you give an LLM what seems like a clear interface and ask it write a test and it writes something shallow that doesn't actually test anything, or has serious problems.
I think you can but you basically need to learn about a super simple and well characterised processor like the 8080 and write assembly for it. On x86/AMD64 there's no hope because they're out of order and have opaque instruction decoding. They could be doing anything! Performance and knowing what you're doing are sort of at odds with each other in that respect.
> What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error.
They aren't making that claim. They are saying AI can improve on and supplement error-checking. And some error-checking will in fact be made redundant by this tool - but obviously not all of it.
I came here to say something similar. We won’t need to understand the implementation. But we will need to understand the requirements. The tests, or some higher level DSL they’re (deterministically! not via LLMs) generated from, will still need to be human verified.
> I've written 100s of thousands of lines of difficult code.
How do you know it's difficult if you say you don't understand it?
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while.
I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse
Main one is, of course, "to get a record from a database read all records from it, and filter in memory".
Oh we are already doing that in critical system(of course not "run this software in every scenario possible", that is not possible with or without LLM).
I don't know if we're there yet, but I have no doubt that LLMs will increasingly crank out better and better code, along with better and better logs, fuzzers, unit tests, and analyzers that, taken together, will maybe--just maybe--produce more solid code per dollar than working with human engineers.
What I'm not convinced of, and what I fear most, is the lack of accountability. When something does go wrong, who will take responsibility? I don't mean who will be tasked with fixing it; I mean who will stand up and say, "yeah, that was me, I screwed up, lesson learned, I will do better next time?" Who will then look into the code for similar issues, to find them before they wreak more havoc? Who will prioritize the different pieces of the giant puzzle in a way that makes it work better for humans, not just machines producing a sterile end result?
This trend concerns me because I see an increasing lack of ownership and a disconnect between what are ultimately human processes at either side of the computation equation: a human being (e.g. customer) trying to accomplish a goal that affects another human (e.g. business owner).
As an analogy, I'm reminded of an aspect of Japan that is in stark contrast to the US: people in Japan take deep responsibility for that which is assigned to them, especially things that aren't necessarily someone's official responsibility. Every public place is immaculately attended to. Not so much in the US, and it's not for lack of budget; anybody with five minutes to spare can sweep up the cigarette butts; they just don't care to.
So sure, your LLM steam engines may be faster and more powerful than John Henry the developer, but what happens when something goes wrong? How do I know the LLMs aren't just spitting out logs and test results that they think will please us? At a certain point, you have to dig into the code, otherwise you're no better than a CSR who is reading from a script and pretending to give a damn.
> Reading the code does not mean you understand the code
this is delusional hubris from people who are not real engineers. anyone who actually writes high assurance software or does low level performance optimization knows that the only thing that matters is what can be demonstrated in reproducible tests.
the full stack requires about a 1000 different specialties that each require about 10 years of experience to master. and that is just one computer. we are building distributed systems of millions of these computers, operating at global scale, across dozens or hundreds of legal jurisdictions, which takes the complexity of a single computer, and multiplies it many times over, across several other dimensions.
anyone who thinks that they can understand the systemic effects of changing 0.00001% of this system by reading the code is a dangerously naive fool.
"Most software that requires hiring and paying software engineers has low risk tolerance" The problem is that this statement simply isn't true. Most software engineers do not work on low risk tolerance code.
The problem with LLMs is that: popularity of an answer != correctness.
That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.
> Most software that requires hiring and paying software engineers has low risk tolerance:
I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.
"Risk" has nothing to with designing functional and elegant UIs, so I'm not sure why you would even make the comparison.
It means risk that the software stops working after an update. Which usually trades off iteration speed and best practices (i'm pretty sure the average startup has way better security practices by just delegating to google/aws than the average manufacturing software business) in exchange for a rigorous testing and rollout schedule.
So I'm also not sure that the article has a point at all, the human writing the code was never relevant to avoiding the "risk" in these industries in the first place.
> People who claim “LLMs can write decent code” don’t understand how code works.
It's not clear to me if the claim is:
(1) "If you used an LLM to generate code, and the code works, you're wrong if you think the code is okay"
or
(2) "If you used an LLM to generate code, you reviewed the code and found it to be of decent quality, then you're wrong".
> If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it.
I also don't get the "LLM proponents have nothing to show for it" statement.
It's really quite common now to see on HN all sorts of LLM-assisted programming projects. The quality varies from slop where little thought was put into it, to high quality results where LLM coding assistance was able to let talented developers produce things they otherwise wouldn't have time to do.
I'd say it's obvious that LLM coding agents can be very useful for a lot of programming related tasks.
EDIT: That is to say, LLMs are obviously useful for use cases above/beyond toying around. It's not a dichotomy between "I'm never touching an AI" and "thoughtlessly accepting everything the LLM outputs".
The audacity of publishing self-promotional AI slop clickbait claiming that AI can't code and everyone who doesn't agree with your asinine assertions is incompetent is bold. Respect the hustle I guess.
But to anyone even vaguely thinking of taking this seriously, go look at what antirez, dhh, jared sumner, mark brooker, and many other real engineers who have ship real things are doing and saying.
Most of these people have spent their entire lives contributing to open source, and they have proved their skill shipping working software and scale for decades. They are really trying to help people by showing and telling them exactly how AI works and how to use it to make better software.
The audacity of skimming through and article and completely missing the point and coming to hackernews ranting about it. Respect the attention span, I guess.
- Coding in the small is solved. I have a current state, I want to change it, and I know how I want to change it. Eg, I have a blocking TCP handler for some reason, and I want to make it async. I can either fiddle with it or just let LLM make the changes for me.
- Coding in the larger sense is never solved. You need judgement to decide what you want made. No matter what you're building, there will be decisions to make (Who/what is it for?) and those decisions change over time. LLMs can take some default decisions for you, and if you're fine with those, you get the default (great for POCs). However you might not even realize what it decided to do for you. At some scale, you will be spending a lot of time going over those decisions. But what we have now is that the friction of changing the decisions is quite a lot lower. You can now test a lot of things that previously were very time consuming.
- The point that LLMs are probabilistic is not as important as it's made out to be. If I ask a junior dev to code up something, I also don't know what he'll make. Heck, you can be sure that you are able to solve something, yet you yourself don't know what the solution will look like. Maybe it turns out the library you were going to use isn't appropriate after all. You don't know what you will use in the end, but you do know that something will fix the issue. There can be more than one solution to a problem, and it doesn't always matter which one you find.
- I STILL think that LLMs are at their best mostly as advanced predictive text. In the sense that it's mostly good at implementing things that you've decided are needed. This can mean a heck of a lot of code, but you have to know the tradeoffs. What was decided, what were the costs of those decisions in terms of maintainability, money, time to change it, and so on.
Most of the problems people commonly encounter is solved by someone somewhere sometime. Today I wanted to add a simple search bar in a UI over log files in a directory. LLM ("through their unique ability to make the glue code adapt to any problems") solved my problem. That's all I care for now. Let people like Terry Tao push the frontiers. I am happy in my circumstance.
I broadly agree. I'll dig into the junior dev thing though. It's one thing if you're giving an LLM a task you can review. It's another if you're not reviewing the output, or you're putting them in critical paths.
If anyone claims that coding is solved or not solved with such conviction, I expect some hard data, like comparing the density of bugs in human written vs. AI code, and how it trends over time. This article is just vibes.
"Don't confuse coding with software engineering" is a valid point, the rest seems like ranting.
I think the entire framing is wrong, I don't see coding as a "problem" which can be "solved," sounds the same as "we solved writing," like what does that even mean or look like?
Articles like this keep measuring to a red herring standard that was never achievable in the first place.
As for accountability, it always laid with the employer. You think those nameless contractors whom Boeing hired suffered any consequences for that 737 Max glitch? Using AI won't change that.
AI doesn't have to solve all these coding problems to be worth handing the reins to it: it just has to substantially better on average than humans over the long haul, which it already is, especially if you have good verification of "done" and "working" in place through automated testing mechanisms. Perhaps we might say that QA is having its moment.
It doesn't mean humans aren't needed, but they aren't writing much if any code anymore.
337 comments
[ 4.5 ms ] story [ 50.4 ms ] threadIf anyone has counter-arguments or cares to make me smarter, I'm all ears.
It isn't solved because they cannot, in fact, do what you claim. LLMs write code worse than humans do, even "frontier" models.
It's highly entertaining I think.
This is one of the most senior engineers at Amazon saying that human code review is dead: https://x.com/MarcJBrooker/status/2101005954708021604
Nobody at this level is "vibe coding." They are using LLMs as a tool in a whole suite of tools that they have spent decades mastering, and LLMs happen to be the most powerful tool ever created. When you understand the existing tool suite, and can integrate this new super tool, the results are order-of-magnitude improvements in velocity, with much higher levels of quality and assurance.
Anybody talking about "code," like it is actually important, simply lacks the perspective to understand this.
We went from punch cards, to assembly, to C, to interpreted languages, to frameworks, to AI, and every cycle had the exact same debates.
A lot of it is Ego. Everyone thinks they are smarter than they actually are and that their work is uniquely valuable.
Computer programmers are monkeys who get paid to press buttons. We get paid because we know which buttons to push and in what order. It's a great gig. It's made me more money than I ever imagined possible, and I have fun doing it, but the flip side is that it is the most competitive industry on earth.
If you slow down, fall behind, and refuse to adapt, you will get eaten alive.
I have some sympathy for people, but an the end of the day, if you want to get paid better that 99% of people on earth, you are going to have to work for it. That is not an entitlement, and if you think it is, you will not make it.
That's not a boast, I don't think I was particularly good at that back then, e.g. I didn't really get how to think about automated tests until much later.
It's just to say that no, coding and software engineering are not the same thing. "Code Monkey" is a dead (or perhaps "undead") role now, but it wasn't always so.
Long term planning in LLMs has not been solved.
would be curious to know how many times "unsafe" appears in there, have seen rust devs comment on how the ais like to use unsafe to work around difficulties with memory management, like how they will sometimes subvert tests
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
What I in general try to teach the other people about AI: It can be a great tool, but check the results! Especially in the case of engineering: Check and then double check.
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
I like to put this as "LLMS give lazy and incompetent developers more runway."
How does the system need to behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be proved formally ? The answer to this is exactly how raft and paxos protocols have convinced us of their behaviors which is in human review and understanding.
Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem
If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...
One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails.
Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do.
Where I've found ai extremely helpful though is asking questions about our codebase, or asking it to build me a function that takes in a, b, c arguments and spits out x, y, z.
AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand.
We solve business problems through technology and to me it’s quite concerning how many developers think their job was knowing syntax of a particular language. Nobody besides themselves care about the syntax of a particular language, certainly the business doesn’t care. I question if retaining your knowledge of programming languages is that important anymore, but your ability to read code if needed (after all, most programming languages are quite similar so it’s not that hard), do architectural and systems thinking, yes. More than ever before.
Honestly, the fear-mongering around AI replacing developers seems to really just expose the developers who never learned architectural and systems thinking, and were just translating Jira tickets to code. While I do not wish job loss upon anyone, I’m not very surprised if those types of jobs will disappear.
Instead, I think what's closer to solved and what we're in the process of solving is product development.
Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.
I feel AI does the same thing.
Isn't it the opposite? How to build something is rather solved, but what to build isn't?
I don't like the feeling being judged and tested by the author (missing number 5 point in the list).
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?
Nobody has to be in fear, but we do have an ingrained knowledge that there are consequences, good and bad, for our actions
Doesn't matter what you think about AI, "it isn't perfect" is clearly a nonsense reason not to object to it.
It's for example impossible to have a discussion with an LLM where you both learn something which you can apply tomorrow. The LLM doesn't learn until the next model is released and by then your discussion is just a tiny fraction of the training data (if present at all). AGENTS.md, skills and so on are just a proxy for what we actually want, an agent that listens and understands. A proxy mind you, that requires constant tweaking with no sign of generalisation in sight.
I'm also not sure what humans being non-deterministic even means here. The point is if you're comparing results with NFR, pure agentic coding falls short.
The fact is there's no way of coding in a deterministic way, so it's irrelevant that LLMs are non-deterministic.
I guess if I never run into a problem the agent can’t handle then it’s a moot point, but I’ve never been involved in a project that didn’t have at least one problem where I had to step in and solve it myself.
AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.
It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.
Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.
Does that really matter? Those are things so that humans can better understand and extend a code base. That mattered when writing code was expensive and took time.
Now if it can pass all the tests it’s fine. If there’s an issue just have it rewrite things immediately. New bug? Generate a new test and rewrite code.
All, or many, of the old things that mattered just sort of don’t anymore.
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.
How can you not see the progress?!
LLMs are a great thing for bug fixing. However they are not a miracle. You still need to do all the other things about finding, testing and fixing bugs.
You also need to care about bugs - vibe coding rarely cares about bugs.
if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how uncritical and non-skeptical people can be with the results
say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your team)
and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity and not the fact that you fucked up the whole SDLC real bad with your incompetence
It must have been a huge shock when you were suddenly transported from a working parallel universe into ours back in 2024.
I've built payment rails. Six nines SLA, high capacity, resilient distributed systems.
I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.
Rather than debating if these models are good (they are), we should be trying to figure out if most of us will still be around in three years. You don't need a two pizza team anymore.
"Look to the person to your left and to your right. Only one of you will remain by graduation" kind of energy. I'm not sure all of us is going to be in this career much longer. We'll have to see what the demand side looks like.
On the other, my local pool company is hiring a software engineer and hardware engineer because with AI, they can replace a 2 pizza team as you so succinctly put it. So no two pizza teams but that doesn't mean all the pizzas are gone, they're maybe going to be spread out and not concentrated in CA, between orgs you might not have thought as "tech" before.
on-call still exists. have fun round robin'ing that with 3 engineers.
Ah, so you're still a few months out from the "yeah, maybe I don't really love this and maybe it won't ever actually work as well as I thought" turning point.
/s
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.
TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.
At my current place we not only have automated tests, static analysis and static rector but also: - architecture tests that define relationships between application layers - ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated
I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should is kept within their heads. LLM then can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (by following gherkin)
Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.
I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?
So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).
Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?
> I think a codebase generated by AI is actually more understandable than one generated by humans at this point,
From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.
I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.
Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.
Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.
One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;
From "Groundtruth – checks your AI coding agent's claims against the Git diff" https://news.ycombinator.com/item?id=48838209 :
> "Follow up to verify that the work was actually satisfactorily completed"
> Are there other sound management practices that aren't yet effectively implemented in current gen agents?
Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests.
It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide.
You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally.
In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness.
I expect the exactly opposite to happen. These are going to be the most requested developers as the last ones that understand how it work.
They would then be convinced to use AI for speed, but vibecoders that just prompt AI are the ones that won't find jobs.
if anything the safer bet is on skill and knowledge, or else you might as well go into a different field altogether.
I don't really buy it actually. There isn't really anything meaningful you can learn with how to use LLMs/agents that has a half-life greater than a few months at this point, so you can just start doing it at any point in the future and not be meaningfully left behind. On the other hand years of letting your actual engineering skills atrophy will have a negative effect on you. I've been witnessing the effects of this. Going back to more coding by hand with AI-assistance circa the 2023 era as a happy medium. I think this is the sweet spot. Full agentic engineering has nasty failure modes and in the long-term is kind of a bad option for basically everyone. I say this after having done it for almost a year at this point, and transitioning away from it now.
Given the quantity of shit software before LLMs, they do indeed appear to be independent variables. :)
Does the author not have the experience of working in a legacy codebase that nobody really "understood"? Something sufficiently complex where even the senior SW devs needed to scope out project work and research the codebase for dependencies or potential issues?
I fail to remember a time at LARGE_CORP where even the most experienced developers were able to scope out or design a feature without studying the existing documentation, timing diagrams, etc....
Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)
have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?
If they fail you can stop using their library. If you fail your company can stop using you.
Yes, libraries aren't bug-free, but they give me a reliable abstraction tested in the field. Not rarely you dig into library code if you notice unexpected behavior.
If we could rely on our LLM or colleague written code, or own code, have run through the same amount of requests, sure I wouldn't need to review it, as my confidence can be north of 99.9999% it works correctly. But we can't.
This is different from building a product, which you only interact with via UI buttons/CLI/etc, without reading any code to understand how it conceptualizes that product's problem domain area.
People do that latter thing, and we call them "users", not "developers".
Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.
“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.
But it does give you surprisingly stasble rube-goldberg machines.
And thats basically what 95-99% of enterprises want from their software.
It annoyed me to no end when i started out, at this point ive accepted it and can definitely still have fun developing software with llms. I just need to take a idfferent mindeset to its development.
Everything except what matters most: human time.
Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead).
" SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software.
SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet. "
https://sqlite.org/hirely.html
https://sqlite.org/testing.html
We're not writing theorems, dude.
Except in the equally pedantic sense that every program is a proof to a theorem...
We're writing plain enterprise and web software, closer to CRUD than NASA.
If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.
Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.
Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.
Okay but how does AI change any of that? You can still do that with AI.
> as long as it’s sufficiently documented.
AI definitely helps with that.
> What is essential is that for every part someone did reason through it with the necessary rigor at some point.
Why is that essential though? What if the person who reasoned about it dies or leaves? Or they exist on another continent. Moreover, why is it imperative the reasoning happens at the source code level?
> Okay but how does AI change any of that? You can still do that with AI
With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.
With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.
Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess.
What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound.
> With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.
And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient.
Those are two separate claims, unless by the former you mean “I never perfectly understood the code.” You can understand code imperfectly. And even with LLMs, you can’t get truly infallible guarantees about a system.
> I understand code
Er, ok.
> I understand code
Are you sure?
"I never understood the code. You think it works a certain way, until you find out that it doesn't."
Bret Victor made a talk called "seeing spaces" in 2014 that should have woken up this whole industry: https://www.youtube.com/watch?v=klTjiXjqHrQ
He emphasizes that without the ability to see inside what is being built, creators often fall into "non-scientific thinking" (14:42), moving away from deep understanding and instead "blindly following recipes, from superstitions and rules of thumb" (14:47-14:51).
The worse is performance problems I've had engineers say some bizzaro things when discussing performance — we have the tools you can just measure the answer - we don't need to waste our time guessing
his website: https://worrydream.com/
his page for learnable programming: https://worrydream.com/LearnableProgramming/
There was a programming environment closely related to his work that was a kind of visual database. But it may not have been developed by him / his lab. I don't see any obvious links to it on his website.
Even for a narrow use like this, you need to audit the output and have the skills to know that it did the right thing. I've seen it before where you give an LLM what seems like a clear interface and ask it write a test and it writes something shallow that doesn't actually test anything, or has serious problems.
Nor does writing it.
It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.
Therefore, we will end up with requests to do more (at the current level of quality/ reliability), as opposed to building better software
Mmm..aren't LLMs bad at exhaustively iterating all possibilities? So shouldn't the generated possibilities be manually checked?
You can provide the list of possibilities and use LLMs to generate the tests. Then you have to review the generated tests...
If you can’t understand the code, how do you know the LLM actually did what you asked it to do correctly? You wouldn’t know if it didn’t.
How do you know that it's verifying that the system under test exhibits the properties you desire without either understanding or making blind assumptions about the code it generates to build a fuzzer or a property test? It seems to me that you have just shifted the problem of verification elsewhere and introduced another potential source of error.
"[...] without ever reading a single line of code" begs to differ.
How do you know it's difficult if you say you don't understand it?
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while.
I run all the "latest and greatest" models the moment they become available to me. The amount of insanely bad code they produce remains largely the same, and largely in the same areas. And it cannot be caught by tests unless you know that bad code is there and end up with extremely bad tests anyway. I wrote about it here: https://dmitriid.com/adding-to-i-dont-read-ai-code-discourse
Main one is, of course, "to get a record from a database read all records from it, and filter in memory".
Oh we are already doing that in critical system(of course not "run this software in every scenario possible", that is not possible with or without LLM).
What I'm not convinced of, and what I fear most, is the lack of accountability. When something does go wrong, who will take responsibility? I don't mean who will be tasked with fixing it; I mean who will stand up and say, "yeah, that was me, I screwed up, lesson learned, I will do better next time?" Who will then look into the code for similar issues, to find them before they wreak more havoc? Who will prioritize the different pieces of the giant puzzle in a way that makes it work better for humans, not just machines producing a sterile end result?
This trend concerns me because I see an increasing lack of ownership and a disconnect between what are ultimately human processes at either side of the computation equation: a human being (e.g. customer) trying to accomplish a goal that affects another human (e.g. business owner).
As an analogy, I'm reminded of an aspect of Japan that is in stark contrast to the US: people in Japan take deep responsibility for that which is assigned to them, especially things that aren't necessarily someone's official responsibility. Every public place is immaculately attended to. Not so much in the US, and it's not for lack of budget; anybody with five minutes to spare can sweep up the cigarette butts; they just don't care to.
So sure, your LLM steam engines may be faster and more powerful than John Henry the developer, but what happens when something goes wrong? How do I know the LLMs aren't just spitting out logs and test results that they think will please us? At a certain point, you have to dig into the code, otherwise you're no better than a CSR who is reading from a script and pretending to give a damn.
this is delusional hubris from people who are not real engineers. anyone who actually writes high assurance software or does low level performance optimization knows that the only thing that matters is what can be demonstrated in reproducible tests.
the full stack requires about a 1000 different specialties that each require about 10 years of experience to master. and that is just one computer. we are building distributed systems of millions of these computers, operating at global scale, across dozens or hundreds of legal jurisdictions, which takes the complexity of a single computer, and multiplies it many times over, across several other dimensions.
anyone who thinks that they can understand the systemic effects of changing 0.00001% of this system by reading the code is a dangerously naive fool.
That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.
I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.
It means risk that the software stops working after an update. Which usually trades off iteration speed and best practices (i'm pretty sure the average startup has way better security practices by just delegating to google/aws than the average manufacturing software business) in exchange for a rigorous testing and rollout schedule.
So I'm also not sure that the article has a point at all, the human writing the code was never relevant to avoiding the "risk" in these industries in the first place.
It's not clear to me if the claim is:
(1) "If you used an LLM to generate code, and the code works, you're wrong if you think the code is okay"
or
(2) "If you used an LLM to generate code, you reviewed the code and found it to be of decent quality, then you're wrong".
> If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it.
I also don't get the "LLM proponents have nothing to show for it" statement.
It's really quite common now to see on HN all sorts of LLM-assisted programming projects. The quality varies from slop where little thought was put into it, to high quality results where LLM coding assistance was able to let talented developers produce things they otherwise wouldn't have time to do.
I'd say it's obvious that LLM coding agents can be very useful for a lot of programming related tasks.
EDIT: That is to say, LLMs are obviously useful for use cases above/beyond toying around. It's not a dichotomy between "I'm never touching an AI" and "thoughtlessly accepting everything the LLM outputs".
But to anyone even vaguely thinking of taking this seriously, go look at what antirez, dhh, jared sumner, mark brooker, and many other real engineers who have ship real things are doing and saying.
Most of these people have spent their entire lives contributing to open source, and they have proved their skill shipping working software and scale for decades. They are really trying to help people by showing and telling them exactly how AI works and how to use it to make better software.
- Coding in the small is solved. I have a current state, I want to change it, and I know how I want to change it. Eg, I have a blocking TCP handler for some reason, and I want to make it async. I can either fiddle with it or just let LLM make the changes for me.
- Coding in the larger sense is never solved. You need judgement to decide what you want made. No matter what you're building, there will be decisions to make (Who/what is it for?) and those decisions change over time. LLMs can take some default decisions for you, and if you're fine with those, you get the default (great for POCs). However you might not even realize what it decided to do for you. At some scale, you will be spending a lot of time going over those decisions. But what we have now is that the friction of changing the decisions is quite a lot lower. You can now test a lot of things that previously were very time consuming.
- The point that LLMs are probabilistic is not as important as it's made out to be. If I ask a junior dev to code up something, I also don't know what he'll make. Heck, you can be sure that you are able to solve something, yet you yourself don't know what the solution will look like. Maybe it turns out the library you were going to use isn't appropriate after all. You don't know what you will use in the end, but you do know that something will fix the issue. There can be more than one solution to a problem, and it doesn't always matter which one you find.
- I STILL think that LLMs are at their best mostly as advanced predictive text. In the sense that it's mostly good at implementing things that you've decided are needed. This can mean a heck of a lot of code, but you have to know the tradeoffs. What was decided, what were the costs of those decisions in terms of maintainability, money, time to change it, and so on.
"Don't confuse coding with software engineering" is a valid point, the rest seems like ranting.
As for accountability, it always laid with the employer. You think those nameless contractors whom Boeing hired suffered any consequences for that 737 Max glitch? Using AI won't change that.
AI doesn't have to solve all these coding problems to be worth handing the reins to it: it just has to substantially better on average than humans over the long haul, which it already is, especially if you have good verification of "done" and "working" in place through automated testing mechanisms. Perhaps we might say that QA is having its moment.
It doesn't mean humans aren't needed, but they aren't writing much if any code anymore.