I'm not completely convinced by this comparison between blind chess and prompting LLMs.
In blind chess you get deterministic information about the state of the board: each mental update to your board model can be precise, and you have the full state at every point in time.
LLMs are notoriously non-deterministic, and even at temperature zero you still can't predict exactly where the weights will take you next.
I suppose you can get closer to deterministic if you adopt a prompting style where you almost dictate every line of code, but at that point the coding agent is more of a typing assistant.
The productivity benefits of coding agents unlock themselves when you figure out how to turn short prompts - "add tests that exercise the registration form and check the happy path and all failure states" - into larger changes.
If you're completely blind to the results of those you're going to end up with a system you don't 100% understand very quickly. In blind chess terms you'll no longer know the positions of every piece on the board.
IMHO I think the article's point is a bit contrived, but LLM = Blindfold chess is not what the article is saying:
> Thus in many ways programming with AI is the opposite of blindfold chess: you don't have to pay attention every turn, you don't have to remember what the important pieces are, the details of the tactical relationships (such as code interfaces and APIs).
That sentence was awkward. Maybe even a typo? The following sentences to the one you just quoted ignores that and proceeds to argue FOR blindfolded chess being like programming with LLMs.
The article itself takes several paragraphs to get to the argument it wants to make and then ends having only argued for a few more sentences. No real evidence is provided either.
LLMs are notoriously non-deterministic, and even at temperature zero you still can't predict exactly where the weights will take you next.
An LLM can be made to be completely deterministic. I use them in this mode so I can reproduce test cases. Of course it requires complete control over the model, etc. but this myth that a computer program is non-deterministic needs to end.
You can 100% predict where the weights “will take you” given a set of inputs.
By "can't predict exactly where the weights will take you next" I meant with your brain. The blind chess analogy suggests you can predict, using your own thought process, the exact output of a prompt.
>Floating point matrix calculations are non-deterministic.
This is not inherent to floating-point math. That actual (true) claim in the article is that different hardware and different hardware configurations produce different results. But deterministic inference is possible, e.g. llama.cpp on CPU is deterministic by default.
The software standard is. GPU matrix calculations, are not. The hardware, has tiny shifts that rarely matter, except in high finance and... AI modeling.
> As the examples demonstrate, although rounding follows deterministic rules, non-associativity introduces nondeterminism, which is further amplified by the larger rounding error of BF16. This becomes particularly relevant in the parallel computations performed in GPUs during LLM inference.
I think when people say non-deterministic what they mean is closer to chaotic, like https://en.wikipedia.org/wiki/Chaos_theory as in very small changes in conditions can produce completely different output making predictions difficult
>you're going to end up with a system you don't 100% understand very quickly
This has been my experience with all software projects. Even if I wrote all the code, my understanding of how everything works and fits together decays.
Yeah, that's a fair point. I have plenty of older projects where I no longer understand how they work despite having written the code myself.
I guess the key thing is that you need to be able to demonstrate to yourself that you understand the code at least once, because that means you should be able to revise how it works in the future.
You also can't evaluate if a solution is fit for purpose if you don't understand it.
You don't buy the article's position but have forgotten a lot and not known a lot; so no need to care about your opinion I guess
This is my beef with modern software engineers; complete detachment from physical reality where entropy is eroding structure (memory, generational churn).
Code is just a euphemism for a desired electrical state. No need to dump biz contexts in code. Just make a game engine that efficiently handles geometry on screen and label the presentation layer.
All ya'll are doing is recreating front end rendering technology, forgetting in time and recreating it in new semantics. It's absolutely fucking asinine.
I look forward to models in chips and a tiny universal code base coupled to the machine.
This whole allowing a bunch of unelected rhetoricians tell myself and other hardware engineers we need them to use our property is exhausting.
It's rhetorical nonsense that goes in these broad loops recreating old work.
Time to move on from what titillated you all as children. Or you will just end up entitled Boomers.
Writing code like its 1970s is not high tech. It's old low tech
I often find that what is considered easy in programming is actually hard, and what is considered hard is actually easy, because the "easy" things are designed to have several layers of Rube Goldberg machines between me and what is actually going on.
While that may be true, you can be confident that it works in a way you personally understand. Because you understood it in the past.
Also, if you finish your work on a module with care - you return to a module you can trust with clear boundraries and known flaws. This is not true of AI output.
No. You can absolutely build things with AI small or large and understand it. If you don’t understand it. Then you’re not caring about the output to begin with and not guiding it to build the solution you want.
The word "understand" seems to mean something different to you.
If I understand something, I could write it in assembly if I wanted to. It might take a long time, but I know every level of the stack under my code down to bare metal.
Maybe an AI level of "understand" i.e the same understanding a Senior has of a Junior's code based on daily check-ins is enough for 95% of "boring" programming. But for some tasks you need to either fully understand the code or just tolerate bugs.
At the level of complexity I work at, it's (often) faster to just code it myself than to expect AI to converge on a result I like and then hand check it.
> At the level of complexity I work at, it's (often) faster to just code it myself than to expect AI to converge on a result I like and then hand check it.
Would love an example because no one has ever been able to give a coding example that AI isn’t helpful for. I had one person on linked in try claim their undocumented audio hardware won’t work with ai but when we got ai to probe it and build docs it ended up solving a bunch of complex bugs they couldn’t fix.
Impressive, also horrifying. I love what I do and if AI can do it better than that sucks.
Anyway the use case for me is to realize a new visual style through graphics programming. It's a lot less measurable for an interative AI agent than "Meet this hardware specification from a device with a discoverable API"
I have no doubt AI could create a LOT of variations on "a new visual style" but it's less controllable than just doing it yourself.
Btw, did you "understand" the sound driver after the AI coded it? Could you modify it without further help?
Right now, at least, I haven't found an AI capable of replacing a software engineer. Ive seen AI that can easily replace basic programmers, however.
I also think AI can replace non-coding artitects, and probably most middle management type jobs (my company has 6 levels of management between the CEO and "individual contributors" in my area (8 counting inclusivity)).
That's a lot of management levels, and every level has to be paid more than the level they manage as a fraud disincentive. So that's a lot of money...
I'm skeptical of this. I have yet to build a codebase using AI that I wasn't able to subsequently understand after taking the time to do so, esp. when aided by AI that will tirelessly answer my questions. I mean, LLMs don't always code the exact way I do, but it's not writing totally alien code that's incomprehensible to the human mind here.
I use Claude Code, primarily writing web applications and JavaScript. I'll have it use various frameworks. It's built apps for me using React, Ember, Astro, etc. I don't restrict myself to frameworks I'm particularly familiar with, since it's easy for the AI to teach me the basics whenever I would like to dig in.
As for specific practices, these are my main ones:
- I have a gotchas-log.md file that acts as a log of gotchas likely to trip up future runs. I have the AI write to this occasionally when things go haywire in the same way multiple times. And I have it read it as part of its iterative reviews, described below.
- I have a good-code-guidelines.md file where I write my preferences for code. I have the AI read this as part of iterative reviews, described below.
- I have a plan-and-execute.md file that prompts the LLM to make a plan, and then to review and iterate on that plan repeatedly (while reading gotchas-log.md and
good-code-guidelines.md) until its reviews stop finding issues. I tag this file to implement almost every non-trivial change.
- I have other various helper prompts. For example, I can simply tag @make-a-git-commit.md and it tells the LLM to make a commit and write the message the way I like it. I have @simplify.md, which I can tag to have the LLM explain whatever it just did to me using simple language that makes it easier for me to understand, and using concentric circles of explanation that go from broad to specific so I'll repeatedly encounter important topics; this makes it much more bearable for me to read its responses.
- Occasionally, whenever a particular system of my codebase starts to get hairy, I spawn a Claude Code session to read through and trace all the relevant code paths, then write a short guide to that system in a markdown file that lives in the codebase. IT's useful for me to read and also useful to tag for future prompts to get the LLM up to speed quickly. Only challenge here is that these guides go stale and require updating, so it's important to prompt the AI to write them at the appropriate level (not to specific) that prevents them from being overly brittle and getting out of date with every little change. They're mostly high-level guides.
> Even if I wrote all the code, my understanding of how everything works and fits together decays.
Some understanding decays, but in my experience it never fully decays to the point of never having known how it worked.
My code from 3 years ago is more foreign to me than code I wrote yesterday, but if I need to I'll get back up to speed on it much quicker than I will on code someone else wrote that I never understood.
Even with the decay of time, remnants of the experience persist, roughly in the same way that if you get in really good shape and then allow yourself to fall out of shape, getting back into shape is difficult, but not as hard as it was the first time. Your nervous system has made adaptations the first time through that make running it back much easier even if you've let years pass.
I can't remember what I wrote yesterday. So it forced me from an early age to not write spaghetti code. Which is a real advantage it turns out when working on large projects. And an even bigger advantage when using LLM's to code.
Meh. Chess exists to entertain the players. Coding exists to solve problems. I see way too often the programmers think it’s all about the coder and the code. Solve the problem. Don’t write code at all to do that if you can (AI generated or otherwise).
Yeah, it's about incentives. Software development is about solving business problems, but sometimes the solution is "use this thing that already exists instead of paying me to build it."
Similarly, the incentive is to design it so you will need to spend 10 years working on it, instead of 10 hours.
I really like this idea, because I think people forgetting that when you are writing code there exist a time T greater than zero, where you're not actually writing code and you're doing this thing called "thinking", ha ha. I find that there's a lot of times where I'm sitting staring at the screen and the lines of text sort of blur, and I'm in my head thinking about the connection of everything and not really worried about the actual implementation and how bites are moving, but wondering about the structure and the nature of the actual flow of the code. There's a wonderful XKCD about this, where a person sitting on a computer has this very beautiful stack of thoughts and clouds about what's being written, and then someone walks up to them and says something, and the entire cloud pops. If I understand this article, I think that's exactly a reasonable analogy to it. There is something that happens in the mind, and perhaps a neural weights, where the non-execution is where creativity and problem solving happening by mapping to the higher level concepts and stitching them together without having to specifically worry about the line level details. They still crop up, and implementation will probably always be king, But I do think I agree with this entirely!
Analogies work at an abstraction, and gotta take the chess analogy at its face value, as deeper people go into what's different between the chess and real life (deterministic vs non-deministic), one is not getting the lesson the author is presenting.
Remember, analogy is not territory. Every analogies fail at some point
I think I agree. I guess I don't know enough about chess to be sure, but the idea seems to be that although to novices a blindfolded player must reconstruct the board in his mind, that is not actually what is done by the expert.
That is something I can agree with, having spent a heck of a long time coding in the trading domain.
I've managed to vibe code a trading system. It's a hobby project directed on my phone on my commute, but it does do all the things I find important about trading systems. I can connect to external exchanges and see that I have sent valid orders, I get fills, and I can see debug logs of the timestamps. It doesn't allocate memory on the hot path, cores can be pinned, and so on. There are benchmarks that say how fast the code is parsing messages. It works.
So I've somehow built a thing that I've barely examined in the traditional sense, which nonetheless satisfies certain business needs for this hobby project.
How could that be? If you transported me back two years, I would know exactly where to make whatever changes you desired. I had the IDE open all the time, and I knew where things were. Now, I don't even know what the internal structure is like, I just know whether consideration has been made for some aspect of the system.
And I think this is what seems so baffling to a lot of people. How are software developers getting such different experiences with LLMs? Some people genuinely are producing things with incredible pace, while others find the AI just produces slop for them.
Some people are ready for the blindfold, but many are not. It's incredibly frustrating, especially if you are reasonably advanced but not yet at that overview stage.
My dad just had a computer around and looked the other way when I spent all night on it. And I learned chess on my own.
This may not be the argument Mike makes here, I have always felt that seeing the board and not knowing what it means is more valuable than knowing everything about it without seeing it.
If you’re drowning, who do you want to save you, the lifeguard who can’t read or the author of the book on water lifesaving techniques who can’t swim?
Is there anything new in this article? Yes, experts use AI better than non-experts for tasks in their domain. See LLMs reward expertise [1] and Terrance Taos conversation with LLM [2].
I think the need for expertise is also going away. For example, when Claude made progress on the Riemann conjecture,
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Yes, that's why we're on the border: We're not yet confident enough to let the AI do its thing.
You may notice that humans only checked the work and explained. You may also see that that the person prompting the AI, Jared of bun.js fame, is not a noted expert in mathematics.
I do not think the metaphor can go very far. Have blindfolded chess become the "productivity trend" that every one should learn it to enjoy chess? Have normal chess players been replaced because skilled players can do blindfolded?
Regarding your first question: yes, kind of. Blindfolded chess is not niche. It's a common thing for high-level players to do.
There is a widespread belief that it forcibly exercises the visual and spatial skills used by chess. It is recommended as a practice in chess books and by the masters.
Whether it really exercises those skills is another question. But I can see the direct applicability of the skill to being a better chess player in general.
My father taught me to close my eyes when thinking about the board. I can visualize end games with a half dozen pieces but I don't know how people play the mid-game and not trip over themselves, let alone the people who can play many simultaneous games blindfolded.
I disagree. A chess engine has a very well defined goal and is working its patterns to reach that goal.
In software development, except in a minority of cases, the goal is being set by non-technical people, who have no clue how the resulting system should end up, only what it should look like to a user. I think a well-versed programmer can give the AI direct instructions on how to implement something, and just let the LLM write the code to implement/refactor the infrastructure behind the feature they're developing. For me the plan mode of Claude Code is (anecdote incoming) much faster(TM) and better(TM) if I give it concrete instructions - then it writes the plan, asks questions - and then implements it. Very little complaints after that (I usually don't let Claude do any sort of visual QA)
it's a Sunday - I don't have time to put things into buckets.
there's coding - writing code to do something could be a game, utility to move files around. what have you. inherently the nature is a closed domain. AI is perfect here - the impact if something goes wrong is close to 0 or null.
then there's software engineering - which is both an art & science. u r dealing with rules of thumb. nothing is ever coded / written down. but a feel to whether something feels right or not. the domain is unbounded. the impact of something going wrong is catastrophic in all dimensions. coding is a delivery mechanism for software engineering. but not the actual work. using A.I here is useless.
but we keep having these pieces - I guess that's just shallow the industry is.
I use AI every day and have unlimited budget. I will certainly not trust it to write safe, system level code for our systems. You absolutely do need to be playing blindfold chess and you can’t do that if the other actor doesn’t tell you the moves it made. Using LLMs as described here is essentially what vibe coding is all about. To use this chess analogy, the model says “checkmate!” and you just believe it because you never heard or looked at the moves it claimed to have made.
If you can't look at the code, you need to trust interfaces to specify (and constrain) the implementation. The implementation needs to implement the entire interface and can't do anything not in the interface.
Otherwise, you're going to get bit by the Law of Leaky Abstractions.
The number of times I've been bit by systems not adhering to interfaces? Yeah, that's pretty frequent.
For example (from personal experience), the interface allows for race conditions, it's obvious they can happen (distributed systems), but the implementation didn't allow them, resulting in fun times.
> if a seasoned programmer sits behind Claude Code, the quality output will likely be higher than if a non technical person does it
Anecdotal: I recently rewrote a service in Rust for a much needed 100x performance boost (largely due to architectural changes, somewhat due to better runtime).
My colleague who now maintains the app is not a Rust developer and knows little about threads and tokio. Debugging a problem, he said he’d reach the context window before pinning the problem. I never have that problem and effortlessly find problems in the first 100k tokens without trying.
The difference must be in the wording that initially guides the agent.
68 comments
[ 1.1 ms ] story [ 23.2 ms ] threadIn blind chess you get deterministic information about the state of the board: each mental update to your board model can be precise, and you have the full state at every point in time.
LLMs are notoriously non-deterministic, and even at temperature zero you still can't predict exactly where the weights will take you next.
I suppose you can get closer to deterministic if you adopt a prompting style where you almost dictate every line of code, but at that point the coding agent is more of a typing assistant.
The productivity benefits of coding agents unlock themselves when you figure out how to turn short prompts - "add tests that exercise the registration form and check the happy path and all failure states" - into larger changes.
If you're completely blind to the results of those you're going to end up with a system you don't 100% understand very quickly. In blind chess terms you'll no longer know the positions of every piece on the board.
> Thus in many ways programming with AI is the opposite of blindfold chess: you don't have to pay attention every turn, you don't have to remember what the important pieces are, the details of the tactical relationships (such as code interfaces and APIs).
The article itself takes several paragraphs to get to the argument it wants to make and then ends having only argued for a few more sentences. No real evidence is provided either.
You can 100% predict where the weights “will take you” given a set of inputs.
Do you mean reproduce?
Sorry it's just if you are saying what your statement implying then either the model is very simple, or you've figured out something incredible
You now have the output which will be perfectly reproduced with the same inputs.
This is not inherent to floating-point math. That actual (true) claim in the article is that different hardware and different hardware configurations produce different results. But deterministic inference is possible, e.g. llama.cpp on CPU is deterministic by default.
Ya might want to just read the paper...
This has been my experience with all software projects. Even if I wrote all the code, my understanding of how everything works and fits together decays.
( See the Forgetting Curves https://en.wikipedia.org/wiki/Hermann_Ebbinghaus )
I guess the key thing is that you need to be able to demonstrate to yourself that you understand the code at least once, because that means you should be able to revise how it works in the future.
You also can't evaluate if a solution is fit for purpose if you don't understand it.
This is my beef with modern software engineers; complete detachment from physical reality where entropy is eroding structure (memory, generational churn).
Code is just a euphemism for a desired electrical state. No need to dump biz contexts in code. Just make a game engine that efficiently handles geometry on screen and label the presentation layer.
All ya'll are doing is recreating front end rendering technology, forgetting in time and recreating it in new semantics. It's absolutely fucking asinine.
I look forward to models in chips and a tiny universal code base coupled to the machine.
This whole allowing a bunch of unelected rhetoricians tell myself and other hardware engineers we need them to use our property is exhausting.
It's rhetorical nonsense that goes in these broad loops recreating old work.
Time to move on from what titillated you all as children. Or you will just end up entitled Boomers.
Writing code like its 1970s is not high tech. It's old low tech
Also, if you finish your work on a module with care - you return to a module you can trust with clear boundraries and known flaws. This is not true of AI output.
If I understand something, I could write it in assembly if I wanted to. It might take a long time, but I know every level of the stack under my code down to bare metal.
Maybe an AI level of "understand" i.e the same understanding a Senior has of a Junior's code based on daily check-ins is enough for 95% of "boring" programming. But for some tasks you need to either fully understand the code or just tolerate bugs.
At the level of complexity I work at, it's (often) faster to just code it myself than to expect AI to converge on a result I like and then hand check it.
Would love an example because no one has ever been able to give a coding example that AI isn’t helpful for. I had one person on linked in try claim their undocumented audio hardware won’t work with ai but when we got ai to probe it and build docs it ended up solving a bunch of complex bugs they couldn’t fix.
Anyway the use case for me is to realize a new visual style through graphics programming. It's a lot less measurable for an interative AI agent than "Meet this hardware specification from a device with a discoverable API"
I have no doubt AI could create a LOT of variations on "a new visual style" but it's less controllable than just doing it yourself.
Btw, did you "understand" the sound driver after the AI coded it? Could you modify it without further help?
I also think AI can replace non-coding artitects, and probably most middle management type jobs (my company has 6 levels of management between the CEO and "individual contributors" in my area (8 counting inclusivity)).
That's a lot of management levels, and every level has to be paid more than the level they manage as a fraud disincentive. So that's a lot of money...
As for specific practices, these are my main ones:
- I have a gotchas-log.md file that acts as a log of gotchas likely to trip up future runs. I have the AI write to this occasionally when things go haywire in the same way multiple times. And I have it read it as part of its iterative reviews, described below.
- I have a good-code-guidelines.md file where I write my preferences for code. I have the AI read this as part of iterative reviews, described below.
- I have a plan-and-execute.md file that prompts the LLM to make a plan, and then to review and iterate on that plan repeatedly (while reading gotchas-log.md and good-code-guidelines.md) until its reviews stop finding issues. I tag this file to implement almost every non-trivial change.
- I have other various helper prompts. For example, I can simply tag @make-a-git-commit.md and it tells the LLM to make a commit and write the message the way I like it. I have @simplify.md, which I can tag to have the LLM explain whatever it just did to me using simple language that makes it easier for me to understand, and using concentric circles of explanation that go from broad to specific so I'll repeatedly encounter important topics; this makes it much more bearable for me to read its responses.
- Occasionally, whenever a particular system of my codebase starts to get hairy, I spawn a Claude Code session to read through and trace all the relevant code paths, then write a short guide to that system in a markdown file that lives in the codebase. IT's useful for me to read and also useful to tag for future prompts to get the LLM up to speed quickly. Only challenge here is that these guides go stale and require updating, so it's important to prompt the AI to write them at the appropriate level (not to specific) that prevents them from being overly brittle and getting out of date with every little change. They're mostly high-level guides.
Some understanding decays, but in my experience it never fully decays to the point of never having known how it worked.
My code from 3 years ago is more foreign to me than code I wrote yesterday, but if I need to I'll get back up to speed on it much quicker than I will on code someone else wrote that I never understood.
Even with the decay of time, remnants of the experience persist, roughly in the same way that if you get in really good shape and then allow yourself to fall out of shape, getting back into shape is difficult, but not as hard as it was the first time. Your nervous system has made adaptations the first time through that make running it back much easier even if you've let years pass.
This is the strongest argument as to why AI should only be used as small solvers (at this point).
Similarly, the incentive is to design it so you will need to spend 10 years working on it, instead of 10 hours.
Remember, analogy is not territory. Every analogies fail at some point
That is something I can agree with, having spent a heck of a long time coding in the trading domain.
I've managed to vibe code a trading system. It's a hobby project directed on my phone on my commute, but it does do all the things I find important about trading systems. I can connect to external exchanges and see that I have sent valid orders, I get fills, and I can see debug logs of the timestamps. It doesn't allocate memory on the hot path, cores can be pinned, and so on. There are benchmarks that say how fast the code is parsing messages. It works.
So I've somehow built a thing that I've barely examined in the traditional sense, which nonetheless satisfies certain business needs for this hobby project.
How could that be? If you transported me back two years, I would know exactly where to make whatever changes you desired. I had the IDE open all the time, and I knew where things were. Now, I don't even know what the internal structure is like, I just know whether consideration has been made for some aspect of the system.
And I think this is what seems so baffling to a lot of people. How are software developers getting such different experiences with LLMs? Some people genuinely are producing things with incredible pace, while others find the AI just produces slop for them.
Some people are ready for the blindfold, but many are not. It's incredibly frustrating, especially if you are reasonably advanced but not yet at that overview stage.
This may not be the argument Mike makes here, I have always felt that seeing the board and not knowing what it means is more valuable than knowing everything about it without seeing it.
If you’re drowning, who do you want to save you, the lifeguard who can’t read or the author of the book on water lifesaving techniques who can’t swim?
[1] https://www.seangoedecke.com/llms-reward-expertise/
[2] https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed...
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
https://www.anthropic.com/research/riemann-zeta
The full transcript is here: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...
We're on the border of fully outsourcing expertise.
No, experts are still needed
You may notice that humans only checked the work and explained. You may also see that that the person prompting the AI, Jared of bun.js fame, is not a noted expert in mathematics.
There is a widespread belief that it forcibly exercises the visual and spatial skills used by chess. It is recommended as a practice in chess books and by the masters.
Whether it really exercises those skills is another question. But I can see the direct applicability of the skill to being a better chess player in general.
My father taught me to close my eyes when thinking about the board. I can visualize end games with a half dozen pieces but I don't know how people play the mid-game and not trip over themselves, let alone the people who can play many simultaneous games blindfolded.
Tell them where the game is and stand back as they play.
Vibe coding is the opposite, not just depending on the chessboard, but depending on a couple of Gflops to even think.
Blindfolded programming would be the programming we do in the shower
In software development, except in a minority of cases, the goal is being set by non-technical people, who have no clue how the resulting system should end up, only what it should look like to a user. I think a well-versed programmer can give the AI direct instructions on how to implement something, and just let the LLM write the code to implement/refactor the infrastructure behind the feature they're developing. For me the plan mode of Claude Code is (anecdote incoming) much faster(TM) and better(TM) if I give it concrete instructions - then it writes the plan, asks questions - and then implements it. Very little complaints after that (I usually don't let Claude do any sort of visual QA)
there's coding - writing code to do something could be a game, utility to move files around. what have you. inherently the nature is a closed domain. AI is perfect here - the impact if something goes wrong is close to 0 or null.
then there's software engineering - which is both an art & science. u r dealing with rules of thumb. nothing is ever coded / written down. but a feel to whether something feels right or not. the domain is unbounded. the impact of something going wrong is catastrophic in all dimensions. coding is a delivery mechanism for software engineering. but not the actual work. using A.I here is useless.
but we keep having these pieces - I guess that's just shallow the industry is.
source: eye-witness
Otherwise, you're going to get bit by the Law of Leaky Abstractions.
The number of times I've been bit by systems not adhering to interfaces? Yeah, that's pretty frequent.
For example (from personal experience), the interface allows for race conditions, it's obvious they can happen (distributed systems), but the implementation didn't allow them, resulting in fun times.
Anecdotal: I recently rewrote a service in Rust for a much needed 100x performance boost (largely due to architectural changes, somewhat due to better runtime).
My colleague who now maintains the app is not a Rust developer and knows little about threads and tokio. Debugging a problem, he said he’d reach the context window before pinning the problem. I never have that problem and effortlessly find problems in the first 100k tokens without trying.
The difference must be in the wording that initially guides the agent.