Nobody is going to read that though, reading will go the way of code reviews. That’s the problem.
Ever since coding agents came out, It felt a little like we were implementing the self improving lisp programs of old. This idea almost gets us there. Could the next step be to make it the program itself?
At least it’s safe
The state of Agent APIs is bad. completions API (supported indefinitely) but does not support reasoning. OpenAI's reasoning API started out well, but is moving in the direction as described well in the article.…
yes it does - https://github.com/hsaliak/std_slop/blob/main/system_prompt.... I've added the same thing to the system prompt of the internal agent I wrote for work, which talks to claude models. I have not gone into the…
Seems to be doing too much, a 1 line in the system prompt is all you need. And it works well enough. “Output tokens are precious, be succinct in your responses. Use ASD-STE100 simplified technical english”
Let all trust this team to deliver AGI.
This graph biases outcomes (A| O) it would be interessting to have an others.. so we can view the whole picture
so the joke was the implication that they are not frontier.
no, this post was written by Demis from Deepmind.
I have my own coding harness (std::slop), in that i focus on a vim centric flow. Basically a few commands open. up my $EDITOR which is vim. 1. /edit => opens in editor 2. /feedback => opens the last llm message in an…
It's been clear for some time that model tool calling is heavily fit to a few common patterns, it's unsurprising that a tool call that looks the same or has the same name, but works differently, is falling back to…
so he's the quintessential brilliant jerk, ok
Gemini fumbled not on the models but on the basics. Gemini 3.1 flash was actually an amazing model to code with and their 20 dollar AI plans had solid value, but they locked it all behind 429s, needless gatekeeping of…
Big Monsanto energy
I’ve been looking for a way to articulate this shift, and your analogy nails it. The value of libraries and infrastructure components in software engineering is eroding fast. I am sure that in many organizations, teams…
I have a custom agent that generates patches like you will with kernel development and I review and merge those in. https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode... My agent forces this workflow by…
I wrote https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode... to avoid the brain rot from just shooting slop. It has helped me stay sane, review code and make changes step by step. I dont go as fast as with…
Its going to be interesting to see how this holds up in production after a release or two.
The problem with vibe coded re-writes is that you basically sign off on understanding the generated codebase at that point. Any historical knowledge of the codebase is gone.
the problem is not the forward pass, its the control/feedback loop when slop is written in response to the forward pass. Perhaps we should give the LLM 2 specs, one designed for the forward pass and another for the…
nobody knows what to build when everything can be built, there is no moat.
Using CRDT gossip to inform scaling is a clever idea. You are on to something there. Perhaps extract it as a core library/concept from the runtime? I feel that would be generally useful!
I have a coding agent https://github.com/hsaliak/std_slop where the sessions are in SQL ledger. So /session [new, clone, deletes, undo] are supported and all sessions are persistent. Cloning lets you 'fork' the context…
I find parallel agents to be an exception rather than a norm. Maybe I’m the problem? For those exceptional cases, opening a few more terminals gets the job done. It’s unclear to me if this needs to be the primary…
Nobody is going to read that though, reading will go the way of code reviews. That’s the problem.
Ever since coding agents came out, It felt a little like we were implementing the self improving lisp programs of old. This idea almost gets us there. Could the next step be to make it the program itself?
At least it’s safe
The state of Agent APIs is bad. completions API (supported indefinitely) but does not support reasoning. OpenAI's reasoning API started out well, but is moving in the direction as described well in the article.…
yes it does - https://github.com/hsaliak/std_slop/blob/main/system_prompt.... I've added the same thing to the system prompt of the internal agent I wrote for work, which talks to claude models. I have not gone into the…
Seems to be doing too much, a 1 line in the system prompt is all you need. And it works well enough. “Output tokens are precious, be succinct in your responses. Use ASD-STE100 simplified technical english”
Let all trust this team to deliver AGI.
This graph biases outcomes (A| O) it would be interessting to have an others.. so we can view the whole picture
so the joke was the implication that they are not frontier.
no, this post was written by Demis from Deepmind.
I have my own coding harness (std::slop), in that i focus on a vim centric flow. Basically a few commands open. up my $EDITOR which is vim. 1. /edit => opens in editor 2. /feedback => opens the last llm message in an…
It's been clear for some time that model tool calling is heavily fit to a few common patterns, it's unsurprising that a tool call that looks the same or has the same name, but works differently, is falling back to…
so he's the quintessential brilliant jerk, ok
Gemini fumbled not on the models but on the basics. Gemini 3.1 flash was actually an amazing model to code with and their 20 dollar AI plans had solid value, but they locked it all behind 429s, needless gatekeeping of…
Big Monsanto energy
I’ve been looking for a way to articulate this shift, and your analogy nails it. The value of libraries and infrastructure components in software engineering is eroding fast. I am sure that in many organizations, teams…
I have a custom agent that generates patches like you will with kernel development and I review and merge those in. https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode... My agent forces this workflow by…
I wrote https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode... to avoid the brain rot from just shooting slop. It has helped me stay sane, review code and make changes step by step. I dont go as fast as with…
Its going to be interesting to see how this holds up in production after a release or two.
The problem with vibe coded re-writes is that you basically sign off on understanding the generated codebase at that point. Any historical knowledge of the codebase is gone.
the problem is not the forward pass, its the control/feedback loop when slop is written in response to the forward pass. Perhaps we should give the LLM 2 specs, one designed for the forward pass and another for the…
nobody knows what to build when everything can be built, there is no moat.
Using CRDT gossip to inform scaling is a clever idea. You are on to something there. Perhaps extract it as a core library/concept from the runtime? I feel that would be generally useful!
I have a coding agent https://github.com/hsaliak/std_slop where the sessions are in SQL ledger. So /session [new, clone, deletes, undo] are supported and all sessions are persistent. Cloning lets you 'fork' the context…
I find parallel agents to be an exception rather than a norm. Maybe I’m the problem? For those exceptional cases, opening a few more terminals gets the job done. It’s unclear to me if this needs to be the primary…