15 comments

[ 0.30 ms ] story [ 27.1 ms ] thread
We found after deploying many enterprise agents that letting the model choose tools at run time can cause problems. The agent holds the credential and sometimes chooses the wrong tool. Using the same tool every time prevents this.
Absolutely in love. I'll be testing this after the daily grind.

I've had nothing but success with converting daily work into "notes" that then translate into runnable, deterministic application CLIs to completely sidestep "what do I need to say to you to make you do the thing??"

I'm interested in how this spreads across enterprise flows, because the issue is always discovery and usability.

How deep do you usually go with CLI composition and layering? Do you tend to find a flat list of commands and sub command help works most? Have you experimented with connecting CLIs Linux-pipe-style as if it was a dynamic application in the OS?

The objectives are to reduce/eliminate as much inference variability as possible. A side benefit is that inference costs collapse as well.

It is used internally for what we call 'compiled workflow agents' where no inference is necessary. An agent composer determines the exact command at design time. That is part of a discovery loop that can introspect the service to construct the command.

More here. https://www.promptone.ai/resources/downloads/

The command structure is provider/topic/command. Taken literally from the oclif framework, with the provider simply being a topic namespace. Haven't needed to pipe anything since in our deployments, the agent makes a call to the gateway via a command API so a pipe wouldn't be possible between commands
"The agent holds the credential"

Huh? It shouldn't. Am I misunderstanding, or is this referencing poor practices?

"Using the same tool every time prevents this"

What does this mean? I looked at the project, im not sure what this means.

In many cases, the agent does hold the credential. When you authorize OpenClaw to read your gMail, OpenClaw has the credential. This is absolutely a poor practice, but common, nevertheless.

As for using the 'same tool', what I meant was that you the agent doesn't have to pick the tool at all. There is just one: the aclif CLI. Not separate tools for Salesforce, Docusign, Workday, etc that the agent needs to learn (and possibly mess up). Just the one aclif tool. Same grammar for all external services. Less agent inference the better.

Finally, alif CLIs support individual auth so a request can use SSO identities and fetch a token from a secrets value. The CLI holds the secret. If you deploy the CLI on a host or gateway, the agent never sees it.

Some really broad assumptions here, and youre being unclear.

"When you authorize OpenClaw to read your gMail, OpenClaw has the credential."

I can only assume you are implying that the execution environment accessible by the model via the harness here, had access to the credential. This is not even broadly true, as there are many single click solutions for deploying gateways that will allow operators to tls inspect and replace secrets in flight, outside of the agent execution environment.

Sure, not doing that is poor practice- but your phrasing is unfair.

"the agent doesn't have to pick the tool at all. There is just one: the aclif CLI"

Okay, so- if thats the only tool, then why are we even talking about credentials? Why are we talking about openclaw? The whole conversation regarding creds being in bad places is predicated on agents having native control over a sandbox- generally through shell. If your use case lets you bake whatever resource access is needed into a single tool- then many other security layers bubble up in value, being that you no longer need to authn/z arbitrary networked calls.

Also, yea sure- there is "one tool", all you've done is abstracted the tools into arguments.

"Less agent inference the better."

Show me the data, then. Show me how this performs better than the alternatives. This sounds like all you've done here is reinvent progressive disclosure?

"The CLI holds the secret. If you deploy the CLI on a host or gateway, the agent never sees it."

Okay, so- brokering, again. How are you solving authz then?

My response was simplified because your question was basic. OpenClaw example was just one that fit the pattern. As for showing the data for advantages, I'll refer you to Anthropic and Cloudflare's Code Mode

https://www.anthropic.com/engineering/code-execution-with-mc... https://blog.cloudflare.com/code-mode-mcp/

We have our own, but will defer to 3rd party evidence.

Reinventing progressive disclosure is definitely part of this, but that's only one element of the approach. To be clear, this is all based on what we use internally, deployed in a particular way within our platform. Will leave it up to others to determine how useful it may be for them.

As for authz, that's what the rest of our system does and is beyond the scope of the aclif effort.

I have to build a cli for every provider? Seems like a lot of effort.
Tell your providers to stop building workspaces with conflicting identifiers and namespaces.
This does not make any sense whatsoever.
How does this differ from the printing press?