311 comments

[ 0.20 ms ] story [ 75.0 ms ] thread
I have a tool wrapper that captures the output of anything and allows the LLM to query it later, to save on tokens. It “smartly” truncates the output (basically like Node’s util.inspect) and allows the LLM to expand truncated content.

It basically is called like “capture some-cli” and it… captures the CLI output, outputting a subset of it + a handle to continue querying.

This for me solves the danger of a tool returning tons of content.

Don't most harnesses already do that for bash commands?
I’ve been trying headroom-ai for this reason. That project also reduces tokens in other clever ways. Have you looked into it?
It seems like OP needs to provide a solution to hiding the credentials from the model in order to suggest CLI-mode only, and also a solution to the problem of agents without shell access.
In coop [0], a VM manager we developped to sandbox agent usage (i.e. give them freedom to do their stuff, but in an isolated environment from your main machine), we try to solve this by having an intermediate proxy (coop-proxy).

Instead of forwarding ANTHROPIC_API_KEY / OPENAI_API_KEY into the guest, we simply run a reverse-proxy within the VM that replace a constant-time token with the real API key on the host.

Works well enough in practice

[0] https://github.com/trailofbits/coop

[flagged]
You know you are getting old when Acronym's change on you.
Vacate entry port, program! I said, move out!
I’m not sure I agree that the frontier just know the apis right now, in my experience trying this there’s still a lot of faffing around trying to figure out the right parameters happily burning tokens and bloating context. Also the cost effective models to use in production for real agentic enterprise work absolutely still need the extra help and will do for at least the next 6 months.
>Recently, a Vercel engineer called on harnesses to send the programming language the client prefers

oh would you look at that, Vercel suggesting to abuse how standard headers have been used for decades so it can send Accept-Language: rust because it's too lazy to ask for standardising an X-Prefers-Lang or anything else, and Shopify is here to shit on the internet too. Great.

Don’t worry, you don’t need to attack them – they do a great job of making themselves look ridiculous with their ignorant conversation. I would be so embarrassed if I had suggested that in public then subsequently discovered that the header doesn’t mean that at all.
Yeah I think a better header is good. The idea isnt bad in concept.
Well Malte came up with AMP so there’s a history of doing odd things
Well, SaaS don't do CLIs for extension APIs.

Plus the performance issues to restarting processes all the time.

Neither of those things are true?

Also you are seriously comparing process startup time with the latency of a network call, or worse an llm call?

There are numerous applications that you don't need and don't want to give shell access to an llm.
i feel like MCP was bad, but people are saying recent improvements have made it worthwhile now? i.e. stateless http
There's probably still value (if you want to call it that) in it as a proxy, both to bypass IP address rate limits and to add necessarily credentials.

There's also another aspect Quite a few API providers provide automatic renewal for MCP server registrations, but not for personal access tokens. This may be less relevant when models just drive the user's browser.

> both to bypass IP address rate limits

Eventually this part will end. Most are in a "move fast and raise our stock value as much as possible" mode, so are fine with infinite auto-scaling to handle the surges MCP traffic is causing right now. But eventually, we'll probably see more per-auth rate limiting and/or heavily restricted MCP usage for non-frontier labs agents (especially if they only want e.g. Chat users, not coding harness users).

It's not just the agent understanding the API, it's locking down the access they have. If I want to give access to an internal service in specific ways that the API doesn't lock down then an MCP that offers very specific queries, with protective controls and transformations in place is very useful.
If one can build a MCP with proper protections, they can certainly do the same for their API/CLI/SDK.
I can build a local MCP that gives restricted access to an API which I have no control over.
(comment deleted)
MCPs are indeed useless, they're very limited in functionality and frequently struggle with large requests or get wedged in bad states.

There is no reason not to use the native API directly.

This doesnt match my experience. Yesterday, I was using Microsoft's Power BI Authoring MCP to make a semantic model from some SQL or CSV files. It was magical.

Microsoft has defined how to do that in the MCP. It's trivial to add the MCP to the machine and reliable in execution.

The alternative would be the model having to get the documentation directly from their documentation website, it sounds like. If this was the case, then MS would likely have great docs and probably support that markdown header... but everything hinges on finding a specific web page on the internet? Seems worse in every way than MCP to me.

AWS has many MCP servers that work extremely well.
But you had to install their MCP server on your computer? That works for developers. Wouldn't it be nice to avoid another install?
As opposed to which alternative? Spending credits to send the work to a non-local AI to fumble through it?
MCP for agents never made sense, especially when the tokens they consume a significant amount of tokens on a single request for a basic action, and sustained usage blows up you token costs.

The spec was poorly designed to begin with. Even saw some folks here thinking it was a good idea to enable MCP directly on a production database for what? Risking exfiltration of sensitive data for bad AI agents.

Given the increased security capabilities of these new models (Mythos, Astra, K3), it sounds like MCP would not be able to justify on making sense from a security perspective and would be a very bad idea to use anyway.

So no thanks and no deal.

Well, think of the near-future when tokens are so cheap that they’re not worth to meter.
(comment deleted)
I’m not sure about some of this — I still think there is some value to MCP as a gateway to private resources when API access doesn’t exist.

But please don’t try to redefine the Accept-Language header. These things are well defined for a reason and redefining things in the fly for LLMs is how we got into the current mess. For all of the cruft that W3C has, I think that working with standards committees could help the AI vendors here.

This, please refrain from using standard HTTP headers in non-standard ways. Why not use something like X-Accept-Programming-Language so that the semantics of Accept-Language can remain for human language, since the docs site is going to return text in a human language as part of the response anyway.
MCPs are winning because within the ChatGPT and Claude apps, there are Plugin stores. These plugins are one-click installation MCP servers, with support for authentication. This is what business users are using.
Exactly this. It’s a very effective way to integrate your app into Claude/openai.

I was anti-MCP at one point when it was eating up a substantial amount of context in Claude code. That’s largely been fixed now.

From my perspective, they are a great way to wrap an API for agent consumption. I can see a future where every major commercial or service website (think airline websites) have an MCP your agent can use to check flight status, rebook, or check you in.

To be fair, after visiting some trade shows recently, a lot of these “parasitic” LLM startups are not very convincing from a business-use case perspective. (They’re not really parasitic, more like a remora attaching itself to a shark, the shark being OpenAI or Anthropic).

At IMTS at Chicago the skepticism amongst the visitors towards AI and pure-AI startups was at an all-time high. More than once I heard, “Oh yeah I visited that AI booth and they couldn’t explain what value they would add to my business.”

So in that sense, while MCPs are “winning”, I also see them as a significant part of the AI “bubble”, specifically solutions looking for problems backed by VC money.

> Recently, a Vercel engineer called on harnesses to send the programming language the client prefers, so documentation sites can serve more specific examples. For example, adding Python could prioritize docs for the Python SDK instead of sending something generic.

I know this is pedantic, but IMO that should just be in the URL if the resource is going to be totally different. Accept-Language is already a bit weird for the same reason in my mind, but I think the intention behind it is the resource itself is attempting to communicate the exact same resource. Obviously, two different languages from two different cultures are going to have different interpretations of the same direct translation, but the intention is the service has at least tried to avoid that as much as possible.

Adding programming language into that same concept just makes it seems like you're serving both /docs/typescript/vx/... and /docs/python/vx/... from /docs/vx/... despite them (in theory) having many more differences in between implementations/context than that would imply.

Agents should read sitemaps. Does anyone's harness do specifically that when looking up documentation?

This article entirely misses the value that MCP brings today.

Sure, there's almost no reason to use MCPs if you are running a full-blown terminal agent (Claude Code, Codex, Meta Muse, OpenClaw etc) with unfettered internet access - just let it call APIs directly.

If you want to operate something that's less YOLO than that, you'll find yourself wanting:

1. Control over exactly which external services it can access

2. A way to handle authentication that doesn't allow the agent to directly access API keys

3. A sensible UI to allow users to connect and authenticate further services

4. Strong audit logging for what's going on

MCP makes all of that so much easier to provide.

Thinking MCP is obsolete because full coding agents don't need it misses out on all of the other things we might want to build.

Exactly, in our company, we have built MCPs that simplifies interactions with internal tools we use a lot, which saves time and tokens. Sure, we could let the agent poke and fumble around with a not-so-ideal API too, but it makes sense to formalize it and give the agents quick access to what we want it to fetch 99% of the time.
The simplified interactions are how it should have been designed in the first place.
It's difficult to come up with the optimal workflow on the first try. So an API should typically start out flexible even at the expense of complexity, in order to enable experimentation, and then you can optimize to make the common case simple, once you know what the common case is.
Sure but that common case should go into the core APIs, not an MCP.
Even if your agent has internet access, why waste tokens having it re-discover and re-implement its own API client each time? It doesn't make any sense.
This is my favorite position, the advantage of an MCP tends to lie more in token optimization.
On the contrary, you burn more tokens with MCP.
This is not true.
I guess it depends on the implementations.

When I use the Atlassian CLI vs their MCP server, I tend to see something like half the token burn with the CLI with more accurate results.

That could be an Atlassian issue but that's the results I'm seeing.

yeah MCP contains a provision for some blocktext for instructions. If your MCP has walls of text in those instructions, it will burn more tokens that a terse MCP.
How do you think the LLM becomes aware of which mcps are available? Vibes?
It depends on the harness. But most use a tool_search tool.
yes and using that tool does what with tokens?
(comment deleted)
You seem confused.
I'm not. You can build an MCP and throw a ton of junk in there and bloat the shit out of its token cost. So, really, "it depends".
In what way?

Not sure whats the norm nowadays, but it used to be MCP descriptions were loaded in from the start.

In any case, to be cheaper the `cli --help` command needs to more noisy than the json description.

Finally, and the really big one: cli can be composed with `grep`, `jq` , etc.

For me, I have a service that has multiple vectors of what would be called an API - PowerShell Modules, WMI, RESTFUL Web Services, some are available, some can do some things, some are more direct, some are not allowed with enterprise security etc.

Either way, I don't do what you suggest. I have self-learning rules and have the models build a well-rounded API engine once, then re-use it with query scripts through skills. Its portable and flexible in many environments.

No one is saying to wrap your local power shell modules in an MCP. That would be pointless and stupid.
> re-implement its own API client each time

Is using curl considered re-implementing your own api client each time?

Making the model use curl is even stupider. It still needs to do the same amount of work to lookup and understand the API contracts (assuming those are even public). But now it also needs to juggle the auth flows and marshalling at the tool call level.
You said ""re-implements an API client" which means you either don't know what curl is or don't know what an api client is. So I'd chill on calling things "stupider" when you don't have a strong grasp on the words you're using.
What is "implementing".... I'd define it as "figuring out the workflow, and storing it in a way that allows reuse". This could be a program written in assembly, or rust, or even python. Or it could be a shell script that calls curl. Or it could just be a set of tokens in the current session. Outside of the computer it could even be a set of processes people do, or a mechanical device.

If it's just a set of tokens in the current session, well then next session it has to figure out the workflow, and then store it in a way to use in the next session.

Seems like maybe you should take your own advice.

How does the definition of "implementing" change the fact that curl quite literally is an api client? It's an implemented api client.

> Seems like maybe you should take your own advice.

Seems like maybe if you have to use your own custom definition of a word in order to support a point you might not have one. lol.

Nonsense, curl is a tool for executing a single http request.

Many apis require several requests to get things done. (one to auth, one or more to fecth resource ids, one or more to modify resources, etc).

That would be a series of curl calls with logic applied to the output of each call to curl. An api client just does those things in a single function call. The steps are the same, but in one case the AI has to figure out each curl call and implement the logic, rather than just call the function.

lol Id say the nonsense is tying multiple calls to the definition of an api client.

Irony died in this comment thread.

By your logic 'ls' is a file manager, 'grep' is a search engine, and 'echo $X >> /proc/sys/$Y' is a settings manager.

It makes sense though. It's the same logic that allows you to say promting an AI with "do a simple thing for me" makes you a programmer, and prompting an AI with "what is an api" allows makes you knowledgable about computers.

In your mind, is netcat also an API client? When normal people talk about api clients, they're referring to an sdk or a cli which provides a simpler interface for a specific API.
You're not implying netcat and curl are synonymous are you?

And When "normal" people talk about api clients? lol This site is wild.

Alright dude ... Either you have major comprehension issues or you're arguing in bad faith. But I'm done responding to you.
Did you know that computer programs, once written, can be stored to disk once and run many times?
I love the level of sassiness this debate brings forward in people
People like you can't seem to comprehend that there are use-cases other than your little local claude code workflow.
Sure, but if you have a staff of 10000, you're doing this 10000 times. Unless you share that in a common place, and now you're re-inventing the thing we're claiming not to need.
A CLI doing all this is still a better UI for the agent though.
What does that even mean?
totally makes no sense. so weird after all this time people don’t “get” the value of MCPs, so weird
What if the agent doesn't have a CLI?
You hobble the expressiveness of the LLM and reduce its capability.

Think of an agentic harness as like a kind of body for the LLM. It gives it primitive inputs (read_file, web_search or whatever) and primitive outputs (edit file, respond to user, etc). Give it a command line environment (in a locked down sandbox, with as few or as many tools as you prefer), and you've given it a toolbox. It can do a whole lot more, faster and more efficiently. It can compose tools together. It makes fewer transcription errors manually shifting data around. It can tame verbosity with good protections in the harness and access to grep, sed and awk.

It's really up to you how useful you want your agent to be.

Yeah but you're still assuming it's running on someone's machine with a CLI to even use.
Someone can be OpenAI/Anthropic/whomever.

If you don't have something running somewhere, you don't have an agent, you don't have a harness. You've got a token generator, an LLM from the 2024 era.

That's where you are completely wrong. The point of MCP is that you can have an agent and a harness without running raw CLI or Python commands. Very common for relatively lightweight loads that involve shuffling data around between APIs, often run in a tiny serverless task.
Sure, you have an LLM which can invoke functions. I will say that without storage and composition, you're asking for hallucination.
What if I don't want my agent to have terminal access though? I get the feeling many people here simply never worked with smaller models, those will confuse cli args really quickly once context expands, and then you have no idea what damage they might do. With MCP's, they get just the access they actually need. Is the Principle of least privilege just not something we want to apply anymore?
CLIs, run in a sandbox as tight as your preferred choosing, live in an ecosystem, where, via pipes and redirection, input and output can be easily manipulated. An agent can do similar things but more laboriously (and less token efficiently) via Python or similar but it would still live in a sandbox somewhere.

Going without the sandbox means hobbling the LLM. It can do things directly but is less able to construct ad-hoc programs to deal with looping, conditionality, tame verbosity, connect tools together, and so on.

It's a choice to not give the LLM an environment. As you say, it can be necessary if you're using dumb models. I don't find it particularly worth the trade most of the time.

If not even OpenAI can properly box in their models I definitely won't trust myself to do so with the very limited time available to me, and instead just use a standard that's already defined, and proven to work.
By OpenAI, I presume you mean Irregular - these guys https://www.irregular.com/about ?

These guys are the common factor, the guys running the evals that let all the AI agents out, it looks like.

1. Control can be done via CLIs --> api key based access controls. We have been doing it forever. 2. I think this really only applies to Oauth based MCPs. Many server support api key based auth, stored as files --> security is still flawed imo. 3. This can be done via apis/clis too --> not something unique to MCP iimo 4. Same thing, not unique to MCP --> api servers can also be logged

MCP doesn't inherently make this easier. its still requires engineering maintaincence.

All of these things are perfectly possible with a plain REST API with an Open API spec and using some standard auth options, and an AI client that implements a reasonable “make api request tool” (just like the AI clients implement MCP today).

I think the real value of MCP is that it allowed companies to say “we’re doing AI!” When they built an MCP server. Just saying “use our api” was a lot less exciting.

It probably also helped cut through politics at companies where non-technical people didn’t want to open up user data with an API, but they did want to do AI.

Hah, I made that same point last December: https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-...

> For a while it also felt like MCP was a convenient answer for companies that were under pressure to have “an AI strategy” but didn’t really know how to do that.

I've since come back to MCPs, because I want to build my own agents without first having to solve the problem of effectively sandboxing Bash.

Yeah, from a design perspective MCP upsets me, because it’s a poorly designed standard and creating a good one could have been much easier.

But you’re right, since clients don’t have a nicely sandboxed “make api request” tool, it’s basically the way to go for a lot of use cases.

I think the new 07-28 spec is quite decent
"without first having to solve the problem of effectively sandboxing Bash"

Hopefully this is easier as time goes on. Of course- also policy on the egress

> without first having to solve the problem of effectively sandboxing Bash

"Sandboxing bash" is a problem that has been solved a zillion years ago already. Take your pick of any of the dozens of battle-proven solutions.

Which solution do you recommend?

Bonus points if it's available on both macOS and Linux and doesn't come from a random unmaintained GitHub repository with a note in the README that says "don't run this in production".

"Battle-proven" until an LLM decides it really needs to escape the sandbox you put it in and eventually succeeds.

For personal work, I run Codex in a VM that contains only what's necessary to do software development. Could it escape the VM? Sure, if there's a zero-day in VMWare Workstation.

Yeah, I'm using a pile driver when I really probably just need a hammer, but I've seen too many horror stories, and I don't trust guard rails. Even if there was an option to limit Bash calls to read-only operations, I would be 0% surprised to eventually run into "You're absolutely right! `rm -rf / --no-preserve-root` was a write operation! That's totally on me."

MCP is one of those things that is "too good enough".
Understood but it seems like effectively sandboxing cash is a very very important problem for the industry to solve!

Would be a much more robust and general solution of the problem of controlling and auditing agentic access to sensitive information.

It's such an important problem that it is sucking all available VC money into an exponentially-growing number of startups promising to make sandboxed agents safe and usable. In other news, MCP exists.
Honest question: which startups are trying to write a sandboxed-by-default easy to configure bash meant to be safely used by agents?
It's not "bash", it's containers/VMs/whatever that are isolated from the host and run the agent, which can access a shell to do work.
What do you mean by sandboxing bash? Isn’t this about just having a tool like curl or Postman?

Implanting an MCP client in your agent code isn’t all that different from calling requests or whatever

I mean the ability to have an agent run commands in a Bash shell without allowing them access to any file or environment variable visible to the user on that computer, and without allowing them uncontrolled internet access.
Program specific permissions (separate from the user operating them) are part of the Linux permissioning system already right? That doesn’t seem like an issue to me.

I guess the main problem would be finding an API client that can easily plug into your harness, with a nice UI for turning specific APIs on and off.

it's mostly true but the mcp also installs the knowledge of that REST API in a standard way so that a user can ask "what's projected revenue this month?" and it'll know how to hit your company brain and answer
Or just write good API docs that humans can use too.
most users i'm dealing with are not doc-reading developers. even getting them to tell claude to use tool X is pretty hit or miss whereas claude already knowing what tool to use is 100% hit with correct mcp tool descriptions.
Pretty much, MCP is still a bad idea.

LLMs perform significantly better and faster when you strap them to plain old apis/and an open api spec with a search tool.

My current MCP design is… grab a fastapi spec shove it into fastmcp, shallow wrapper, search tool for the full schema.

Oh boy so exciting I just wrapped an api spec for no reason and have to host infra for the translation layer. If only we invented api gateways.

But I am Mr. AI now.

How do you handle credentials safely?
Just leak them to the inference providers, obviously /s

If you have self hosted models and/or self hosted APIs, maybe you don’t need MCP to provide a gateway to a secure resource.

If neither of those things are true, you need an authenticating gateway/proxy or a target API that supports single use credentials (and get the model to generate a call to use them).

We can argue whether MCP is a good authenticating middle layer, but not whether one is required.

What about just handing the agent a token with limited time to live and constrained access permissions?
> For a while it also felt like MCP was a convenient answer for companies that were under pressure to have “an AI strategy” but didn’t really know how to do that.

MCP was a convenient answer for companies that had spent the last few years shutting down APIs because allowing API access bad.

Couldn't agree more. MCP is just Tool Use and the terminal agents all have embedded tool uses like WebSearch, Bash, Grep, etc and those are just MCP by another name. CLI's called by a model are just Bash Tool usage calls. Bash tool is just the most open ended broad MCP you can expose and what you gain is less context bloat (no specialized tool descriptions, just Bash) and what you lose is control over the agent -- until you setup a rigorous set of governing permissions on the Bash Tool.

I have a fleet of sandboxed Claude Code instances running and they share files with each other. The files are stored on AWS but they don't have access to AWS at all -- they can't see the access keys. In fact they don't know the files are on AWS. Instead they have a set of MC tools for listing/uploading/downloading from an internal filesystem URI handler (ie agentfiles://somefile.json) and the outer orchestrator of the Claude Code instances takes the MCP requests and does the actual file manipulation on Claude's behalf. The LLM seems to adapt quite well to this strange, arbitrary filesystem and I get to keep these agents fully compartmentalized. And I have tool request logs and logs in the outer orchestrator for full auditing of the agents. MCP is a really natural fit for this kind of stuff.

> what you gain is less context bloat (no specialized tool descriptions, just Bash)

Couldn't they train the understanding of a specific tool set directly into the model instead of needing it to be in context? Like isn't that basically what happens now with the Bash tool?

Is there a reason we need to rely on such a high level of access for something that should really only ever be cleaning up the project directory, hitting the 'Run Test' button and authoring some Git commits?

Perhaps a little out of date now, but I found claude was better (trained?) with the gh command line than with the github mcp. For a $corp internal tool ... I don't really see how and LLM would be able to be trained on it.
The spec describes Resources, Prompts, Tools, and Elicitation.

In practice, I believe Tools represent 95%+ of what people actually use MCP for. I've not seen an MCP with Resources or Prompts that seems to have widespread use of those features, and I don't think I've ever seen anything implement Elicitation.

Exactly. A lot of people complaining about MCP are doing so because their only interactions with LLMs are via big batteries including code harnesses and don't understand what kind of (usually much more domain-specific) agentic systems are being built. For example, "CLI vs MCP" doesn't make any sense whatsoever if the agent doesn't have access to a CLI!

MCP suffers from its harebrained choice early on to load everything into context up front.

The other problem with MCPs is that the harnesses don't auto-reconnect if you've enabled access. Connection status should be persistent
> Sure, there's almost no reason to use MCPs if you are running a full-blown terminal agent (Claude Code, Codex, Meta Muse, OpenClaw etc) with unfettered internet access - just let it call APIs directly.

I don't think this is a valid statement too. I'll explain why.

A MCP server represents those APIs that agents and coding assistants can call.

If you feel a need to provide data and services to agents through an API and feel so strongly about it and so compelled to implement your own APIs with the express purpose of being consumed by your agents, wouldn't it make sense to develop an API that is designed purposely for agents using a protocol designed to meet their needs and simplify their work?

Because that's what MCP is all about.

Nowadays, with the improvements in tool-calling and the dissemination of agent skills, MCP's value proposition isn't as clear as back when those weren't a given. But once you face usecases to either centralize your tools across an organization, manage access, and be able to audit it's usage, right now there is no alternative to MCP.

Agents are perfectly capable of using the same APIs designed for humans.
Real hard to not think-

"This is a job for a CLI/terminal, which is absolutely a million times easier to learn today, thanks to AI."

Feels like MCP is still a product of -- well, "AI as a product" brain, which I have no love or use for. Give EVERYONE ALL the tools.

So how do you make that CLI tool work in the context of a web client?
Exactly. There are so many replies here of the form “MCP sucks. If you just do <all the things mcp does a different way> you don’t need mcp at all.”

Well yes.

I think the point of those comments is this could have been much simpler for everyone if the model providers had given us some tools around existing standards.

Instead they invented something new/weird/complex standard, and then every client implemented slightly differently (and different parts of it).

> Exactly. There are so many replies here of the form “MCP sucks. If you just do <all the things mcp does a different way> you don’t need mcp at all.”

See my previous reply to simonw elsethread.

It's not "If you just do <all the things mcp does a different way> you don’t need mcp at all", it's that many of us were already doing that in a CLI prior to LLMs. You think only github and amazon had CLI clients?

> there's almost no reason to use MCPs if you are running a full-blown terminal agent

I disagree. Adding "https://mcp.linear.app/mcp" and having everything happen (discovery, usage, updates to the API, etc) without having to install or configure anything else locally is a big deal.

Funny you mention Linear’s specifically, I just posted about moving to a CLI instead because of the MCP’s egregious token-usage [1]. While this doesn’t discount the points you mentioned, I think the context savings (which can be huge, I hadn’t listed all differences in that post) outweigh them specifically in Linear’s case. It is just too inefficient in that regard.

Edit: this of course says nothing about MCP vs API/CLI in general. It’s just a bad implementation by Linear.

1: https://thebiglog.com/links/linear-cli-instead-of-linear-mcp

This surprises me a bit. I've found much better token usage with a proper/efficient MCP as even with a deep /skill defining usage, parsing MCP results is generally just better/more efficient than parsing CLI results. I say this having written a CLI tool explicitly for harness usage, and leveraging MCPs for the same.

I'm sure there are bad MCPs and great CLI tools that parse poorly/well via harness, but I'd be curious on an better research study.

Your comment doesn’t contradict mine and mine doesn’t contradict yours. Linear’s MCP is just (very?) inefficient. I explain the source of the difference in the last two paragraphs.
How much are using Linear if token usage is a problem? Sounds like a case of straining at a gnat and swallowing a camel.
In essence: MPC servers are the sandbox
True but a pretty shitty sandbox and we need a less leaky and more general sandbox solution ASAP.
>pretty shitty sandbox

Well, speak for yourself. My MCP servers are pretty solid.

There's nothing in the protocol making them inherently poor other than perhaps popularity, causing a swarm of people vibe-coding things they don't understand.

I've taken to sandboxing my entire agent in a Docker container. I wrote a tool that pretends to be an ACP client but is actually making Docker containers, copying files I specified in, bind-mounting, etc, and then proxying ACP via websocket to an agent in the container (except the ACP terminal/FS commands, those happen in the container).

It works well, though there is some leakiness around paths. I opted to make it place/mount files at the same path as on the host so paths are the same (as opposed to manipulating the ACP messages to modify paths on the fly, that felt messy and buggy).

Configurable networking is on my list for the future, but I haven't decided whether to start with IP-level firewalls or if it's better to start with a proxy and firewall rules to force traffic to it. IP firewalls suck for APIs that might have semi-dynamic IPs.

[1] https://github.com/SethCurry/abyss

MCP definitely has its niche in certain environments. It's good for a specific kind of constrained problem; not so constrained that you could solve the problem with just Node.js + fetch call to LLM API but not so open that you'd want to let the AI agent directly invoke any service it wants over the web. The latter is what happens if you give the agent access to curl.

Though I agree with OP's point of view that MCP was overhyped for too many use cases. Complex agent-system integration problems are usually better solved with just AI agent + curl + SKILL.md. It's just way more flexible.

It's another variant of the 'fat client/thin server vs thin client/fat server' debate. Some people want rigid, thin (e.g. web-based) frontends with the LLM doing work in secret behind the scenes. Others want fat, versatile frontends through which the LLM can interact with the user's own environment.

I've always been a fat client guy and this time is no exception. I doubt the constrained approach is going to lead ground-breaking innovation. I also wish companies would treat SKILL.md + curl as the main mechanism for agent tool calling as opposed to MCP. MCP is niche.

That niche environment is all SaaS-to-SaaS integrations. The client doesn't want to have users struggling to get every integration working on their platform. The service providers don't want to have to support non-standarized behavior by every client.
> This article entirely misses the value that MCP brings today.

Well, me too. I see the only advantage over a CLI app being an agreed-upon convention for syntax (not semantics).

I have a CLI interface to my webapp, not an MCP.

> If you want to operate something that's less YOLO than that, you'll find yourself wanting:

> 1. Control over exactly which external services it can access

Access control is not built into my CLI, it's built into the WebApp.

> 2. A way to handle authentication that doesn't allow the agent to directly access API keys

My CLI takes credentials from the environment, which it uses to talk to endpoints. The caller sets the environment, then calls the CLI program. The caller provides no way for anyone sending it input (the Model) to request or retrieve environment variables.

> 3. A sensible UI to allow users to connect and authenticate further services

The parts of my WebApp that relays or re-requests to other third-party services handles access control.

> 4. Strong audit logging for what's going on

Not sure what this is supposed to mean: the WebApp already has auditing logs.

> MCP makes all of that so much easier to provide.

Sure; I'm considering writing a purely deterministic shim for MCP around my CLI. The semantic/information is the same, the only difference is syntactical in nature.

> Thinking MCP is obsolete because full coding agents don't need it misses out on all of the other things we might want to build.

It may as well be; coding agents are at one end of the control spectrum - run in bash, do anything/everything (so they can leak credentials to the model, or the harness). But a CLI app doesn't have to allow bash. My "harness" (using the term very loosely) can securely call other programs without giving its own caller a RCE via bash.

So you've built your own agent harness that allows the model to call only your CLI but doesn't allow the model to run "env" and view the environment variables itself?

Sounds to me like MCP with a slightly different interface.

> So you've built your own agent harness that allows the model to call only your CLI but doesn't allow the model to run "env" and view the environment variables itself?

Yes, with the difference being it can call other CLIs, just not arbitrary CLIs.

> Sounds to me like MCP with a slightly different interface.

It is, except that it is not limited to being called from harnesses. Also usable from bash (automated scripts), or even humans if hey want to run it on the command-line.

TBH, my webapp(s) had this prior to 2020, because it made automation simpler so I could write shell scripts to do various things on the WebApp.

Still don't get what's the advantage over just adding a special API key or a wrapper cli or whatever mechanism that achieves all of that without being a "protocol" and with all the context bloat. Like the github cli is a good example. You give it proper auth keys etc. and for sure there's some telemetry in there about usage as well. If there isn't then it's easy to do from the api side too.

Even if you don't own the code or infra, like say a frontend team who wants models to test out the backend apis and do something. Well in that case how do regular devs do it? Do they also get unfettered access in the past? Surely there's still some mechanism you can repurpose for agents to use?

I'm not trying to argue I'm just saying I didn't catch on the first time mcp was a thing and I still don't know what it's doing now.

If a full coding agent can access a CLI tool. that agent can almost certainly access the API keys being used by that tool. They can go as far as decompiling binaries, or rewriting them to log the key before it is used.

If you are worried about a prompt injected agent stealing your keys, that's a problem.

(There is a way around that: you can use an HTTP proxy that inserts those credentials but otherwise lives outside of the agent's realm of influence. MCP is a whole lot easier though.)

> If a full coding agent can access a CLI tool. that agent can almost certainly access the API keys being used by that tool.

So, don't do that then?

Why do you need to use a full coding agent as the interface between the model and the CLI tool?

A 10-line program can do the wrapping of any existing CLI program so that environment is not leaked to the model, while providing the CLI program with the environment as well as restricting what programs can be called to a whitelist.

If you CLI program is echoing its keys in the response, or the endpoint is echoing keys back, that's not a problem that can be solved with MCP anyway.

So you're building a custom harness here that provides tools, and you're wiring up your custom harness to effectively do a subprocess execution of a CLI script for every tool call the model request?

One reason to switch to MCP here would be to avoid the overhead of forking a new process for every tool call, and to enable maintaining state between tool calls.

(That performance overhead is so trivial as to not be worth caring about, but the state thing may be useful - keeping a stateful browser session running between tool calls is harder with a CLI, for example.)

> So you're building a custom harness here that provides tools, and you're wiring up your custom harness to effectively do a subprocess execution of a CLI script for every tool call the model request?

Well, yeah. Subprocess execution is on the order of double-digit milliseconds. The "wiring up" is maintaining a whitelist of what tool commands map to which executable. It's a lookup table with very little maintenance required.

> One reason to switch to MCP here would be to avoid the overhead of forking a new process for every tool call, and to enable maintaining state between tool calls.

I feel like I am taking crazy pills :-/

The overhead of forking, on the ancient machine I call my desktop, is at most double-digit milliseconds. The state isn't being tracked by the MCP server anyway, it'll be tracked by the harness and/or the model, no?

My main reason for adding MCP support is so that existing callers that want to use my WebApp(s) can just use it without needing any changes on their side.

IOW, I am going to add it at some point, but not for the reasons you give. I'll add it to be compatible.

You didn't address my comment about state. If you're driving something like Playwright you need a way to maintain state between tool calls.

(Oddly enough I did solve that with my own CLI tool for running browsers during my "who needs MCP" phase, but it's a bit of a nasty hack that involves leaving files with PIDs lying around: https://github.com/simonw/rodney#directory-scoped-sessions and https://github.com/simonw/rodney/blob/a842432246f39775ccb14f...)

Why not just give the agent a short lived token with limited access rights?

We have reached the point where we need to control agent access the same way we control human access to systems.

As someone who has never used MCP, or allowed an AI to directly talk to APIs, putting something between a model and a service just seems like common sense. If nothing else, it presents an opportunity to tightly control what the AI is allowed to do with external services, especially if the MCP code is outside of the AI's context. I would go as far as to say it's a necessary security measure.
well if you realease an api you should make sure its usable for ai out of box without 'something in between'. You should always assume ai might be directly calling your api.
And if you don't have a public API, you should always assume AI will be using your UI, reverse engineering your private API, and driving it directly.
I absolutely love when they do that unprompted. I once asked a Claude work agent to go pull comparable apartment listings and somehow it’s subagent reverse engineered like rentcafe or whatever’s private API to get the data. All unprompted.

So yes. Absolutely assume AI agents acting on behalf of their humans are finding all the token efficient ways to get at your sites data.

Ai can't use any api out of box. It requires thing called "harness" to use anything.
The service itself needs to control what the agent can access.

Relying on an intermediary to provide access controls, and that agents will never access the service directly, seems dangerous and naive.

What? Why? More is more but you're arguing a security proxy is a dangerous and naive pattern?
What’s the advantage of MCP over a regular old security proxy?
And it's easy for the service to do that by hosting a remote MCP server.
Agent doesn't have api key to access service directly
The better solution for security is sandbox execution. Most agents used in practice, if given terminal access, can find ways to get around most of the MCP restrictions.
> MCP makes all of that so much easier to provide.

Easier being the key!

Right now it really does feel like early web days where like 5% of the population is adopting things, but most people just get confused and don't participate.

This is why I'm pretty excited about WebMCP in particular!

WebMCP aligns more stakeholders than pure MCP - which seems heavily biased towards model providers. It also seems like it has the potential for a much cleaner ux.

I feel like this would build a false sense of security.

Authentication and access controls must be able to withstand an agent with full shell access. And auditing must be on the server side to provide a full picture of all activity from all clients, whether human or agent.

We must treat agents as clever humans and secure and audit data access accordingly. The era of thinking we can handle agent access to sensitive information differently from human access has passed.

My point about MCPs here is that they provide a way to make those secrets and API keys deterministically inaccessible to the agents - even agents that's have a shell execution environment.

That's the opposite of a false sense of security.

just use an api gateway (ie., a reverse proxy for apis), eg., envoy. This should also be connected to observability and finops style management anyway rather than leaving it to model providers.
I noticed the security aspect to be one of the main selling points in corporations for going MCP (and one aspect that is underrepresented here).

However, maybe I'm missing something but how is defining access rights in an application layer the agent has access to more secure than defining the access rights at the target application itself?

Sure, hiding the filesystem via encapsulation tricks is one way to go, but what is the real-world usage-style here vis-a-vis all MCP usages? I'd wager 1-10% are spending this extra effort, while 90% of corporations basically run Shims so that they can expose APIs for which they have no/bad access rights management to LLMs.

Without tools LLMs can do nothing more than emit text. You are probably assuming a harness is provided to the LLM with shell access but that is absolutely not necessary for LLM usage depending on the use case. For interesting stuff, you do want to provide some tools, but nothing with the power of a terminal if you are worried about security. A MCP server is perfect to securely provide the LLM with some controlled power exactly because it can do nothing at all other than call tools that go through the MCP server and can therefore be scrutinized, audited and ensure credentials are not visible to the LLM.
Also MCP allows exposition, which helps an autonomous agent understand the semantics of the API.

That's the killer feature from my PoV. I just point my agent at a URL and suddenly it knows when, why and how to use it.

Congratulations you reinvented man pages!
The article also doesn't contemplate the "skill distribution problem" which will be addressed by skills-over-mcp.

With the plain ol' API solution, you still need some way for agents to fetch the instructions provided by the service. Of course there are answers for this, but they are not standardized.

With MCP, all you have to do is provide the harness with a single URL and all context can be bootstrapped the same way for every service.

The article misses the point of MCP but he is not wrong that it was a mistake (in some ways)

The mistake of MCP was building it in such a way that it needed to be on a separate process from the API. The statefullness of MCP was such a massive detour for the industry that we will be cleaning up after it for years.

Now that MCP is stateless we can start building what is actually useful: extending API’s for agents.

When people ask me if they should do MCP today, I say absolutely. But because the Oauth protocol side of MCP is very good and is a net positive for all public API’s

This is not a dig at MCP, the original vision of MCP was much different from how the community used it.

> The statefullness of MCP was such a massive detour for the industry that we will be cleaning up after it for years.

I remember discovering this when i wrote my first MCP server. It was like "huh? why would they do that? what use case did they have in mind?". We've spent decades making "internet shit" as stateless as possible on the backend because making it stateful is expensive and complex if you want to have any reasonable scalability. I mean good luck trying to host a stateful service on any kind of commodity serverless "scale-to-zero" infrastructure here in 2026.

Maybe it's because these AI-labs are used to statefulness. I mean LLM-based sessions are hugely stateful if you want any kind of reasonable caching to happen and caching is the only way you can economically scale out LLM's. Seen from that perspective it kind of makes sense why they'd look at MCP and think "hey, why not make this stateful on the backend as well". Statefullness just part of their DNA.

All of that can be solved in better ways than MCP. 1+2+4 belongs to sandboxing. 3 to sandbox UIs. Ideally baked into the next generation OSs.

MCP is the wrong abstraction for all of that. Skills are better as pluggable interfaces, and I predict[0] chat and agentic loops will converge eventually (Anthropic already did this correctly; OpenAI, it's your turn) and skills marketplace will replace MCP in its current form.

[0] Where "predict" = "hope". The best solutions are often not the winners.

Even for coding agents. MCP is the only unified interface INTO harnesses. This pattern has not struck yet for most, but will soon in the coming months given the MCP spec. MCP has been about models calling tools., it is about to become some form of inverse and MCP will be the only actual way to integrate into the models. Simple example is sending an event to the model (today models have to poll)... these coming updates will cement MCP as permanent infrastructure.
What matters is local server or not. Using a remote server means exposure.
Agree. MCP also helps in some esoteric use cases like interacting with legacy windows applications over COM. I created an MCP for Outlook 2019 desktop version and COM was the most straightforward way to let my agent interact with my mails for classification.
MCPs just misses what AI is. If you think MCP is a good idea you are confused what is going on. AI is beyond MCP. MCP is SOAP of AI.
> The MCP Industrial Complex

Was that a real thing? I mean it must've been for it to be mentioned there, but, rephrased: what was the scale of that?

How many individuals were involved in that? 1? 10? 100? 1000? 10000? 100000?

MCP lead to one good thing though

a lot of websites that never bothered to provide a REST API are now exposing MCP server because it has become popular, and you can use those servers to write normal automation for yourself, without plugging in any LLM etc

Also dynamic oidc registration (client.dev) has became better supported on oidc idps to support MCP
You have to throw in “load bearing” in the request though to avoid suspicion
The same thing is going to happen at the app level on Google and Apple platforms. Android tried to have modular cooperating software, but both the early dominance of iPhone in the developer community, and developer interests in having monolithic apps made that effort fail.

Now both Google and Apple are making app interfaces legible and tool calling discoverable. This will change apps in interesting ways.

I don't think MCP is a bad idea, but using them incorrectly is.

CLI tools are great if you always use the same environment. But try using them from your iPhone, and they simply won't work; a remote MCP will work seamlessly.

HTTP APIs solve a different problem. APIs are designed to be predictable and consistent, so the client always knows the response shape in advance. The MCPs are designed to be dynamically discovered. This lets agents connect to new and unknown ones.

Trying to give APIs extra responsibilities so they can replace MCPs would just create more confusion. It's like creating an MCP server but calling it an API.

This issue isn't just about communication methods and technical details. It's also about “standards”, and it will become increasingly important over time. As we begin to integrate AI into everything-for example, into banks...
MCPs are a bad idea because they encourage token burn at runtime when a deterministic program that uses the API should be used.

Of course Mario, Sammy, Jensen et al. would be for this.

Some people can't appreciate that their local Claude code workflow isn't the only MCP usecase. If you don't need it, then don't use it.