195 comments

[ 0.20 ms ] story [ 48.9 ms ] thread
(comment deleted)
so much MCPs in the FTA yet not a line about what MCP actually is.
It's just like API-s, but with built in documentation.
If you don't know MCP by now that's between you and google
I appreciate their reluctance towards MCP, but /something/ is better than nothing.

It’s suboptimal for the reasons the author outlines: but so is USB-C. So is NVME, so is HDMI.

We use these hugely successful technologies in spite of their flaws because they’re widely compatible and easy for the end user.

That’s why MCP is everywhere. It might not be performant, robust and uniform but it WILL get better over time.

And I’d much rather have the broad MCP ecosystem that we have now than seven or eight different “optimal” ways of plugging in an LLM to something useful.

Agreed. I personally don't use any MCP but I think it was a good move.

I hope they'll do the same and eventually add native support for ACP (https://agentclientprotocol.com/get-started/introduction) which, on the contrary, I use quite.

Arguably, adopting ACP might even help Pi’s case, in that it could escape the terminal interface into one of many wrapper GUIs. TUIs inherently tend to limit your userbase to those who know what a terminal is…
The "general adoption agent" for Earendil is their other product Lefos, that is based on Pi and uses email as the interface.
Good points.

About what Pi/Earendil is going for, I can't really say but a while ago they created quite a stir in the Pi community for adding a trust system[0] which for many (me included) went against the loudly advertized "yolo" phylosophy, at that time I speculated it was a move to make it more palatable for the general population (whatever that actually means), so I'll stay optimist for ACP adoption for now.

[0]: https://pi.dev/docs/latest/security#understand-project-trust

I’m not sure; the integrated GUI seems like a major differentiator for them.

Pi’s agent is supposed to be simple, and a simple ACP agent is like a couple hundred lines of code. Making a system that allows UI plugins is way harder.

Also not sure if you’ve seen but you can get ACP from Pi with https://github.com/svkozak/pi-acp It bridges Pi’s RPC mode to ACP, works okay but not amazingly. My thinking level selector in Zed has never worked with it but everything else I use has worked (I’m sure other things don’t but I must not use them).

FYI Pi can do RPC over stdio (though of course it's a bespoke thing, rather than standardised). I use Pi every day; never used the TUI (I forgot it even had one).
You could already use MCP perfectly well on pi via extensions.

I'm not so sure about this move, or the general inclusion of code mode in the core editor as one of pi's main selling points was its minimal nature.

Except there already was "something" that the creators of MCP just ignored.

The ai clients could have given us a way configure them with OpenAPI specs for APIs that support Oauth2. That would grant it access to use the API specified in the spec, the harness would walk you through the Oauth flow and securely store the token, and then inserts some tools for discovering the API methods and making requests to it in the context. The harness would handle inserting the auth token once the agent had formulated a request (basically, exactly the same way MCP is used today, except with all you need is an OpenAPI spec).

And we'd have a rich of ecosystem of half decent REST APIs, instead of janky new standard that's half implement slightly differently by every harness/chat client.

Yeah I'm not a huge AI bull but I've just never understood why MCP needs to exist when OpenAPI could've just been extended
Anyone found a good file upload solution for MCP? Or is the best practice to use HTTP to upload files outside MCP?
There is a MCP SEP that outlines multiple ways, that I hope will sooner or later be accepted: https://github.com/modelcontextprotocol/modelcontextprotocol...

While we are waiting on that to become stabilized, we implemented a inspired/co-evolved way to do that in our tool[0], where you mark individual fields in the request/response schema as being file payloads, so that file exchange can be properly orchestrated by the harness and doesn't pollute the context. We just do inline base64 uploads of the required payloads, which in practice we've seen to work quite will until ~100MB files (which is otherwise also the size limit we usually recommend for file processed).

It's annoying that it's not stabilized yet, but for most bigger customers we've seen, they implement 80% of the MCP servers they connect in-house, so doing adjustments to the tool surface, and metadata has been less of a pain for them than we expected.

[0]: https://erato.chat/docs/features/mcp_servers#file-support

I have had to hack a couple of workarounds in https://github.com/rcarmo/memento to do uploads, and there's a draft going around, but the general practice in enterprise MCPs seems to be to do it "out of band" and have MCP tools to hand-over storage handles/URLs so the MCP server can do the imports itself "safely".
I just implemented this with oauth - basically I have a tool that returns a signed URL that the client can PUT the file to. Works in ChatGPT Work mode. My use-case was to have the client generate an image and upload it.
> And while we could have just wired up the metadata to enable better MCP extensions, we also think that MCP with Codemode solves quite a few of the issues that it traditionally had.

There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.

It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.

Cloud products based orchestrations with proper security mechanisms configured, don't have shell access and should only communicate over proper network mechanisms.

Rootless immutable containers without shell access, or SaaS products from multiple vendors with WebAPIs as the only touch point.

Yeah, I'm probably over-indexing on local use cases due to my own preferences, prejudices, biases etc. For a coding agent like Pi it does seem reasonable to expect some kind of shell access though, unless some people are using it as a CLI chat client with MCP?
I happen to think codemode is useful, but not the full answer. I have a long and skewed history with chaining things in MCP and built a dozen or so enterprise ones (see https://taoofmac.com/space/blog/2026/04/29/2341 for notes) and it all falls back into the trade-off between agent scope/context and tool coverage: If you are using a coding agent it will have no trouble sorting out any tool regardless of how many are exposed (it's just a matter of either progressive tool disclosure or good tool metadata, since the coding agent will just go at it and expend whatever tokens are needed), whereas in a "normal", limited, scoped agent that has only a few things it needs to do (like handling a ticketing system) codemode is pretty much overkill.

Pi is primarily a coding agent, so yeah, code mode makes sense, but I've found that better MCP design saves everyone a lot of trouble and would also probably have improved the thing's reputation overall (I personally am not fond of the line protocol, would rather have protobuf and more typing, but it is what it is).

Honestly, the provided argument for it is rather weak. They are basically adding a way of running scripts that are contained within harness to execute harness's own tools (that's the Codemode). A coding agent can already compose any arbitrary logic by invoking shell scripts (or python scripts, or node scripts), etc - so this is just entirely unnecessary in the core, from my perspective.

If you feel that Pi has been drifting away from its original vision, try hax (https://usehax.dev/) - you might like it.

> Key Features: > > Respects your terminal — Streaming Markdown and live tool output, reflowed for display in the terminal. Only redraws the current streaming line or the input area, native scrollback is preserved. Does not take over or mess with your terminal.

Thank you. I've been frustrated by harnesses hijacking the terminal and breaking basic features such as scrolling and text selection.

It even sends BEL when the agent completes, which makes so much sense, yet Pi never implemented it.

I'm definitely going to use it over the next few days and hopefully make the switch.

As far as I can tell MCP is just "we bothered to document our api in a programatically readable way".

Just generate CLI tools, with docs, from MCP servers on demand.

In enterprise integrations, that is just not an option. MCP has pretty much taken over there.
It's exactly that, but I don't see the issue. API+Docs under a single URL looks like a win for me. It also warrants a new name.
Theres transport and context management. A cross agent status cache with semaphore over a testing harness driving a browser that can rewind and retry is much easier for agents to drive from mcp than selenium or whatever is fashionable these days. Didn't replace unit testing per se but to rca and fix it's a much better tool.
This is somewhat similar to HuggingFace smolagents where the model writes code that calls tools, instead of emiting json to describe the tool call per turn. Here Codemode is one tool that the model calls when it needs to compose many tool calls, especially MCP ones. Is what i understand of this.
I feel the same way about needing support for sub-agents, those feel pretty foundational to me.

I suspect that a smart model driving multiple dumber models for work and then using sub-agents with the same smart model for adversarial review will be a pretty common pattern.

Personally, I got a bit confused about Pi having most of that stuff as plugins since I remember how much of a mess Eclipse was where so much was just loosely fitting together plugins and just went with OpenCode since it covers most of my needs out of the box. Guess that might also be a sign of me getting older, because my IDEs and desktop environments are all closer to stock too.

Some things are impossible to just tack on or work around though, like MCP, while other things, can be done by just composing stuff.

Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.

What would you say is a good harness with subagents?
I found the MCP extension for Pi to work fine.
I had the model write up a pi extension for logging the transcript to the syslog, and all on its own initiative it spawned a subagent to generate test output. It decided to spawn a lighter model for this trivial task, apparently unaware that I can only fit one model at a time so my llama-server ended up thrashing to the lighter model then back to the original model. My own personal n=2 semi- rogue swarm
I did the same, also, the fact the tools evolve so fast, I dont want to waste time on a particular one while it might be obsolete next week. So either it works now, other I pick something else.
Hah, it's the complete opposite for me :D. In Claude Code I disabled all sub-agents stuff, disabled nearly all tools but Bash, Edit, Write and WebSearch and replaced WebFetch with my own tool that doesn't summarize anything because the results were always worse with sub-agents, they always lack the necessary context and weaker models summarize bad. I also replaced the system prompt with my own that cuts a LOT of tokens, agents don't need a 10k+ system prompt anymore.
That's interesting! You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy? Also what about the cache read costs for longer context sizes?

Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:

  /autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).

Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/

> You don't have cases where the main session has important planning stuff but the actual work to execute has so much crap in it that context compaction will probably dig into the important plan stuff too much and make it too lossy?

For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system.

What’s the difference between sub agents and separate sessions?
I'd assume separate sessions are not aware of each other, sub-agents are spawned by an orchestrator agent?
Almost correct, separate sessions can communicate with one another. In my case, planning and coding communicate by writing files locally.
And subagents don't always communicate with each other, AFAIK that is the most common case.
> For that is it not better to have separate sessions for planning stuff and doing actual work?

Personally, seems like too much effort for something that would still need to (and fail to) have some sort of a link between the two, so I could go from the planning over to implementation and back easily. In reality, that'd get lost in the noise of dozens of sessions - I mostly just want the harness to help me do work and otherwise get out of my way, not make me dance around it. Ergo, the more context management it handles, the better!

Eh, I find managing sessions in Pi extremely easy. In fact, I customized how it mages session as part of my workflow, and how agents in different sessions communicate with one another.

It's part of why I really dislike to work in Claude Code, and found it too unwieldy. There I have to keep dancing around it to manage the context in a sensible way.

separate session need more hand holding, i guess?
running local models, now with Qwen3.8-Flash-Next, they have 256k, but when they get up there their speed is just too slow. So i've taken https://github.com/Tarquinen/opencode-dynamic-context-prunin... and started improving it. It already worked well to get a lot of mileage out of just taking tool calls, code modifications, etc, and dumping them in favor of a summary.

But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage.

I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again.

It's amazing how resilient it is if you give it a good plan. The work flow has basically been:

1. Write up an implementation document for some new set of features.

2. Rewrite the implementation as a TDD document

3. Set it to work.

The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done.

> taking tool calls, code modifications, etc, and dumping them in favor of a summary

Doesn't that destroy the cache? I find that caching significantly sped up my Qwen, especially on said larger contexts.

Yes-ish. The compressed summary sits atop the cache stack, so if we got to 96k, it'll take say 30k, compress to 10k, and that 66k+10k is the new stack, so 66k is still cached and retrievable.

So it is designed like a heap, where we're taking raw context off the heap, compressing it, and putting it back on the heap. So cache during compression is mostly unperturbed, since we're rarely digging all the way to the bottom of the stack, but that could happen.

Eviction though is cache busting; but again, I'm valuing the session's roadmap as the valuable product and context size slows computation size, so I have to bust the cache to sacrifice immediate re-processing for longer term compute speed up.

Because that's faster than getting to the end of the context (remember, every 1k adds to the compute time of the next 1k). So speed at 200k is much lower than at 100k. It's also local, so I'm only paying time+watts for the trade off. As far as I can tell, speed is not being lost since if I let the context grow, the kv cache doesn't help with the compute throughput.

So, yes, but it's "smart"; we're only busting it at the top of the context, so rebuilding it isn't from the bottom up, it's just at the top. Those summaries sink on the heap until you get to the eviction limit, and then, they're evicted, and we rebuild from some intermediate place in the heap.

The benefit of it all is I can have lots of projects, and keep a single session that tends to have the context necessary to avoid having to write AGENTS.md or other context bloats. Set large implementation goals and come back to them as needed, etc. I've had it running like this for awhile and it seems Qwen3.8-Flash-Next has no trouble understanding the rolling window.

I never hit compaction, most of my sessions are 150-300k tokens long with the longest being around 700k. Using sub-agents means that they can't use that cache, have to re-read everything and now multiply this for every sub-agent you call and it just wastes money/tokens. I also don't like how intransparent sub-agents are, I can't follow what they are doing and I can't really steer them. Claude Code has the /agents view but it's clunky and awful to use.

GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).

I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.

I'm only using the 20$ plans.

I admit I'm not a heavy user of subagents but isn't one of the standard use cases for subagents to run a single command, take the output, summarize it for the main agent and pass it up instead of polluting the context?

When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.

That sounds like you just NIH'd mr Zechner's Pi Coding Agent. Those were basically its founding design: yolo-mode security, simple design, minimalist system prompts, plug-in based for anything fancy (even sub-agents and web).
Yep, I've used Pi in the past (see my other comment https://news.ycombinator.com/item?id=49908276) but I don't like the direction it went (selling out, forgetting their principles/throwing them out). And since Anthropic wants you to use Claude Code with their sub I just switched to CC again. I've only used Pi for a few months though when GitHub Copilot gave you 300 requests for like 10$.
oh-my-pi is a fork of Pi that adds a lot of this stuff

https://github.com/can1357/oh-my-pi

I haven't tried it much though, can't vouch how well it works.

Exceptionally well is my takeaway. It’s the only harness I am using these days. I was previously using pi and codex mostly, but also the ones built into editors like zed, vscode, and the jetbrains IDEs.

On top of that, somewhat unrelated I’ll agree but still, it has support for vim keybindings

Oh thank you for mentioning it has vim mode that's been my only gripe and I had no idea it had it. don't know if it's new or not never noticed the setting.
Pi vs omp is hotly debated within my friend group. It has most things you could want, ready out of the box, but also a lot of things you'd never want and it's constantly 5% broken. Some people love that trade, others don't.
I used it for a while, and I thought it was quite hard to follow what was going on. And in the end it just went off the rails anyway, though that might have been a GPT-6 kind of thing.
It's actually the main reason I chose Pi.

I did create some extensions where it spawns sub agents for specific tasks, especially when I want to keep the context clean or when I really want to offload a piece of work to a cheaper model. And for that I have a high degree of control over, I know which model is being used for each subtask.

I find Claude Code too unwieldy for my tastes. Pi's philosophy of being very light on features nut highly flexible for customization, clicked very well for the way I work.

An article pretending to address its title, but just beating around the bush.
I had no idea pi didn't support MCP! I'm a new user, I just started messing around with it. I was getting my tooling up and running and tried to get one of my database MCPs working (Which, in retrospect, seemed a little painful - but I guess I was under the assumption that it was my responsibility to build + maintain those connections).

Another retrospect note, "No MCP" appears to be the first icon on their front page - not sure how I missed that.

Imagine my surprise reading this!

Pi doesn't even support a permissions model. It's extremely barebones, you're expected to customize basically everything.
... and after using it for a while you might end up forgetting most of the extra stuff you thought you needed in the first place. That's what happened to me and I've never looked back and still am a happy Pi user (https://a.l3x.in/ai if you're curious)
I have already basically removed any extensions I had installed. There are still a couple, but I am going to remove those as well. Except maybe a web search one as that is something I use often. Going to remove all skills and anything else as well. I am a simple man when it comes to agent programming.
I was enticed by the promise of it being barebones but I still found it too bloated

So I cut out a lot of functionality in my own fork

https://github.com/Pyrolistical/mi

I took “expected to customize basically everything” to heart

> The first thing to remember is that the world is not static

And you didn’t remember that when you said no to MCP?

No, no MCP for now?

Generally, I use skills with a CLI tool instead of MCP and tools. Usually in most cases I also get a coding agent to generate the CLI tool.

I find this approach is easier to debug and I can also use the tool myself to ensure it's working well.

Good. I too I'm not a fan of MCPs, but these days I do find them useful. In Claude Code I connected to my company's MCP which made Claude Code infinitely more useful for everyday work stuff
So Pi is also accruing cruft now :(
You can turn those tools off. In fact, that is what I am doing right now in https://github.com/rcarmo/piclaw until I am positive the new MCP stuff has full parity with the MCP adapter I've been shipping for the past six months or so.

I'm actually pretty happy that they did it, since 90% of what I have to integrate in enterprises is MCP-driven (it's a security and auth boundary that has become pretty much mandatory for any third-party agents wanting to reach into corporate data) and this lets me use Pi directly. Am just being cautious about the first version, because, well... it's a first version, and I like my tools stable.

(I actually played around with the idea of using QuickJS myself for codemode, but since I rely on Bun that gives me the ability to use other things... never got around to do it though.)

To be clear: absolutely not the plan and nothing is loaded by default that was not loaded before.
Ok! I trust that you and the maintainer will steward the project properly. It's just that I really like Pi as is and am a natural worrier. I'm also not sold on Jev(-likes), so that reasoning rung a bit hollow to me.
I think Jev is worthy of your reconsideration, the OP article has a pretty good use-case for it right within Pi.

It's the cost efficiency that's a big deal. Jev is an excellent cheap "first pass" model.

I'm still confused about what this codemode is. Models have been trained to chain bash and other typical unix tools well. They're so good at that to an uncanny level. Why do we want to not utilize this ability? Is it just a permission management issue in case you don't want the model to use shell directly?
Codemode is a fancy name some MCP authors coined for the practice of providing scripting/method chaining for their MCP tools. It's generally implemented by providing some kind of code execution tool, the LLM calls it with a script, and the MCP server runs it in a sandbox.

It's pretty effective because of the reasons you noted, but there's a composability problem since each MCP has its own sandbox and can't call into the other ones.

IIUC Pi offer a workaround for this, the harness runs the sandbox and populate it with the MCP tools, that way the composability problem is solved and every MCP do not have to implement their own sandbox.

My understanding is that code mode is supposed to be implemented by the harness, not the MCP provider. You chain multiple MCP providers as well as other harness provided tools inside the sandbox.
My timeline might be wrong (I remember a cloudflare article mentionning the "in MCP" case), but anyway yes there's tools to do it in the harness now and it's the better idea.
the conversation seems to dwell on things you could substitute Bash for but the real need stems from completely opaque systems that nothing can reach but which are now getting MCP support. This is where being left out of having MCP support will hurt. I'm still quite happy to let all the harnesses compose bash commands to their hearts content (inside their sandboxes ...)
As many others have commented: good, I too am an mcp hater.

However- in my testing, mcp is really quite fast, and its pretty much free at this point- with frontier models. Context rot is, from what ive tested, not as much of a concern now. I genuinely was not able to hillclimb skills/extensions to beat out the speed of mcp in some cases I've been testing.

I'm a bit bumped when I first saw pi is moving from bash to codemode and also adding MCP, since I thought that loses the purity and simplicity of the "use bash for everything" philosophy. However, after reading more about it, I realized codemode is just a slightly enhanced version of bash: more complex, sure, but likely more robust, secure, and efficient. For those who, like me, don't get the point of this change, here is how I understand it.

The first tool execution runtime in harnesses are direct tool calls with JSON or XML, such as the Read and Edit tools. As an escape hatch, we have Bash tool that allows arbitrary code execution on the host running the agent. The downsides of using bash (on the host) as the main tool execution runtime are:

- Syntax and obvious errors only surface at runtime

- Unergonomic orchestration of parallel and background tasks

- Verbose command output cluttering context

- Dependent on the host environment, packages versions, etc.

- No security measures by default.

To me the last point is the biggest inherent weakness, usually mitigated by creating a dedicated unprivileged user or running bash in a sandbox.

Note that direct tool calling is kind of the polar opposite on these points: syntax errors are caught early, orchestration can be done with some wrapping tools, command output is controlled, and most importantly they are more sandboxed. On the flip side, they obviously have way less power, necessitating Bash tool in the first place.

Codemode is the middle ground between these two extremes. It actually can be derived simply by one idea: what if we replace Bash by another language that can be checked for obvious errors, i.e. type checked?

Everything else falls out from there:

- Any language would do, but I think TypeScript fits the balance between safety, speed, conciseness, and popularity in training data.

- If we use TypeScript, might as well run it in a sandbox as JS runtimes have been designed with this in mind for 20 years

- Orchestration comes for free from the JS runtime. It's not more powerful, just more ergonomic.

- Since the tools are controlled by the harness and not dependent on the host, cloud agent becomes easier.

- With this in place, MCP are not very different from a tool provided to this sandboxed runtime.

Overall I find the benefits compelling enough, but we'll see if the heavily-RLed models these days will use it effectively.

The core idea of Codemode that I understood is - The script doesn't actually write any code itself; rather, it acts as a workflow automation and program management tool. It pulls data from an issue tracker, delegates the analysis to a model, and synthesizes the results to help manage Pi's development priorities.

Codemode isn't replacing MCP; it's fixing MCP's biggest flaw—its lack of composability.

But I want mayors, animals, gas reservoirs and data lakes. Can Pi offer in game purchases?
Interesting - I didn’t even notice it’s not supported. Must have been added via extension because I definitely have mcp active on pi.

The idea of using jev as a cheaper faster subagent for specific use cases is interesting. Will have to experiment with that!