Show HN: Relay – a self-hosted LLM gateway with smart routing and request pacing (github.com)
This started when I was trying to string together several providers’ free tiers. I kept hitting rate limits at different intervals, which broke some of my agent clients. Some agents were also getting greedy with shared resources, so I needed a way to manage how they used the available capacity.
That led to Relay’s queue-first approach. Sometimes it’s better to wait a second or two for your preferred model than immediately fall back to another one. Relay queues and paces requests against configured provider limits, aiming to make use of available capacity without repeatedly hitting rate-limit errors or needlessly falling back to worse models.
Relay’s classification model also looks for signals that a request needs specific capabilities, such as coding or more complex reasoning, while classifying the request’s main intent before routing anything.
One thing I’ve obsessed over is keeping that decision layer cheap. Relay’s built-in classifier runs in the single digit millisecond range. It has a deliberately narrow job: classifying LLM requests and helping decide where to send them. It probably won’t be playing DOOM, but that’s a trade off I’m happy with for a routing layer. In the routing tests I've run so far, the built-in classifier is considerably faster than Laya while producing broadly similar routing decisions. Working on getting Jev up and running, and will report back to see how that stacks up as well.
The community version is available now, with a public repo, a built-in dashboard, and a local classifier. It’s written in Go, and you can run it with npx @anchorshell/relay or build it from source. The gateway itself is lightweight; it also ships with small classification models you can use locally.
There’s also a hosted version with a free tier if you don’t want to run it yourself. It offers our more capable classification models, along with team features, and separate limits for individual agents.
5 comments
[ 2.4 ms ] story [ 25.4 ms ] threadOn the first request in a conversation, we pick the model/provider based on the request, and then keep that route sticky for the rest of that conversation. So we’re not bouncing a large context between providers every turn and constantly blowing away the cache. If that provider fails and we have to fail over, then yeah, the next provider has to prefill the context again. There’s no real way around that.
If you need to handle a separate task mid-conversation, the cleaner pattern with Smart Groups is to spin up a sub-agent with its own context. That request gets evaluated separately, and then that route stays sticky for that task too.
Also, smart routing is only one part of Relay. You don’t actually need to use dynamic routing at all to get value from the rest of the gateway, but it is a cool offering.