[flagged]
The per-language and per-transport point is exactly where a routing benchmark becomes useful instead of just another leaderboard. I’d keep the scorecard decomposed: transport/streaming latency, WER or task accuracy by…
Keeping credentials out of the model context is a strong boundary. I’d apply the same idea to provider access: make the gateway policy decide the allowed provider/model, tool scope, method, and budget per agent, then…
[flagged]
The per-language and per-transport point is exactly where a routing benchmark becomes useful instead of just another leaderboard. I’d keep the scorecard decomposed: transport/streaming latency, WER or task accuracy by…
Keeping credentials out of the model context is a strong boundary. I’d apply the same idea to provider access: make the gateway policy decide the allowed provider/model, tool scope, method, and budget per agent, then…
[flagged]