With setup you mean HW or the SW stack? We used to run GLM-5 class models but have now changed to smaller ones as we're able to serve more concurrent users with our limited hardware (DeepSeek-v4-Flash-0731,…
Yes, agreed. I meant in general it's usually new models I'm seeing the biggest issues with. I hope the refactors on model arch/config will help with this so model changes have a smaller blast radius. It's a lot of…
My team runs open models for devs at our company, mostly on H200s, and I'd also say yes, if you want to always stay on the bleeding edge (and not use their model-specific images they publish before it lands in a…
A part of me wishes the open source community would focus making research and industry-backed initiatives like the vLLM Semantic Router rock solid. Then I'd spend less time every month checking if this or that new model…
Yeah, I think also efforts like Docling and similar are showing that smaller, specialized models with the right tooling might be more effective (and efficient) at this than throwing everything at Opus. They don't seem…
With setup you mean HW or the SW stack? We used to run GLM-5 class models but have now changed to smaller ones as we're able to serve more concurrent users with our limited hardware (DeepSeek-v4-Flash-0731,…
Yes, agreed. I meant in general it's usually new models I'm seeing the biggest issues with. I hope the refactors on model arch/config will help with this so model changes have a smaller blast radius. It's a lot of…
My team runs open models for devs at our company, mostly on H200s, and I'd also say yes, if you want to always stay on the bleeding edge (and not use their model-specific images they publish before it lands in a…
A part of me wishes the open source community would focus making research and industry-backed initiatives like the vLLM Semantic Router rock solid. Then I'd spend less time every month checking if this or that new model…
Yeah, I think also efforts like Docling and similar are showing that smaller, specialized models with the right tooling might be more effective (and efficient) at this than throwing everything at Opus. They don't seem…