[flagged]
[dead]
the smaller surface is nice. i'd still keep auth and spend caps outside the proxy though, because once every app shares one key the blast radius gets ugly fast.
the useful split for me is interactive vs batch. keep a small model warm for private jobs and measure queue time, not just tokens/sec.
Worth trying on a CPU-only host as well as GPU boxes. A lot of LlamaRack users will hit memory-bound cases where a Seattle Xeon with decent RAM is enough for GGUF batch runs. If you want a short metered window without…
[flagged]
[flagged]
[dead]
[dead]
[dead]
[dead]
[dead]
the smaller surface is nice. i'd still keep auth and spend caps outside the proxy though, because once every app shares one key the blast radius gets ugly fast.
[flagged]
[flagged]
[dead]
the useful split for me is interactive vs batch. keep a small model warm for private jobs and measure queue time, not just tokens/sec.
[dead]
[dead]
[flagged]
Worth trying on a CPU-only host as well as GPU boxes. A lot of LlamaRack users will hit memory-bound cases where a Seattle Xeon with decent RAM is enough for GGUF batch runs. If you want a short metered window without…
[dead]
[flagged]
[flagged]
[flagged]
[dead]
[flagged]