[dead]
[flagged]
handy, but the gap most of these filters have is that "fits in VRAM" doesn't mean usable. context length blows up the KV cache fast, a 7B that fits at 2k tokens will OOM at 32k. factoring context len + quant into the…
[dead]
[dead]
[dead]
[dead]
[flagged]
[flagged]
[dead]
[dead]
[flagged]
[dead]
handy, but the gap most of these filters have is that "fits in VRAM" doesn't mean usable. context length blows up the KV cache fast, a 7B that fits at 2k tokens will OOM at 32k. factoring context len + quant into the…
[flagged]
[flagged]
[dead]
[dead]
[flagged]
[flagged]
[dead]
[dead]
[flagged]