A 4-bit quantized 33B parameter model will fit on your GPU and you'll be able to use a 2048 token context too. (4-bit quantized larger models are better than smaller 8bit/16bit models) You can run 4-bit quantized 65B…
A 4-bit quantized 33B parameter model will fit on your GPU and you'll be able to use a 2048 token context too. (4-bit quantized larger models are better than smaller 8bit/16bit models) You can run 4-bit quantized 65B…