1 comment

[ 0.32 ms ] story [ 14.4 ms ] thread
KVQuant: Towards Enabling 10 Million Context Length For LLM Inference through KV Cache Quantization