1 comment

[ 3.4 ms ] story [ 13.2 ms ] thread
New work from the vLLM team that disaggregates prefill and decoding to maximize goodput (throughput subject to latency constraints) in LLM serving