Throughput Is Not All You Need: Maxing Goodput in LLM Serving via Disaggregation (hao-ai-lab.github.io) 5 points by zhisbug 2y ago ↗ HN
[–] zhisbug 2y ago ↗ New work from the vLLM team that disaggregates prefill and decoding to maximize goodput (throughput subject to latency constraints) in LLM serving
1 comment
[ 3.4 ms ] story [ 13.2 ms ] thread