Continuous batching enables 23x throughput in LLM inference (anyscale.com) 2 points by richardliaw 3y ago ↗ HN
0 comments
[ 1.9 ms ] story [ 12.6 ms ] threadNo comments yet.