1 comment

[ 3.1 ms ] story [ 14.3 ms ] thread
We study how to approximate the famous shortest-job-first scheduling in LLM inference!