1 comment

[ 2.3 ms ] story [ 13.2 ms ] thread
Scale LLM serving with programmable cross-engine serving patterns, all in a few lines of Python