Pipeline-parallel LLM inference across GPUs on separate machines (github.com) 5 points by ngaut 2mo ago ↗ HN
0 comments
[ 4.5 ms ] story [ 17.5 ms ] threadNo comments yet.