Pipeline-parallel LLM inference across GPUs on separate machines (github.com) 5 points by ngaut 2mo ago ↗ HN
0 comments
[ 5.3 ms ] story [ 13.0 ms ] threadNo comments yet.