Sequoia: Speculative decoding boosting LLM inference by 8-10x (infini-ai-lab.github.io) 3 points by fgfm 2y ago ↗ HN
0 comments
[ 3.3 ms ] story [ 11.4 ms ] threadNo comments yet.