Medusa: Framework for Accelerating LLM Generation with Multiple Decoding Heads (sites.google.com) 1 points by azeirah 3y ago ↗ HN
1 comment
[ 3.2 ms ] story [ 143 ms ] thread