Transforming LLMs into parallel decoders boosts inference speed by up to 3.5x (hao-ai-lab.github.io) 7 points by snyhlxde 2y ago ↗ HN
0 comments
[ 3.5 ms ] story [ 18.0 ms ] threadNo comments yet.