Fusing a 27B ternary LLM's whole decode step into one CUDA kernel (twitter.com) 3 points by Jr23_xd 2mo ago ↗ HN
1 comment
[ 1.1 ms ] story [ 12.0 ms ] thread