Skipping 90% of KV dequant work speeds up LLM decode by 22% (github.com) 1 points by pidtom 5mo ago ↗ HN
1 comment
[ 2.9 ms ] story [ 17.0 ms ] thread