Re-quantizing a local LLM 14x faster by skipping the tensors that didn't change (andreaborio.substack.com) 8 points by andreaborio 3mo ago ↗ HN
1 comment
[ 3.2 ms ] story [ 9.4 ms ] thread