[–] free_bip 2y ago ↗ Specifically this is Llama2, not Llama3, was a bit disappointed from that. Also wasn't totally clear from the article - will this actually increase GPU inference speed / decrease GPU memory usage?
1 comment
[ 1.9 ms ] story [ 12.9 ms ] thread