Introduction to Flash Attention – improving the efficiency of LLMs (hopsworks.ai) 4 points by javierdlrm 2y ago ↗ HN
1 comment
[ 0.28 ms ] story [ 14.7 ms ] thread