Native Sparse Attention: Hardware-Aligned and Natively Trainable (arxiv.org) 2 points by teepo 1y ago ↗ HN
0 comments
[ 2.5 ms ] story [ 10.8 ms ] threadNo comments yet.