B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams (iaroslavelistratov.github.io) 2 points by magoghm 13d ago ↗ HN
[–] xiphias2 12d ago ↗ Pretty cool, it's missing how these concepts map to ThunderKittens code (my favorite CUDA toolkit)
[–] peter_d_sherman 12d ago ↗ One of the best AI GPU Attention Kernel how-to-implement articles I've read in a long time!Tremendous work, tremendous effort was placed into writing this article -- and it shows!Well done!
3 comments
[ 1.7 ms ] story [ 15.4 ms ] threadTremendous work, tremendous effort was placed into writing this article -- and it shows!
Well done!