Understanding Multi-Head Latent Attention (From DeepSeek) (shreyansh26.github.io) 2 points by shreyansh26 7mo ago ↗ HN
[–] shreyansh26 7mo ago ↗ A short deep-dive on Multi-Head Latent Attention (MLA) (from DeepSeek): intuition + math, then a walk from MHA → GQA → MQA → MLA, with PyTorch code and the fusion/absorption optimizations for KV-cache efficiency.
1 comment
[ 4.8 ms ] story [ 10.6 ms ] thread