Fleet: Hierarchical Task-Based Abstraction for Megakernels on Multi-Die GPUs (arxiv.org) 10 points by matt_d 1mo ago ↗ HN
[–] joshuakelleyds 1mo ago ↗ Nice read! Cool to see more GPU programming models that expose chiptet topology instead of treating it as a flat execution. Reminds me a little of Cerebras' CSL although far from as extreme as that.
1 comment
[ 3.3 ms ] story [ 18.6 ms ] thread