Gemlite: Towards Building Custom Low-Bit Fused CUDA Kernels (mobiusml.github.io) 47 points by un_ess 2y ago ↗ HN
[–] apsec112 2y ago ↗ Weird that they don't mention Triton? I only skimmed it, but I'm not sure what the pros and cons would be vs. Triton, which is the tool I'd use if I wanted custom quantized inference kernels.
2 comments
[ 2.7 ms ] story [ 15.4 ms ] thread