1 comment

[ 3.5 ms ] story [ 7.4 ms ] thread
TLDR; Make 1x1 convolutions sparse, write fast Sparse Matrix Multiplication kernels, get a nearly 2x speedup with smaller models.