TL;DR: By instantiating a CUTLASS s16816gemm_f16 kernel with S=3 pipeline stages, 256×128 threadblock tiles, and epilogue vectorization width 4, I achieved 31.9 TFLOPS FP16→FP32 on an RTX 4060 (AD107) at N=8192 — 14.3% faster than the same GPU's cuBLAS baseline of 27.9 TFLOPS, and 52.7% of the 60.55 TFLOPS theoretical tensor-core peak.
1 comment
[ 0.29 ms ] story [ 10.3 ms ] thread