1 comment

[ 3.6 ms ] story [ 9.6 ms ] thread
LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython.