1 comment

[ 5.7 ms ] story [ 18.9 ms ] thread
Heavily optimized single C file that can train the same model as Karpathy's microgpt to lower loss in under a second on a single Mac core.