broyojo

↗ HN profile [ 18.9 ms ] full profile
Karma
0
Created
()
Submissions
0
  1. This runs the smallest llama2.c checkpoint (stories260K) inside Scratch/TurboWarp by compiling C inference code into Scratch blocks using llvm2scratch. The model is quantized to Q8_0 and packed into Scratch lists. If…