1 comment

[ 2.8 ms ] story [ 66.7 ms ] thread
Run a 27B reasoning model locally on a 16GB M2 Mac. Ferrox + ternary quantization delivers 5.9GB models without sacrificing much quality.