1 comment

[ 1393 ms ] story [ 1847 ms ] thread
The model was quantized to 8, 4, 2, and 1 bit. Characteristics:

• Q8: 8-bit 1.56 TB, lossless

• Q4: 4-bit, 1.51 TB

• Q2: 2-bit: 861 GB

• Q1: 1-bit, 594 GB

The smallest Q1 model keeps 78.9% accuracy, while being almost 3 times smaller than the original one. Instruction for running the model is in the model's card