Faster Mixtral inference with TensorRT-LLM and quantization (baseten.co) 2 points by tikkun 2y ago ↗ HN
[–] tikkun 2y ago ↗ TensorRT-LLM plus int8 quantization = cheaper, minimal quality loss. Some nice tables in the post
1 comment
[ 3.6 ms ] story [ 11.9 ms ] thread