A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM (marcobambini.substack.com) 4 points by marcobambini 1mo ago ↗ HN
[–] armchairhacker 1mo ago ↗ Current speed is “approximately one third of a token per second” [–] marcobambini 1mo ago ↗ Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.
[–] marcobambini 1mo ago ↗ Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.
2 comments
[ 0.24 ms ] story [ 13.2 ms ] thread