Making models smarter and cheaper at the same time

2 points by Wetime ↗ HN
https://arxiv.org/abs/2607.14431

1 comment

[ 3.9 ms ] story [ 10.0 ms ] thread
Byte-exact KV grafting stores verified reasoning on disk. Replaying it lets a frozen 12B LLM beat 31B models at 8,700x less energy.