Making models smarter and cheaper at the same time 2 points by Wetime 1mo ago ↗ HN https://arxiv.org/abs/2607.14431
[–] Wetime 1mo ago ↗ Byte-exact KV grafting stores verified reasoning on disk. Replaying it lets a frozen 12B LLM beat 31B models at 8,700x less energy.
1 comment
[ 3.9 ms ] story [ 10.0 ms ] thread