2 comments

[ 0.33 ms ] story [ 11.3 ms ] thread
Here you can see a 26B model hacked using a trained KV cache bank, responding like Gemma. The latency I am seeing is <127 ms.
That's awesome I came from the jev in python thread. is the model shareable? do you have any more writing on this? would love to read more about it! cool demo anyways.