1 comment

[ 5.8 ms ] story [ 9.7 ms ] thread
A local model generating 20 tokens/sec today could potentially reach 40 tokens/sec in many scenarios.