4 comments

[ 2.8 ms ] story [ 19.8 ms ] thread
This is pretty sick. 7tokens/second is not fast but still seems useful.
+1 for pairing Samosa with Chāt :D :D
Very interesting. I am always curious on llama.cpp and vllm. All the best for samosa inference engine. One reason I am not running local models on my Mac now a days, is My Mac book pro is getting heated more. Otherwise would love to run couple (at least 1 to 3) of local models continuously.
any way to run samosa in opencode?