I will give it a look!
Coming back to this a few hours later, I've decided to add a section to explain this to the blog post. Thank you for flagging it.
I haven't, but that's just because it really isn't my personal usage pattern.
A literal manual typo. Good catch, will fix.
The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in…
Hehe I really am just working it out as I go along - I promise it is fairly painless. Hugging Face allow you to specify your machine and then browse models that fit. And then you can just vibe out 'oh this one is a bit…
It is not
Running it depends on RAM, which is what I wrote, bandwidth is important for speed. I chose my words carefully, but you are absolutely right.
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?
I'm the author - hello! Added to the post! Qwen averages 325 tok/s in processing prompts, and 34 tok/s in token generation. That isn't instant, but it's quick enough that I never really think about it.
I'm the author - hello! I talk about it in the blog post - knowing what's being run, knowing where it's being run, and not having anyone else control it.
I will give it a look!
Coming back to this a few hours later, I've decided to add a section to explain this to the blog post. Thank you for flagging it.
I haven't, but that's just because it really isn't my personal usage pattern.
A literal manual typo. Good catch, will fix.
The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in…
Hehe I really am just working it out as I go along - I promise it is fairly painless. Hugging Face allow you to specify your machine and then browse models that fit. And then you can just vibe out 'oh this one is a bit…
It is not
Running it depends on RAM, which is what I wrote, bandwidth is important for speed. I chose my words carefully, but you are absolutely right.
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?
I'm the author - hello! Added to the post! Qwen averages 325 tok/s in processing prompts, and 34 tok/s in token generation. That isn't instant, but it's quick enough that I never really think about it.
I'm the author - hello! I talk about it in the blog post - knowing what's being run, knowing where it's being run, and not having anyone else control it.