17 comments

[ 200 ms ] story [ 1905 ms ] thread
[flagged]
But what good is 2B for? I didn’t see a use case tbh
Fine-tuning on a niche task I guess! E.g. a bunch of LoRAs that can be easily swapped depending on the task!
2B's hallucination will be glorious I feel
This is the problem. There isn't any use-cases for this other than being a toy. Even on the web.
Will Google add performant 7B support? 7B's quality is good enough for a serious Web APP.
(comment deleted)
It only says 'affor affor affor...(repeated 100+ times)' for me.

I am not sure how to debug this.

But is it saying "'affor affor affor...(repeated 100+ times)" really fast?

But actually, did you download the linked model from Kaggle ("gemma-2b-it-gpu-int4.bin") and upload it into the demo? It is working fine for me out of the box.

yes, I loaded the 1.35GB .bin file.
Are u using 'gpu' one? what's your platform. my mbp2019 seems good
Try a different prompt
nope, all my prompts lead the the same answer.
this can happen if you try to feed a cpu model to gpu inference. make sure you download the gpu model if you are using gpu accelerated ml inference.
This looks far faster than TVM's one?!