[–] karmasimida 2y ago ↗ But what good is 2B for? I didn’t see a use case tbh [–] reqo 2y ago ↗ Fine-tuning on a niche task I guess! E.g. a bunch of LoRAs that can be easily swapped depending on the task! [–] karmasimida 2y ago ↗ 2B's hallucination will be glorious I feel [–] rvz 2y ago ↗ This is the problem. There isn't any use-cases for this other than being a toy. Even on the web. [–] lccccc 2y ago ↗ Will Google add performant 7B support? 7B's quality is good enough for a serious Web APP.
[–] reqo 2y ago ↗ Fine-tuning on a niche task I guess! E.g. a bunch of LoRAs that can be easily swapped depending on the task! [–] karmasimida 2y ago ↗ 2B's hallucination will be glorious I feel
[–] rvz 2y ago ↗ This is the problem. There isn't any use-cases for this other than being a toy. Even on the web.
[–] lccccc 2y ago ↗ Will Google add performant 7B support? 7B's quality is good enough for a serious Web APP.
[–] qrian 2y ago ↗ It only says 'affor affor affor...(repeated 100+ times)' for me.I am not sure how to debug this. [–] mjks 2y ago ↗ But is it saying "'affor affor affor...(repeated 100+ times)" really fast?But actually, did you download the linked model from Kaggle ("gemma-2b-it-gpu-int4.bin") and upload it into the demo? It is working fine for me out of the box. [–] qrian 2y ago ↗ yes, I loaded the 1.35GB .bin file. [–] hackerrrrx 2y ago ↗ Are u using 'gpu' one? what's your platform. my mbp2019 seems good [–] magemgem 2y ago ↗ Try a different prompt [–] qrian 2y ago ↗ nope, all my prompts lead the the same answer. [–] miohtama 2y ago ↗ It thinks it’s a pokemon. [–] impjdi 2y ago ↗ this can happen if you try to feed a cpu model to gpu inference. make sure you download the gpu model if you are using gpu accelerated ml inference.
[–] mjks 2y ago ↗ But is it saying "'affor affor affor...(repeated 100+ times)" really fast?But actually, did you download the linked model from Kaggle ("gemma-2b-it-gpu-int4.bin") and upload it into the demo? It is working fine for me out of the box. [–] qrian 2y ago ↗ yes, I loaded the 1.35GB .bin file. [–] hackerrrrx 2y ago ↗ Are u using 'gpu' one? what's your platform. my mbp2019 seems good
[–] qrian 2y ago ↗ yes, I loaded the 1.35GB .bin file. [–] hackerrrrx 2y ago ↗ Are u using 'gpu' one? what's your platform. my mbp2019 seems good
[–] magemgem 2y ago ↗ Try a different prompt [–] qrian 2y ago ↗ nope, all my prompts lead the the same answer.
[–] impjdi 2y ago ↗ this can happen if you try to feed a cpu model to gpu inference. make sure you download the gpu model if you are using gpu accelerated ml inference.
17 comments
[ 200 ms ] story [ 1905 ms ] threadI am not sure how to debug this.
But actually, did you download the linked model from Kaggle ("gemma-2b-it-gpu-int4.bin") and upload it into the demo? It is working fine for me out of the box.