Ask HN: Is it feasible to run a model on device for complete privacy?
Tried Gemma, Qwen and a few others. Need vision and larger context windows for an application I am working on. Results were quite poor Gemma 4E2B probably the best of the ones I tired but still fell apart and keep hallucinating with ~5000 tokens. Cloud based models had no problems even even Gemini 3.1 Flash-Lite and GPT-5.4 mini do a lot better and a way faster.
2 comments of 6
[ 7.1 ms ] story [ 54.9 ms ] thread