Ask HN: What's the best LLM model that on a 24 GB VRAM GPU? 3 points by max93 3mo ago ↗ HN What’s the best model right now that outperforms Qwopus3.6-27B-v2-MTP-GGUF 8-bit on a 24 GB VRAM GPU? Looking for real reviews. I found 4 bit not usable in production.
[–] jr_isidore 3mo ago ↗ Good question. I was told in 2024 to get an RTX 3090 (24 GB VRAM), so I did, and nothing on HuggingFace was usable.
[–] sds357 3mo ago ↗ qwen3.6 27b and 35b work for me. I run them on proxmox with a an old P40 gpu, and ollama
3 comments
[ 4.0 ms ] story [ 28.4 ms ] thread