> I know multiple startups that use LLMs as their core bread-and-butter intelligence platform instead of tuned but traditional NLP models It seems like LLMs would be perfect for start-ups that are iterating quickly. As…
GPU RAM quantity isn’t typically correlated to inference rate. Precision/quantization levels do affect model size, which will affect inference rate. However, I would expect a smaller model to be faster (less RAM).
If you think this is wild, see the PaLM 2 paper with 2.5 pages of 2 column attributions. https://arxiv.org/pdf/2305.10403.pdf
> I know multiple startups that use LLMs as their core bread-and-butter intelligence platform instead of tuned but traditional NLP models It seems like LLMs would be perfect for start-ups that are iterating quickly. As…
GPU RAM quantity isn’t typically correlated to inference rate. Precision/quantization levels do affect model size, which will affect inference rate. However, I would expect a smaller model to be faster (less RAM).
If you think this is wild, see the PaLM 2 paper with 2.5 pages of 2 column attributions. https://arxiv.org/pdf/2305.10403.pdf