Show HN: Selfhostllm.org – Plan GPU capacity for self-hosting LLMs (selfhostllm.org) 7 points by erans 1y ago ↗ HN A simple calculator that estimates how many concurrent requests your GPU can handle for a given LLM, with shareable results.
[–] erans 1y ago ↗ I also added a Mac version: https://selfhostllm.org/mac/ so you can know which models you can run on your Mac and get an estimated tokens/sec.
[–] harshnigam 1y ago ↗ I see it doesn't take GPU performance into consideration when showing the estimates. H100 and A100 are performing the same. Am I doing it wrong?
3 comments
[ 3.1 ms ] story [ 20.7 ms ] thread