1 comment

[ 4.0 ms ] story [ 16.5 ms ] thread
This calculator estimates the GPU memory needed to run LLM inference. Select the model size and precision (FP32 - FP4) to get a quick memory range estimate.