Model memory estimator
A transparent first estimate. No mystery hardware score.
Enter the total stored parameter count and nominal bits per weight. The result estimates weight storage in binary gibibytes. It is not the total RAM or VRAM needed to run a model.
Before quantization metadata, context/KV cache, runtime buffers and other applications.
The calculation
Total parameters × bits per weight ÷ 8 gives bytes. Dividing that result by 1,073,741,824 converts bytes to GiB. For example, 8 billion parameters at a nominal 4 bits requires approximately 3.73 GiB for the weights under this simplified calculation.
Why actual memory differs
Quantization formats store scales and other metadata, and may mix precisions across tensors. Inference also needs memory for the runtime, intermediate values and cached context. Larger context windows and concurrent requests can increase that additional footprint. CPU/GPU splitting and offloading change where memory is used.
Use the actual download size and runtime documentation for a more concrete baseline, then measure your intended workload. This tool does not predict tokens per second, verify model quality or recommend a GPU purchase.