Local AI hardware guide
How Much VRAM Do You Need to Run an LLM Locally?
There is no single VRAM number that works for every model. Use model size as a first estimate, then account for context, runtime overhead, and the rest of your system.
If you are choosing a graphics card for local AI, VRAM matters. But a model's parameter count or download size does not tell you the complete amount of memory it will use while running.
The right amount depends on the model, its file format and quantization, the context length, the software, and whether some work is handled by system RAM instead. That is why "this model needs exactly X GB" can be misleading without those details.
What VRAM does
VRAM is the dedicated memory on a graphics card. During inference, the system may place some or all model weights in VRAM so the GPU can process them. The application also needs memory for other runtime data, including the context used to handle a prompt and conversation.
If the model does not fit completely on the GPU, some software can use system RAM or split work between the CPU and GPU. That may let the model run, but the speed can be lower and the exact behavior varies by runtime and hardware.
Use model size as a rough starting point
A simple lower-bound estimate for weight storage is:
Approximate weight memory in bytes = parameter count x bits per parameter / 8
For example, a 7-billion-parameter model represented at 4 bits per parameter has a theoretical weight payload of about 3.5 billion bytes. This is only a rough estimate: real files have metadata and additional structures, and runtime memory is needed beyond the weights.
The model's published download size can be a more practical first clue than parameter count alone, but it still does not represent the full memory needed while generating responses.
Why context length changes the answer
Context is the text the model can consider while responding. A longer context can be useful for large documents, coding sessions, or long conversations, but it requires additional memory. A model that loads with a short context may fail or slow down when the context is increased.
Ollama's documentation specifically notes that increasing context length increases memory requirements. The application may also choose a default context based partly on available VRAM, so check the actual setting instead of assuming every install uses the same amount.
Quantization can help, with tradeoffs
Quantization stores model weights using fewer bits. A lower-bit version is often smaller and may fit in less memory, which can make local inference possible on more computers. Different quantizations can vary in response quality and speed, so compare versions for the task you care about.
How to estimate what will fit
- Find the exact model file and its download size.
- Check the model's recommended memory and supported runtimes.
- Account for context length and memory used by your operating system and apps.
- Leave headroom instead of targeting 100% of your available VRAM.
- Test the model and inspect whether it is running on the GPU, CPU, or both.
Do not treat a model as a good fit just because its file size is slightly smaller than your VRAM capacity. The runtime needs memory too. If a model is close to the limit, try a smaller quantization, shorter context, or a smaller model.
Need help matching hardware to a workload? Check your PC with AIPCFit and compare the result with the model's own requirements.
The short answer
There is no universal VRAM requirement for local LLMs. Start with the model's weight size, then account for context, runtime overhead, and available system memory. A realistic test on your own setup is the best way to confirm both fit and speed.