Local LLM compatibility

What LLM can I run locally?

AIPCFit checks your PC hardware against local model data to help you understand which models are a practical fit for Ollama or LM Studio. Fit can depend on GPU VRAM, system RAM, model quantization, context length, offloading, and application compatibility.

The short answer

There is no single VRAM number that answers the question.

A smaller quantized model may fit comfortably, while a larger model may need CPU and system RAM offloading or may not be practical for the selected configuration. AIPCFit keeps verified compatibility, planning heuristics, and uncertainty separate instead of treating model file size as an official VRAM requirement.

If you are starting from a specific model and quantization instead of a PC profile, use the LLM VRAM calculator.

Check your hardware

See what fits your PC.

Set or reuse your PC profile, then check model fit separately for Ollama and LM Studio.

My PC

Set your hardware once

Your hardware profile is stored only in this browser and reused across compatibility tools.

Loading saved profile…
NPU detected in profile: AMD Ryzen AI 9 HX 370 · 50 TOPS

Your PC

Check an Ollama model

Choose your GPU, system memory, and model. The result combines verified Ollama compatibility with a conservative VRAM planning estimate.

Using My PC

NVIDIA GeForce RTX 5080 16GB

32 GB system RAM

Compatibility result

Good GPU fit

The selected GPU appears to have useful VRAM headroom for this model.

Ollama: supported

GPU VRAM

16 GB

Model file

5.2 GB

Planning VRAM

~6.7 GB

GPU accelerationcuda
Selected system RAM32 GB
WorkflowLocal chat
Context8K tokens
Context memory pressurelow

Context notes

  • This is a relatively modest context setting.
  • Longer contexts generally require more memory.

What this means

  • The GPU has enough VRAM for the model weights plus the conservative planning headroom used by this tool.
  • Actual VRAM usage still depends on context length, runtime settings, and concurrent GPU workloads.
  • This is an estimate and not a vendor guarantee.
Estimate only. Model file size is source data; planning VRAM includes extra headroom added by this tool and is not an official minimum requirement.

What affects local LLM fit?

Hardware capacity is only part of the answer.

GPU VRAM

More VRAM can allow more model data and runtime memory to stay on the GPU, but there is no single VRAM minimum that applies to every local LLM.

System RAM

System memory matters when model data or layers need to be handled outside GPU memory.

Quantization

Quantized model variants can reduce storage and memory pressure. File size alone is not the same thing as required VRAM.

Context length

Larger context windows can increase runtime memory use, so fit can change when the selected context changes.

Software

Ollama and LM Studio have their own platform and hardware compatibility paths, so AIPCFit checks them separately.

Popular hardware

Check popular GPUs for local LLMs.

Common questions

Can I run an LLM without a GPU?

Some local LLM runtimes can use the CPU and system RAM. Practical fit and performance vary by model, hardware, runtime, and configuration.

Is model file size the same as required VRAM?

No. Runtime memory, context length, offloading, and the implementation all matter. AIPCFit does not present model file size as an official VRAM requirement.

Does more VRAM always mean a faster LLM?

No. More VRAM can make larger configurations possible, but speed also depends on the GPU, CPU, memory path, runtime, and model configuration.

Should I use Ollama or LM Studio?

Both can be useful for local LLM workflows. The better choice depends on your preferred workflow and the compatibility path for your system, so AIPCFit checks them separately.

Evidence and trust

Unknown stays unknown.

AIPCFit separates verified application compatibility from model fit heuristics and does not invent benchmark performance when reliable data is unavailable.