Dedicated GPU memory guides
These Windows/NVIDIA guide pages use dedicated GPU memory as the primary fit budget and keep partial offload separate from a full GPU fit.
DeepSeek · DeepSeek
Review available GGUF variants, quantization choices, and conservative memory-planning levels for running DeepSeek-R1 Distill Qwen 7Blocally.
| Variant | Quantization | Estimated file | Minimum load | Planning memory |
|---|---|---|---|---|
| DeepSeek-R1 Distill Qwen 7B · Q4_K_M | Q4_K_M | 4.1 GB | 5.1 GB | 6.1 GB |
| DeepSeek-R1 Distill Qwen 7B · Q8_0 | Q8_0 | 7.4 GB | 8.5 GB | 10 GB |
GPU compatibility
These pages compare this model's recorded GGUF variants against the dedicated VRAM of selected NVIDIA GPUs. They are deterministic memory-fit checks, not performance benchmarks.
These Windows/NVIDIA guide pages use dedicated GPU memory as the primary fit budget and keep partial offload separate from a full GPU fit.
Apple Silicon pages use conservative unified-memory budgets and Metal-capable local inference terminology instead of treating the Mac as a CUDA workstation.
These are deterministic memory-planning checks, not benchmark claims. Context length, runtime settings, offload behavior, and concurrent apps can change practical fit.