Home/32GB VRAM

Dedicated GPU memory guide

Best local AI models for 32GB VRAM

This guide uses a Windows NVIDIA profile with 32GB dedicated GPU VRAM. It separates full GPU fits, tight fits, and partial offload instead of treating model file size as an official memory requirement.

Dedicated VRAM is not the same as system RAM or Apple Silicon unified memory. For Mac planning, use the Apple Silicon guides so the fit engine uses unified-memory terminology and a Mac-specific conservative budget.

Likely full GPU fit

Good fit

Llama 3.2 3B

Q4_K_M

Llama 3.2 3B · Q4_K_M · 3.7 GB planning memory · 2.7 GB minimum load

Good fit

Phi-3.5 Mini

Q4_K_M

Phi-3.5 Mini · Q4_K_M · 4.2 GB planning memory · 3.2 GB minimum load

Good fit

Phi-4 Mini

Q4_K_M

Phi-4 Mini · Q4_K_M · 4.2 GB planning memory · 3.2 GB minimum load

Good fit

Gemma 3 4B

Q4_K_M

Gemma 3 4B · Q4_K_M · 4.5 GB planning memory · 3.5 GB minimum load

Good fit

Qwen3 4B

Q4_K_M

Qwen3 4B · Q4_K_M · 4.5 GB planning memory · 3.5 GB minimum load

Good fit

Llama 3.2 3B

Q8_0

Llama 3.2 3B · Q8_0 · 5.2 GB planning memory · 4.2 GB minimum load

Good fit

DeepSeek-R1 Distill Qwen 7B

Q4_K_M

DeepSeek-R1 Distill Qwen 7B · Q4_K_M · 6.1 GB planning memory · 5.1 GB minimum load

Good fit

Mistral 7B

Q4_K_M

Mistral 7B · Q4_K_M · 6.1 GB planning memory · 5.1 GB minimum load

Good fit

Gemma 3 4B

Q8_0

Gemma 3 4B · Q8_0 · 6.2 GB planning memory · 5.2 GB minimum load

Good fit

Qwen3 4B

Q8_0

Qwen3 4B · Q8_0 · 6.2 GB planning memory · 5.2 GB minimum load

Good fit

Llama 3.1 8B

Q4_K_M

Llama 3.1 8B · Q4_K_M · 6.6 GB planning memory · 5.6 GB minimum load

Good fit

Qwen3 8B

Q4_K_M

Qwen3 8B · Q4_K_M · 7 GB planning memory · 6 GB minimum load

Good fit

Llama 3.1 8B

Q5_K_M

Llama 3.1 8B · Q5_K_M · 7.8 GB planning memory · 6.8 GB minimum load

Good fit

Qwen3 8B

Q5_K_M

Qwen3 8B · Q5_K_M · 7.8 GB planning memory · 6.8 GB minimum load

Good fit

Gemma 3 12B

Q4_K_M

Gemma 3 12B · Q4_K_M · 9.9 GB planning memory · 8.4 GB minimum load

Good fit

DeepSeek-R1 Distill Qwen 7B

Q8_0

DeepSeek-R1 Distill Qwen 7B · Q8_0 · 10 GB planning memory · 8.5 GB minimum load

Good fit

Mistral 7B

Q8_0

Mistral 7B · Q8_0 · 10 GB planning memory · 8.5 GB minimum load

Good fit

DeepSeek-R1 Distill Qwen 14B

Q4_K_M

DeepSeek-R1 Distill Qwen 14B · Q4_K_M · 10.9 GB planning memory · 9.3 GB minimum load

Good fit

Phi-4

Q4_K_M

Phi-4 · Q4_K_M · 10.9 GB planning memory · 9.3 GB minimum load

Good fit

Gemma 3 12B

Q5_K_M

Gemma 3 12B · Q5_K_M · 11.6 GB planning memory · 9.9 GB minimum load

Good fit

Qwen3 14B

Q4_K_M

Qwen3 14B · Q4_K_M · 12.2 GB planning memory · 10.4 GB minimum load

Good fit

Qwen3 14B

Q5_K_M

Qwen3 14B · Q5_K_M · 13.6 GB planning memory · 11.6 GB minimum load

Good fit

Mistral Small 3.2 24B

Q4_K_M

Mistral Small 3.2 24B · Q4_K_M · 20.3 GB planning memory · 17.3 GB minimum load

Good fit

Gemma 3 27B

Q4_K_M

Gemma 3 27B · Q4_K_M · 22.3 GB planning memory · 19 GB minimum load

Good fit

Qwen3.8 27B

Q4_K_M

Qwen3.8 27B · Q4_K_M · 24.3 GB planning memory · 20.7 GB minimum load

Good fit

Qwen3 30B-A3B

Q4_K_M

Qwen3 30B-A3B · Q4_K_M · 25.7 GB planning memory · 21.9 GB minimum load

Good fit

Gemma 3 27B

Q5_K_M

Gemma 3 27B · Q5_K_M · 26.2 GB planning memory · 22.3 GB minimum load

Good fit

DeepSeek-R1 Distill Qwen 32B

Q4_K_M

DeepSeek-R1 Distill Qwen 32B · Q4_K_M · 27 GB planning memory · 23 GB minimum load

Fits

Mistral Small 3.2 24B

Q8_0

Mistral Small 3.2 24B · Q8_0 · 35.1 GB planning memory · 29.9 GB minimum load

Fits

Mixtral 8x7B

Q4_K_M

Mixtral 8x7B · Q4_K_M · 36.6 GB planning memory · 31.2 GB minimum load

Tight fit

Partial offload

Compare nearby tiers