Home/Guides/Q4_K_M

GGUF quantization guide

Q4_K_M explained

Q4_K_M is a quantization label you will often see on GGUF local LLM files. It describes how model weights are represented. It does not, by itself, tell you the VRAM required, runtime speed, context capacity, or whether a model will fit your complete PC.

Quick answer

Q4_K_M is a quantization type, not a hardware requirement.

Quantization reduces the precision used to store model weights, which can make local model artifacts smaller. Current llama.cpp exposes Q4_K_M as a distinct quantization option and treats Q4_K as an alias for it. GGUF can store quantized tensors together with model metadata in a single model file.

What the label tells you

Separate quantization from runtime fit.

What Q4_K_M does tell you

  • It identifies a GGUF weight quantization preset.
  • It is part of the K-quant family used by llama.cpp.
  • It lets you distinguish one model artifact from other quantized variants of the same model.

What Q4_K_M does not tell you

  • Exact VRAM required at runtime.
  • Tokens per second on your GPU.
  • The maximum context your PC can use.
  • Whether full GPU offload is possible.

Related labels

Q4_K_M is one preset among several.

Q4_K_S

A separate Q4 K-quant preset. Its exact model artifact size depends on the model being quantized.

Q4_K_M

A Q4 K-quant preset commonly found in GGUF model repositories. Current llama.cpp also treats Q4_K as an alias for Q4_K_M.

Q5_K_M

A separate Q5 K-quant preset. The label describes a different quantization level, not a universal VRAM requirement.

These labels should not be turned into universal rankings. Different models produce different files, and practical runtime fit still depends on the complete configuration.

Recorded examples

Q4_K_M models currently tracked by AIPCFit

These are specific recorded GGUF artifacts in the AIPCFit LM Studio dataset. File size is shown as source data only. It is not presented as required VRAM.

Qwen3 4B · Q4_K_M

Q4_K_M
Recorded file
2.5 GB
Recorded context
32,768

lmstudio-community/Qwen3-4B-GGUF

View recorded source

Qwen3 8B · Q4_K_M

Q4_K_M
Recorded file
5.03 GB
Recorded context
32,768

lmstudio-community/Qwen3-8B-GGUF

View recorded source

Qwen3 14B · Q4_K_M

Q4_K_M
Recorded file
9 GB
Recorded context
32,768

lmstudio-community/Qwen3-14B-GGUF

View recorded source

Gemma 3 4B · Q4_K_M

Q4_K_M
Recorded file
2.49 GB
Recorded context
32,768

lmstudio-community/gemma-3-4b-it-GGUF

View recorded source

Gemma 3 12B · Q4_K_M

Q4_K_M
Recorded file
7.3 GB
Recorded context
32,768

lmstudio-community/gemma-3-12b-it-GGUF

View recorded source

Gemma 3 27B · Q4_K_M

Q4_K_M
Recorded file
16.5 GB
Recorded context
32,768

lmstudio-community/gemma-3-27b-it-GGUF

View recorded source

Memory

Why GGUF file size is not the same as VRAM use

Loading a local LLM involves more than its model file. Runtime memory can also be affected by context length, KV cache, GPU offload, implementation details, and other load settings. AIPCFit therefore keeps recorded GGUF file size separate from its own planning heuristics.

Common questions

What is Q4_K_M?

Q4_K_M is a K-quantization preset used for GGUF model weights in the llama.cpp ecosystem. It describes how model weights are quantized; it is not a statement that a model needs 4 GB of VRAM.

Does Q4_K_M mean 4-bit?

The Q4 label refers to a four-bit-class K-quantization scheme, but the complete representation also includes scales and related quantization data. The label should not be converted directly into a final model file size or runtime memory number.

Is Q4_K_M file size the same as VRAM required?

No. A recorded GGUF file size is useful information, but runtime memory can also depend on context length, KV cache, GPU offload, runtime implementation, and other load settings.

Is Q4_K_M always the best quantization?

No universal quantization is best for every model, device, or workload. The useful choice depends on the specific model artifact, available hardware, and the trade-offs you accept.

Can LM Studio run Q4_K_M models?

AIPCFit's LM Studio catalog includes recorded Q4_K_M GGUF variants. Full practical fit still depends on the complete PC, including GPU VRAM, system RAM, CPU support, context, and offload settings.

Evidence and trust

File size is data. VRAM fit is a separate question.

AIPCFit keeps recorded model metadata, planning heuristics, application compatibility, and runtime uncertainty separate. The model examples above link directly to the repositories recorded in the catalog.

How AIPCFit evaluates local AI fit