Hardware fit
Memory planning
Q4_K_M · 3.7GB recommended
2.7GB minimum load estimate
Hardware-fit model ranking
Compare local LLMs by practical hardware fit instead of generic benchmark hype. AIPCFit focuses on what can realistically run on your hardware using stored model metadata, quantization data, memory-planning estimates, VRAM tiers, Apple Silicon memory tiers, and use-case tags.
Hardware-fit rankings use AIPCFit's deterministic compatibility and memory-planning data. Model quality, runtime speed, and benchmark performance are separate questions.
Interactive hardware-fit leaderboard
Rows use AIPCFit's stored model metadata and memory-planning fit logic. They do not use invented quality scores or runtime speed claims.
20 matching models
Hardware filter
Use case filter
Hardware fit
Memory planning
Q4_K_M · 3.7GB recommended
2.7GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 4.2GB recommended
3.2GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 4.2GB recommended
3.2GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 4.5GB recommended
3.5GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 4.5GB recommended
3.5GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 6.1GB recommended
5.1GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 6.1GB recommended
5.1GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 6.6GB recommended
5.6GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 7GB recommended
6GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 9.9GB recommended
8.4GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 10.9GB recommended
9.3GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 10.9GB recommended
9.3GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 12.2GB recommended
10.4GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 20.3GB recommended
17.3GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 22.3GB recommended
19GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 24.3GB recommended
20.7GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 25.7GB recommended
21.9GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 27GB recommended
23GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 36.6GB recommended
31.2GB minimum load estimate
Hardware fit
Memory planning
Q4_K_M · 54.8GB recommended
46.7GB minimum load estimate
Ordering logic
The default leaderboard puts stronger fit statuses first, then earlier practical memory tiers, lower planning memory, model size, and model name as a stable fallback. It does not rank universal intelligence or claim a single model is the right choice for every local setup.
8GB VRAM tier
These models have practical fit signals in the stored 8GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.
Llama
3B parameters · Q4_K_M · 3.7GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Gemma
4B parameters · Q4_K_M · 4.5GB recommended planning memory
12GB VRAM tier
These models have practical fit signals in the stored 12GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.
Llama
3B parameters · Q4_K_M · 3.7GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Gemma
4B parameters · Q4_K_M · 4.5GB recommended planning memory
16GB VRAM tier
These models have practical fit signals in the stored 16GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.
Llama
3B parameters · Q4_K_M · 3.7GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Gemma
4B parameters · Q4_K_M · 4.5GB recommended planning memory
24GB VRAM tier
These models have practical fit signals in the stored 24GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.
Llama
3B parameters · Q4_K_M · 3.7GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Gemma
4B parameters · Q4_K_M · 4.5GB recommended planning memory
Unified memory systems
Apple Silicon uses unified memory, so CPU and GPU share the same pool. Available unified memory, model size, quantization, context, and runtime overhead all affect practical local LLM fit.
Llama
3B parameters · Q4_K_M · 3.7GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Phi
3.8B parameters · Q4_K_M · 4.2GB recommended planning memory
Gemma
4B parameters · Q4_K_M · 4.5GB recommended planning memory
Qwen
4B parameters · Q4_K_M · 4.5GB recommended planning memory
DeepSeek
7B parameters · Q4_K_M · 6.1GB recommended planning memory
Coding models
Coding fit depends on model size, quantization, repository context, local development tooling, and available memory. This leaderboard can filter coding-tagged models, but the dedicated coding page goes deeper on hardware tiers and programming use cases.
Open the local coding LLM guideMethodology
This is a hardware-fit leaderboard. It is not a universal intelligence leaderboard, model-quality ranking, or runtime speed ranking.
Fit depends on GPU VRAM or unified memory, system RAM, quantization, model size, context length, runtime overhead, offload behavior, and the software used to run the model.
Benchmark quality and runtime speed are useful separate concepts, but AIPCFit keeps them separate from memory fit unless the project has structured, sourced data for them.
Read the AIPCFit methodologyPC-specific recommendation
Use the leaderboard when you want a practical model shortlist by tier. Use the PC checker when you want recommendations from your exact GPU, RAM, platform, model size, quantization, and runtime.
Common questions
A local LLM leaderboard compares models that can run on your own hardware. AIPCFit's version ranks practical hardware fit using model metadata, memory-planning estimates, VRAM tiers, Apple Silicon memory tiers, and use-case tags.
8GB VRAM is usually a careful planning tier for smaller quantized models and conservative context settings. Use the 8GB filter to see models with the strongest stored fit signals for that tier.
16GB VRAM gives more flexibility for mid-size models and larger quantizations, but model size, quantization, context length, and runtime overhead still decide practical fit.
More VRAM can allow larger configurations and more context headroom, but it does not guarantee better quality or faster runtime by itself. The model, quantization, runtime, CPU, RAM, and software path also matter.
No. AIPCFit does not invent model-quality ratings or runtime speed numbers. This page is a hardware-fit leaderboard, not a universal intelligence or speed ranking.
Rows are ordered by the strongest practical fit status first, then by lower required memory tier, planning memory, model size, and model name as a stable fallback.
Yes, Apple Silicon can run local LLMs through supported local runtimes. CPU and GPU share unified memory, so available unified memory, model size, quantization, and context length all affect fit.
Use the AIPCFit PC checker to start from your exact GPU, RAM, operating system, and local AI workload instead of relying only on a generic leaderboard tier.