Home/Local LLM leaderboard

Hardware-fit model ranking

Local LLM Leaderboard

Compare local LLMs by practical hardware fit instead of generic benchmark hype. AIPCFit focuses on what can realistically run on your hardware using stored model metadata, quantization data, memory-planning estimates, VRAM tiers, Apple Silicon memory tiers, and use-case tags.

Hardware-fit rankings use AIPCFit's deterministic compatibility and memory-planning data. Model quality, runtime speed, and benchmark performance are separate questions.

Interactive hardware-fit leaderboard

Filter local LLMs by hardware and use case

Rows use AIPCFit's stored model metadata and memory-planning fit logic. They do not use invented quality scores or runtime speed claims.

20 matching models

Hardware filter

Use case filter

#1 · Llama

Llama 3.2 3B

Meta · 3B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 3.7GB recommended

2.7GB minimum load estimate

chatlightweightRAGView model

#2 · Phi

Phi-3.5 Mini

Microsoft · 3.8B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 4.2GB recommended

3.2GB minimum load estimate

chatlightweightcodingView model

#3 · Phi

Phi-4 Mini

Microsoft · 3.8B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 4.2GB recommended

3.2GB minimum load estimate

chatlightweightcodingView model

#4 · Gemma

Gemma 3 4B

Google · 4B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 4.5GB recommended

3.5GB minimum load estimate

chatlightweightRAGView model

#5 · Qwen

Qwen3 4B

Alibaba Cloud · 4B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 4.5GB recommended

3.5GB minimum load estimate

chatlightweightRAGView model

#7 · Mistral

Mistral 7B

Mistral AI · 7B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 6.1GB recommended

5.1GB minimum load estimate

chatcodinglightweightView model

#8 · Llama

Llama 3.1 8B

Meta · 8B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 6.6GB recommended

5.6GB minimum load estimate

chatcodingRAGView model

#9 · Qwen

Qwen3 8B

Alibaba Cloud · 8B parameters

Hardware fit

Good fit8GB VRAM

Memory planning

Q4_K_M · 7GB recommended

6GB minimum load estimate

chatcodingRAGView model

#10 · Gemma

Gemma 3 12B

Google · 12B parameters

Hardware fit

Memory planning

Q4_K_M · 9.9GB recommended

8.4GB minimum load estimate

chatreasoningRAGView model

#12 · Phi

Phi-4

Microsoft · 14B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 10.9GB recommended

9.3GB minimum load estimate

chatreasoningcodingView model

#13 · Qwen

Qwen3 14B

Alibaba Cloud · 14B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 12.2GB recommended

10.4GB minimum load estimate

chatcodingreasoningView model

#14 · Mistral

Mistral Small 3.2 24B

Mistral AI · 24B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 20.3GB recommended

17.3GB minimum load estimate

chatcodingreasoningRAGView model

#15 · Gemma

Gemma 3 27B

Google · 27B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 22.3GB recommended

19GB minimum load estimate

chatreasoningView model

#16 · Qwen

Qwen3.8 27B

Alibaba Cloud · 27B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 24.3GB recommended

20.7GB minimum load estimate

chatcodingreasoningRAGView model

#17 · Qwen

Qwen3 30B-A3B

Alibaba Cloud · 30.5B parameters

Hardware fit

Partial offload8GB VRAM

Memory planning

Q4_K_M · 25.7GB recommended

21.9GB minimum load estimate

chatcodingreasoningRAGView model

#19 · Mistral

Mixtral 8x7B

Mistral AI · 46.7B parameters

Hardware fit

Partial offload24GB VRAM

Memory planning

Q4_K_M · 36.6GB recommended

31.2GB minimum load estimate

chatcodingreasoningView model

#20 · Llama

Llama 3.3 70B

Meta · 70B parameters

Hardware fit

Partial offload32GB VRAM

Memory planning

Q4_K_M · 54.8GB recommended

46.7GB minimum load estimate

chatreasoningView model

Ordering logic

Default order favors practical fit, not hype.

The default leaderboard puts stronger fit statuses first, then earlier practical memory tiers, lower planning memory, model size, and model name as a stable fallback. It does not rank universal intelligence or claim a single model is the right choice for every local setup.

8GB VRAM tier

Best Local LLMs for 8GB VRAM

These models have practical fit signals in the stored 8GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.

Good fit

3B parameters · Q4_K_M · 3.7GB recommended planning memory

chatlightweightRAG
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG

12GB VRAM tier

Best Local LLMs for 12GB VRAM

These models have practical fit signals in the stored 12GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.

Good fit

3B parameters · Q4_K_M · 3.7GB recommended planning memory

chatlightweightRAG
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG

16GB VRAM tier

Best Local LLMs for 16GB VRAM

These models have practical fit signals in the stored 16GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.

Good fit

3B parameters · Q4_K_M · 3.7GB recommended planning memory

chatlightweightRAG
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG

24GB VRAM tier

Best Local LLMs for 24GB VRAM

These models have practical fit signals in the stored 24GB VRAM guide data. Quantization, context length, runtime overhead, system RAM, and software settings can still change real-world behavior.

Good fit

3B parameters · Q4_K_M · 3.7GB recommended planning memory

chatlightweightRAG
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG

Unified memory systems

Best Local LLMs for Apple Silicon

Apple Silicon uses unified memory, so CPU and GPU share the same pool. Available unified memory, model size, quantization, context, and runtime overhead all affect practical local LLM fit.

Good fit

3B parameters · Q4_K_M · 3.7GB recommended planning memory

chatlightweightRAG
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

3.8B parameters · Q4_K_M · 4.2GB recommended planning memory

chatlightweightcoding
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG
Good fit

4B parameters · Q4_K_M · 4.5GB recommended planning memory

chatlightweightRAG

Coding models

Best Local LLMs for Coding

Coding fit depends on model size, quantization, repository context, local development tooling, and available memory. This leaderboard can filter coding-tagged models, but the dedicated coding page goes deeper on hardware tiers and programming use cases.

Open the local coding LLM guide

Methodology

How AIPCFit Ranks Local LLMs

This is a hardware-fit leaderboard. It is not a universal intelligence leaderboard, model-quality ranking, or runtime speed ranking.

Fit depends on GPU VRAM or unified memory, system RAM, quantization, model size, context length, runtime overhead, offload behavior, and the software used to run the model.

Benchmark quality and runtime speed are useful separate concepts, but AIPCFit keeps them separate from memory fit unless the project has structured, sourced data for them.

Read the AIPCFit methodology

PC-specific recommendation

Find the Best Local LLM for Your PC

Use the leaderboard when you want a practical model shortlist by tier. Use the PC checker when you want recommendations from your exact GPU, RAM, platform, model size, quantization, and runtime.

Common questions

Local LLM Leaderboard FAQ

What is a local LLM leaderboard?

A local LLM leaderboard compares models that can run on your own hardware. AIPCFit's version ranks practical hardware fit using model metadata, memory-planning estimates, VRAM tiers, Apple Silicon memory tiers, and use-case tags.

Which local LLM is best for 8GB VRAM?

8GB VRAM is usually a careful planning tier for smaller quantized models and conservative context settings. Use the 8GB filter to see models with the strongest stored fit signals for that tier.

Which local LLM is best for 16GB VRAM?

16GB VRAM gives more flexibility for mid-size models and larger quantizations, but model size, quantization, context length, and runtime overhead still decide practical fit.

Is more VRAM always better for local LLMs?

More VRAM can allow larger configurations and more context headroom, but it does not guarantee better quality or faster runtime by itself. The model, quantization, runtime, CPU, RAM, and software path also matter.

Are these model benchmark results?

No. AIPCFit does not invent model-quality ratings or runtime speed numbers. This page is a hardware-fit leaderboard, not a universal intelligence or speed ranking.

How does AIPCFit rank local LLMs?

Rows are ordered by the strongest practical fit status first, then by lower required memory tier, planning memory, model size, and model name as a stable fallback.

Can Apple Silicon run local LLMs?

Yes, Apple Silicon can run local LLMs through supported local runtimes. CPU and GPU share unified memory, so available unified memory, model size, quantization, and context length all affect fit.

How do I know which local LLM my PC can run?

Use the AIPCFit PC checker to start from your exact GPU, RAM, operating system, and local AI workload instead of relying only on a generic leaderboard tier.