Small
Phi-3.5 Mini
- Provider
- Microsoft
- Parameters
- 3.8B
- Quantizations
- Q4_K_M
- Use case
- Coding
Local coding LLM hardware fit
These recommendations start with models already tagged for coding in the AIPCFit catalog, then use practical memory fit, GGUF quantization, VRAM tiers, and hardware compatibility signals. The best local LLM for coding here means the best practical local fit for programming work on your hardware, not a coding benchmark leaderboard.
Coding-tagged only
Every local coding LLM on this page comes from the existing catalog and includes the coding use-case tag.
Fit before preference
VRAM, quantization, model size, load memory, and runtime overhead determine whether a local LLM for programming is a practical fit.
Speed is separate
A model can fit in memory without being the right speed or quality match for every codebase, runtime, or coding task.
Coding models
Models are grouped by the catalog size class, not by claimed code intelligence. Use this section to compare local coding models by provider, parameter count, quantization options, and available VRAM fit links.
Smaller coding-tagged models that can be practical on lower-memory local systems.
Small
Small
Small
Small
Small
Small
Coding-tagged models in the catalog's medium size class.
Medium
Medium
Medium
Higher-parameter coding-tagged models where fit depends more heavily on VRAM, quantization, and offload behavior.
Larger
Larger
Larger
Larger
Very Large
Coding by VRAM
Each tier filters the existing VRAM fit guide to coding-tagged models, keeps acceptable fit statuses from the existing rules, and deduplicates by model ID. Use these as hardware-fit starting points for searches like coding LLM for 8GB VRAM, coding LLM for 16GB VRAM, or the best coding LLM locally on a larger GPU.
Hardware tier
Hardware tier
Hardware tier
Hardware tier
Hardware tier
Selection guide
Start with model size and quantization. Smaller or lower-bit GGUF variants usually have lower memory requirements, while larger models need more VRAM, system RAM, or offloading headroom.
Check context length and runtime behavior for your coding workflow. Long files, multi-file prompts, retrieval workflows, and IDE integrations can change memory use beyond the model file size.
Fit and speed are separate questions. The fit engine can show whether a local llm for programming is practical for a memory tier, but coding quality and latency still vary by task, prompt, runtime, drivers, and software settings.
Setup paths
AIPCFit has GPU-specific runtime pages rather than a single universal runtime score. Start with your GPU, then compare local model fit for Ollama, LM Studio, and compatibility pages.
Methodology
These are hardware-fit recommendations for local coding LLMs. They are not a coding benchmark leaderboard, model-quality ranking, or universal speed ranking.
Coding quality varies by programming language, repository size, prompt, runtime, context settings, and tool integration. A practical memory fit is the starting point, not the final answer for every developer workflow.