· 7 min read
How to Pick Hardware for Running Coding Models Locally
By J. Jung
- tools
Pick local-model hardware from the model, context window, and number of simultaneous agent sessions you actually intend to run; then prioritize memory capacity over advertised AI TOPS. For a developer who wants a coding agent to work through a real repository at a 64K context length, 24GB of VRAM is a tight fit, 32GB is more comfortable, and a Mac needs enough unified memory that the model is not competing with your editor, browser, containers, and the OS.
Don’t start with “What GPU is fastest?” Start with the command you want to leave running while you work. Ollama’s current coding integration recommends at least 64,000 tokens of context, and its example for glm-4.7-flash says that configuration needs roughly 23GB of VRAM. That is the useful unit of planning: a specific model at a specific context length, not its parameter count in isolation.
Measure the workload before you price a machine
A model’s quantized weight file is only the starting point. Running it also consumes memory for the KV cache, compute buffers, and runtime overhead. The KV cache is why a model that fits at 8K context can suddenly stop fitting after you tell an agent to inspect a large diff, read tests, and keep the last ten tool calls in view.
Use the model and context you plan to use for coding, not a small chat prompt, as the test. If you use llama.cpp directly, its llama-fit-params utility will inspect free memory and print launch arguments that fit the model. It also prints a breakdown for model, context, compute, and unaccounted GPU memory. That beats guessing from a Reddit spreadsheet.
# Run this after downloading the exact GGUF you want to use
./build/bin/llama-fit-params \
--model /opt/models/your-coding-model.gguf | tee args.txt
# Start the server with the fitted settings
cat args.txt | xargs ./build/bin/llama-server \
--model /opt/models/your-coding-model.ggufFor Ollama, pull the candidate model, set the intended context length, run the kind of task you will actually delegate, and inspect whether it stays on the GPU. A useful first experiment is the documented coding-model path below. Do it on the machine you already own before deciding that a new box is necessary.
ollama pull glm-4.7-flash
ollama launch claude
ollama psThe last command matters. If the model spills substantially into system RAM, it may still work, but interactive tool use becomes the part you resent: waiting after every shell command, file read, and test result. It is fine for an overnight task; it is a poor experience for a tight edit-run-review loop.
Buy the memory tier that matches the job
- 16GB VRAM: good for trying smaller quantized coding models, autocomplete-like help, short-context chat, and one focused task. It is not where I would plan a full 64K coding-agent workflow. You will spend time reducing context or accepting partial GPU offload.
- 24GB VRAM plus 64GB of system RAM: the practical floor for a dedicated local coding box. It can be enough for a model whose measured requirement is around 20–23GB, but it leaves little space for longer context, another model, or a second agent session.
- 32GB VRAM: the cleaner target if local coding is a daily tool rather than an experiment. The extra capacity is not just about larger weights; it buys context headroom and reduces the need to tune every run around an out-of-memory edge.
- 64–128GB of unified memory on Apple silicon: sensible when you want one compact machine for development and local inference, or when a single consumer GPU cannot hold the model you want. Treat unified memory as shared memory, though: Docker, browsers, IDE indexing, displays, and the model all draw from the same pool.
- CPU-only with plenty of RAM: acceptable for evaluation, embeddings, batch jobs, or a remote service where latency is not your problem. Don’t buy it expecting a responsive coding agent that continuously reads files and runs tools while you wait.
The uncomfortable answer is that a 16GB card can be technically capable and still be the wrong purchase. If your intended workflow is “ask a model about one function,” it is plenty. If the workflow is “give an agent a ticket and let it retain architecture notes, test output, and a wide diff,” the context cache changes the purchase decision.
Pick a platform for the software you will run
For a Linux or Windows build intended primarily for local inference, NVIDIA remains the straightforward path when your preferred runner expects CUDA. Capacity is still the gating number. NVIDIA’s GeForce RTX 5090 has 32GB of GDDR7, but its reference specification is also 575W total graphics power with a 1000W recommended system power supply. Check the case, PSU, connector clearance, heat, and noise before treating a 32GB card as a drop-in upgrade.
AMD can be a viable capacity-per-dollar route if your runner supports it, but verify the exact stack before ordering hardware. Ollama now enables Vulkan by default for broader AMD and Intel GPU acceleration, which removes some setup friction, but it does not make every model format, quantization, operating system, and agent wrapper equally tested. Install the runner you plan to use on your existing hardware first; don’t make an expensive driver experiment your migration plan.
Apple silicon is compelling when unified memory is the point, not when you expect modular upgrades. Apple’s current Mac Studio line offers up to 128GB unified memory with M5 Max and up to 512GB with M5 Ultra. That makes large local models possible in a small desktop, but configure memory for the life of the machine: unlike a PC GPU, it is not a part you swap next year. It is also a poor fit if your workflow depends on CUDA-specific tooling or you want to reuse an existing desktop GPU.
Do not solve a capacity problem with two random GPUs
Two cards can help, and modern Ollama scheduling explicitly supports multi-GPU placement, including mismatched GPUs. But it is not equivalent to buying one card with the combined VRAM. The runtime has to split layers and coordinate devices; performance and usable capacity depend on the model, interconnect, host RAM, and backend. Build multi-GPU only after one large card cannot meet your capacity target or after you have measured a model that benefits from the split.
Also separate interactive latency from throughput. A second GPU may help a shared local service handle multiple requests, while doing little to make one agent feel less interruptible. If only you will use the machine, put the budget into enough memory for one complete model plus its working context before chasing concurrency.
Use a boring acceptance test
Before the return window closes, run the agent on one medium-sized task: ask it to inspect a few directories, make a multi-file change, run the test suite, fix one failure, and summarize the diff. Use your intended context setting, with your normal containers and editor open. Record time to first useful response, generation speed after tool output arrives, peak memory in ollama ps or the llama.cpp breakdown, and whether it offloads fully to the accelerator.
If it fits only after you cut context, close every other app, or disable the second session you wanted, it does not fit your workflow. Return it or change the model tier. Local inference gets pleasant when the machine has enough memory that you stop thinking about memory; until then, every agent task turns into capacity planning.
Sources & citations
- [1]Ollama: launch local or cloud models with coding tools
- [2]llama.cpp: fit model parameters to available device memory
- [3]llama.cpp: quantization and model memory/disk requirements
- [4]Ollama: new model scheduling and multi-GPU memory handling
- [5]Ollama: Vulkan GPU acceleration and GGUF support
- [6]NVIDIA: GeForce RTX 5090 specifications
- [7]Apple: Mac Studio with M5 Max and M5 Ultra