Dev Tool Experiences
All articles

· 8 min read

Best Local LLM for Coding on a 24GB GPU: 2026 Pick

By Q. Hosseini

  • tools

The best local LLM for coding on a 24GB GPU is Qwen3-Coder-30B-A3B-Instruct, using the 19GB Q4_K_M build distributed by Ollama. It is the best fit because it is a coding- and tool-use-oriented 30.5B-parameter MoE model, yet activates only 3.3B parameters per token; in practice, that leaves enough of a 24GB card for a useful local coding session instead of forcing the model weights onto system RAM.

That recommendation is for the developer who wants an agent to inspect a repository, make a bounded edit, run a command, read the failure, and try again. If you mainly want a fast local reasoning model with more VRAM headroom, choose gpt-oss:20b instead. If you expect to stuff an entire monorepo into context, neither answer is magic: 24GB is enough for a good interactive coding model, not for carefree 128K-token agent runs.

Why Qwen3-Coder-30B is the best 24GB GPU coding model

The hardware fit is unusually clean. Ollama’s qwen3-coder:30b package is Q4_K_M and listed at 19GB. That is not “19GB of GPU use”: runtime buffers and the KV cache still consume memory. But it is a realistic starting point on an RTX 3090, RTX 4090, RTX 5090-class 24GB card, or a 24GB Radeon setup supported by your runtime. The full upstream model has 30.5B total parameters, 128 experts, and eight active experts, with 3.3B active parameters per token. That combination is why it can feel responsive enough for an edit-test loop while retaining more coding-specific capacity than the smaller general-purpose models.

It also has the properties that matter more than a single code-completion benchmark once you put it behind tools. Qwen publishes native 262,144-token context support, a function-call format intended for coding agents, and fill-in-the-middle support. The headline context number is architectural capability, not a promise that 24GB will hold 262K tokens. Treat it as a model that can work with long context on appropriate hardware; on one 24GB GPU, start much smaller and let repository search decide which files earn their way into the prompt.

The less flattering part: Qwen3-Coder-30B is non-thinking only. It will not spend an explicit reasoning phase exploring a tricky design before responding, and it can confidently choose an attractive but wrong cross-file abstraction. It is also not a model for parallel agent swarms on one card. Give it one focused task, a test command, and a small enough context that the weights remain fully on the GPU.

What to run instead: gpt-oss:20b or a smaller model

Pick gpt-oss:20b when headroom is the requirement. OpenAI specifies 21B total parameters and 3.6B active parameters, a 128K context window, configurable reasoning effort, and local operation in as little as 16GB of memory. Ollama’s MXFP4 package is 14GB. On the same 24GB GPU, that extra space can make the difference between a 16K or 32K context experiment and an out-of-memory error. It is a credible default for general code review, debugging, shell-command planning, and tool calling.

But it is not the first pick here because “fits easily” and “best at coding” are separate claims. gpt-oss:20b is a general open-weight reasoning model trained with an emphasis on STEM and coding, whereas Qwen3-Coder-30B is the purpose-built coding option and is listed by Ollama as its recommended local coding model. Use gpt-oss when you value reasoning controls and spare VRAM more than a coding-specialized model; use Qwen3-Coder when the day is mostly repository work.

A smaller 7B–14B code model still has a place: laptop-adjacent machines, autocomplete-like requests, fast throwaway scripts, or a second local service. What it is bad at is the work that causes people to install a 24GB GPU in the first place—following a change across unfamiliar modules, recovering after test failures, and preserving constraints that were only implied by several files. Do not quantize an 80B coding model down to something extreme merely to say it runs; a sensible 30B Q4 model is usually the less annoying engineering tool.

How much context fits on a 24GB GPU?

Start Qwen3-Coder at 8K or 16K context, verify it is entirely GPU-resident, then increase only when the job needs it. The model file consumes about 19GB before the cache and compute allocations. Context is not free: every extra token expands the KV cache, and the runtime needs additional working memory. The practical rule is simple: a 24GB card is a weight budget plus a context budget, not a 24GB model-size allowance.

Ollama now defaults cards in the 24–48 GiB range to a 32K context length, while warning that larger context requires more VRAM. That default can be too aggressive for a 19GB model on a nominal 24GB card, depending on driver overhead and what else is using VRAM. Check rather than guess:

ollama pull qwen3-coder:30b
ollama run qwen3-coder:30b

# In another terminal, confirm the model did not spill to CPU.
ollama ps

Read the PROCESSOR column. You want 100% GPU; CPU/GPU splitting is often the explanation for a local model that technically loaded but now emits code at the pace of a slow CI job. If you need to set a conservative default while testing, start the server with a 16K context:

OLLAMA_CONTEXT_LENGTH=16384 ollama serve

For a lasting per-model setting, use a Modelfile instead. num_ctx controls the context window used by the model. Do not claim a universal tokens-per-second number for this setup: GPU generation speed changes materially with card generation, power limit, driver, context length, prompt ingestion, Flash Attention support, and whether even a fraction of the model spilled into host memory.

Best local LLM setup for coding agents on one GPU

Use Ollama first unless you have a reason to own the inference stack. It gives you a local server on port 11434, a one-command model pull, and an API that agent tools commonly understand. Start with qwen3-coder:30b, keep one model loaded, and ensure the machine has enough system RAM for the model download, runtime, IDE, language server, and test suite. GPU VRAM determines whether inference stays fast; host RAM determines whether the rest of your development environment remains usable.

Then change the workflow, not just the model. Ask the agent to inspect two or three named directories before editing. Tell it the exact test or lint command it must run. Keep generated logs out of the prompt unless the failure points nowhere else. Start a fresh task after a large completed change rather than carrying a stale conversation through the next unrelated bug. This is how a 24GB setup stays useful: retrieval and task boundaries substitute for the context window you do not have.

  1. Pull and run qwen3-coder:30b with Ollama.
  2. Check ollama ps; fix any CPU offload before judging quality or speed.
  3. Begin at 8K–16K context for Qwen3-Coder and increase only after a real task needs more repository state.
  4. Give one testable task at a time: inspect, propose, edit, run the command, summarize the diff.
  5. Keep gpt-oss:20b installed as the lower-memory fallback for longer-context or reasoning-heavy sessions.

Is local coding worth the setup?

Yes when the work is repetitive enough that local availability matters: private repositories, air-gapped or unreliable-network environments, experiments that would otherwise consume metered API calls, and long afternoons of small repairs. It is not automatically cheaper once you count a GPU purchase, electricity, maintenance, and the time spent debugging a runtime. And it is not automatically more private if your agent still calls remote tools, downloads dependencies, or sends telemetry elsewhere. Local inference solves the model-request boundary; audit the rest of the toolchain separately.

A 24GB GPU is also a compromise rather than a destination. It buys a strong quantized coding model, one interactive session, and disciplined context management. It does not buy the largest current local coding models at comfortable quantization, high concurrency, or a license to make agents read every file before every edit. If that constraint sounds acceptable, Qwen3-Coder-30B is the one to install first.

Use your 24GB local model from Cline

If the missing piece is not inference but an agent interface inside the editor, Cline documents a local-model workflow with Ollama, LM Studio, and Atomic Chat. For Ollama, its setup is concrete: select the Ollama provider in Cline Settings, use http://localhost:11434, select the locally loaded model, and enable Use Compact Prompt. That last setting matters on a 24GB card because smaller task context is directly easier to serve.

Cline describes its open-source extension as free for individual developers; local models have no per-request inference charge beyond the hardware you provide. Its pricing page also says you can bring your own API keys or use its provider, while its local-model documentation explicitly covers local inference. That makes a Qwen3-Coder or gpt-oss setup practical when you want to switch between a private local model for routine repository work and another provider for the rare task where the local model is the bottleneck.

Sources & citations

  1. [1]Qwen3-Coder-30B-A3B-Instruct model card
  2. [2]Ollama registry metadata for qwen3-coder:30b
  3. [3]Ollama context-length documentation
  4. [4]Ollama Anthropic compatibility and local-model recommendations
  5. [5]OpenAI gpt-oss release documentation
  6. [6]Cline local-model documentation
  7. [7]Cline pricing