Dev Tool Experiences
All articles

· 8 min read

Best Local Coding Models for Apple Silicon Macs: 2026 Picks

By F. Okonkwo

  • tools

The best local coding models for Apple Silicon Macs split cleanly by unified memory: run Qwen3.6 35B A3B Coding NVFP4 on a 32GB-or-larger Mac, Devstral Small 2 when you want a capable 15GB coding-agent model, and Qwen2.5-Coder 7B on 16GB machines. The practical constraint is not the parameter count in a release post; it is whether the model, its context cache, your editor, Docker, and the test suite can coexist without macOS applying memory pressure.

These are local-first recommendations for developers who want the model weights on the Mac, not a cloud endpoint selected from a desktop app. Ollama’s current Apple Silicon runtime uses MLX and unified memory, which is a useful improvement for Macs—but it does not repeal the memory math. Leave headroom, start with a bounded context, and promote a model only after it can edit, test, and recover from a failed command in one of your real repositories.

Best local coding model for a 32GB Apple Silicon Mac

Pick qwen3.6:35b-a3b-coding-nvfp4 first. The Ollama build is a 24GB download, carries tool and thinking support, and is specifically packaged for coding. It is the current high-ceiling choice for an M-series Mac that has enough unified memory to make a model feel like a working tool rather than a demo. The model is recent, so verify it against your language, framework, and agent harness before standardizing on it; new weights can have rough edges in tool-call formatting and long-running edits.

ollama run qwen3.6:35b-a3b-coding-nvfp4

On a nominal 32GB machine, call this a lower bound rather than a comfortable target. A 24GB model file is not the full process footprint: prompt tokens and generated tokens need KV cache, and macOS still needs room for everything else. Shut down a heavy container stack before asking it to read a large monorepo. If Activity Monitor shows yellow or red memory pressure, reduce the context length or use the 20GB qwen3.5:27b-coding-nvfp4 instead. A model that starts quickly but spends each tool round paging is worse than a smaller one that can finish the test-fix loop.

Why not automatically choose Qwen3-Coder 30B? It remains a sound alternative: its Ollama package is 19GB, has a 256K native context window, and uses a mixture-of-experts design with 3.3B activated parameters. But Qwen3.6 is the newer coding-focused option to test first on this hardware. Keep the 30B model around if its tool behavior is more reliable for your existing prompts; model choice is often more sensitive to the agent’s template than a leaderboard suggests.

Best local coding model for codebase agents on a Mac

Use devstral-small-2 when the workflow is “inspect files, make coordinated edits, run commands, repair what the command found.” Its 24B Q4 build is a 15GB download, has a 384K context window, accepts image input, and is explicitly oriented around tool use, codebase exploration, and multi-file changes. That makes it the practical second pick for a 24GB or 32GB Mac and often the less stressful default when your agent will spend more time in the terminal than composing an explanation.

ollama run devstral-small-2

Its weakness is also obvious in use: it is an agent model, not a guarantee that an agent should get broad permissions. It can confidently follow a bad assumption across several files. Start it in a branch; require a plan for migrations, auth changes, and generated-code updates; and make the test command part of the request. The 15GB Q4 package is the reasonable Mac version. The Q8 package is 26GB and the FP16 package is 48GB, so moving up in quantization can erase the headroom that made the model pleasant in the first place.

Best local coding model for a 16GB MacBook

Use qwen2.5-coder:7b for everyday contained work: explain a stack trace, write a focused test, implement a function from an existing interface, or produce a review checklist for a diff. The default package is 4.7GB with a 32K context window. That leaves room for VS Code or JetBrains, a browser, and normal development processes on a 16GB Mac in a way that 14GB–24GB models generally do not.

ollama run qwen2.5-coder:7b

# Confirm which models are still resident after an agent run
ollama ps

Do not confuse “it fits” with “it replaces the larger model.” A 7B model is good at local transformations with enough surrounding code supplied. It is noticeably less dependable at deciding which four packages need changing, preserving a cross-cutting invariant, or operating a terminal loop without supervision. Keep requests narrow: name the files, state the acceptance test, ask for a diff, and review it. If you need vision input or a much longer advertised context on the same machine, Qwen3.5 9B is a 6.6GB general-purpose alternative—but Qwen2.5-Coder remains the better starting point when source code is the main input.

Which model should run on a 64GB or 96GB Mac?

Try qwen3-coder-next if you have 64GB unified memory and a workload that genuinely benefits from long agent context. Its Q4 package is 52GB, offers 256K native context, and uses an 80B-total/3B-active mixture-of-experts architecture. It is aimed directly at tool-using coding workflows. A 64GB Mac can load it, but “load” is not the same as “has room for a 64K-context agent session plus the rest of your workday”; 96GB or more is the sensible place to expect fewer compromises.

This is not the model to download merely because it has the largest number on the page. For a service with clear ownership and a focused change, Devstral Small 2 or Qwen3.6 will usually get you to a reviewable diff with less startup and memory cost. Use Qwen3-Coder Next when an agent repeatedly needs broad repository awareness and the Mac is dedicated enough that a 52GB resident model is not competing with emulators, local databases, and several containers.

Is GPT-OSS 20B good for local coding on Apple Silicon?

Yes, gpt-oss:20b is worth keeping as a reasoning-oriented second opinion. Ollama’s package is 14GB and supports tool use, structured outputs, and thinking. It is a particularly reasonable choice when the bottleneck is interpreting an unfamiliar failure or planning a change, rather than writing a large volume of code. It is not the automatic winner for codebase editing: the model is broader than a code-specialist, so test it against your repository tasks instead of assuming the name or reasoning mode settles the question.

ollama run gpt-oss:20b

How to configure a local coding model on macOS

Install the current Ollama app, pull one model, and give it a real task before wiring it into an agent. Ollama recommends at least a 64K context length for coding tools; that is useful for repository work, but it is not a setting to maximize blindly on a laptop. Start at 16K or 32K on a 16GB or 32GB Mac, then increase it only if the agent demonstrably loses necessary context. The native context number in a model card describes what the model can accept, not what you can afford to retain concurrently.

  1. Pull one model rather than five: ollama run devstral-small-2 is a sensible 24GB/32GB trial.
  2. Ask it to complete a production-shaped task: find the handler, add a failing test, implement the narrow fix, and run the project’s test command.
  3. Watch memory pressure and use ollama ps after the run. Stop other models before deciding that a candidate is too slow.
  4. Record failure modes, not just success: wrong file selection, invented APIs, malformed tool calls, skipped tests, or a patch that fails lint.
  5. Only then connect it to an IDE agent. Keep approvals on for shell commands and file changes until you have seen the model handle your repository’s conventions.

What local models are bad at

They are bad at being invisible infrastructure. The model is local, but the agent still has file-system and shell privileges you grant it. They are also bad at giant unstructured tasks: “modernize the backend” encourages broad edits with weak success criteria, and no 256K context claim fixes that. Finally, local models are not automatically cheaper if your workflow retries expensive agent loops for half an hour. The reliable local workflow is constrained: a branch, an explicit test command, a bounded task, reviewable diffs, and a larger or hosted model available for the genuinely hard problem.

Run these local coding models with Cline

If you want to test the same local model inside an agent rather than in a terminal chat, Cline is an open-source coding agent that its site says runs in the editor, terminal, and SDK; it supports local Ollama and LM Studio as well as OpenAI-compatible endpoints. Its relevant advantage here is model choice: point the agent at the model that fits your Mac today, rather than tying the workflow to a bundled inference plan.

For individual developers, Cline’s open-source offering is free and its pricing page says inference can be BYOK or purchased at cost, without a subscription or seat fee for the open-source version. That makes it a practical harness for comparing Devstral Small 2 against Qwen3.6 or a smaller Qwen2.5-Coder on the exact repository and approval policy that matter—not a synthetic prompt in a model browser.

Sources & citations

  1. [1]Ollama: Qwen3.6 35B A3B Coding NVFP4 model page
  2. [2]Ollama: Apple Silicon MLX runtime announcement
  3. [3]Ollama: Devstral Small 2 model page and tags
  4. [4]Ollama: Qwen3-Coder and Qwen3-Coder Next model pages
  5. [5]Ollama: Qwen2.5-Coder model sizes and context windows
  6. [6]Ollama: GPT-OSS 20B model page
  7. [7]Ollama: coding-model integration and 64K context recommendation
  8. [8]Cline pricing