Dev Tool Experiences
All articles

· 5 min read

How to Run Cline Fully Local With Ollama or LM Studio

By Z. Saleh

  • tools

Yes: install Ollama or LM Studio, download a local model, then select the matching provider in Cline and point it at localhost. For most developers, Ollama is the quicker repeatable setup; LM Studio is better when you want to inspect model files, loading, and server behavior in a GUI.

This is fully local inference, not a magic security boundary. Cline can still read the workspace and run the terminal commands you approve, and your editor may still have extensions, source control, package registries, and MCP servers that use the network. If “fully local” means an actual air-gapped machine, download the extension, runtime, and model artifacts before disconnecting it.

Start with Ollama if you want the fewest moving parts

Install Ollama, open a terminal, and pull a model. A practical smoke-test model is Qwen2.5-Coder 7B: Ollama lists that variant at 4.7 GB with a 32K context window. That is small enough to establish whether your machine and Cline configuration work, not a promise that it will reliably drive a long multi-file agent task.

ollama pull qwen2.5-coder:7b
ollama run qwen2.5-coder:7b

Exit the interactive chat after it answers once. Ollama normally serves its local API at http://localhost:11434; Cline’s local-model documentation uses that address for its Ollama provider. In the Cline sidebar, open Settings, set API Provider to Ollama, leave the base URL as http://localhost:11434, and choose qwen2.5-coder:7b from the model picker. Local providers do not require an API key.

If the model picker is empty, don’t debug Cline first. Run ollama list; if the model is listed, restart the Cline panel or VS Code and check that nothing has changed the base URL. If ollama run qwen2.5-coder:7b cannot answer a one-line prompt, the agent won’t fix it for you.

Once that works, use a task that exposes bad tool use without risking a repository-wide rewrite: “Read package.json, tell me the test command, then run it. Do not edit files.” A local model that can explain a function but cannot consistently select a terminal command, wait for output, and recover from a failing test will be frustrating in Cline. That is a model limitation, not a reason to keep toggling provider settings.

Use LM Studio when you want more visibility into the runtime

LM Studio gives you the same localhost arrangement, with a UI for downloading models, loading one into memory, and starting the server. Download a chat-tuned coding model in the app, load it, open the Developer tab, and start the server. Its default server listens on http://localhost:1234.

Then set Cline’s API Provider to LM Studio, keep the base URL at http://localhost:1234, and select the loaded model. The non-negotiable distinction is “downloaded” versus “loaded”: a model sitting on disk is not necessarily available to Cline. Verify the server before opening the editor with:

curl http://localhost:1234/v1/models

You should get a JSON model list. If you don’t, start the server with the app’s Developer switch or, if you installed LM Studio’s CLI, run lms server start. LM Studio exposes OpenAI-compatible endpoints, which is why Cline can query it and why the same server is useful for other local tooling.

LM Studio is particularly useful when a run is strange rather than merely slow. Keep lms log stream open while you test Cline. It lets you see requests entering the local server, which separates “Cline never connected” from “the model received a huge tool prompt and produced unusable output.”

Pick a model for agent work, not for a pleasant chat demo

A coding agent needs more than code completion. It needs enough context for the task and dependable tool-call formatting. LM Studio explicitly distinguishes models with native tool-use support from its fallback parsing mode; the fallback can work, but small or non-tool-trained models may emit malformed calls. That failure mode looks like an agent that writes a reasonable paragraph about what it plans to do, then never performs the action.

For a starting point, use an instruct or coding model with tool support, keep the task narrow, and turn on Cline’s Use Compact Prompt setting under Settings → Features. Cline recommends compact prompts for local inference because every extra instruction, tool definition, directory listing, and command result competes for the same context window and slows generation.

Cline’s own hardware guidance is a useful sanity check: 16–32 GB RAM for small or quantized models, 32–64 GB for mid-size models, and 64 GB or more for larger models and bigger contexts. Treat that as capacity planning, not a performance guarantee. Memory bandwidth, GPU offload, quantization, context length, and the rest of your development environment decide whether waiting for a tool call is tolerable.

Don’t start by downloading the largest model page you can find. For example, Ollama’s Qwen3-Coder 30B download is listed as 19 GB, while the 480B local variant lists a 250 GB minimum memory or unified-memory requirement. The bigger model may be appropriate for a workstation built for it; it is not the sensible first diagnostic when your goal is to verify that Cline’s loop works.

Set boundaries before the first real task

Local does not mean automatic approval. Start Cline in a disposable branch and require approval for writes and terminal commands until the model has earned more latitude on your repository. Give it an explicit first job: inspect three files, propose the edit, change one file, then run one named test command. That structure is not ceremonial. Smaller local models lose the thread when you combine discovery, architecture, edits, test repair, and cleanup in one prompt.

  • Use Ollama when you want a scriptable local service: ollama pull, ollama list, then Cline at http://localhost:11434.
  • Use LM Studio when you want to choose and inspect GGUF/MLX-style local models in a UI, load them deliberately, and observe API traffic at http://localhost:1234.
  • Keep the server bound to localhost. LM Studio warns that binding to 0.0.0.0 exposes it beyond your machine; add authentication before putting it on a LAN.
  • Reset the task when context gets bloated. Cline specifically recommends focused tasks and starting a new task as local context grows.

The useful local Cline setup is usually not the one that claims to replace every cloud model. It is the one that handles codebase orientation, contained edits, test runs, and private repositories without a per-request meter or data leaving the machine. Keep a stronger remote option for the stubborn, repo-wide work if you need it—but make the local path boring enough that you will actually use it.

Sources & citations

  1. [1]Cline documentation: Running models locally
  2. [2]Cline documentation: Authorization and local runtimes
  3. [3]Ollama model registry: Qwen2.5-Coder
  4. [4]Ollama model registry: Qwen3-Coder
  5. [5]LM Studio documentation: Local API server
  6. [6]LM Studio documentation: Tool use
  7. [7]LM Studio documentation: Serve on local network
  8. [8]LM Studio documentation: Offline operation