Dev Tool Experiences
All articles

· 7 min read

Self-Hosting an Open-Weight Coding Model for Your Team With vLLM

By C. Wijaya

  • tools

Yes: vLLM is a sensible team-serving layer for an open-weight coding model when you need a stable internal endpoint, predictable data handling, and enough recurring usage to justify owning a GPU service. No: it is not a shortcut around model evaluation, capacity planning, or security; it replaces an API bill with an inference service that somebody now has to run.

Start with one endpoint and a constrained shape

Don’t begin by trying to reproduce every hosted-model feature or by exposing five models to every IDE. Pick one coding model, one internal endpoint, and a context limit that leaves room for concurrent requests. Qwen’s Qwen3-Coder-30B-A3B-Instruct is a reasonable candidate to evaluate: it is a 30B-parameter mixture-of-experts coding model with a published 256K-token context window. That does not mean you should allocate 256K tokens per developer request on day one.

For a two-GPU NVIDIA test box, this is a useful smoke-test shape. The 32K cap is intentional: context consumes KV-cache memory, so treating the model’s maximum advertised context as the default service limit is a quick route to poor concurrency or startup failures.

docker run --rm --gpus all \
  --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
  -e HF_TOKEN \
  -e VLLM_API_KEY \
  vllm/vllm-openai:latest \
  --model Qwen/Qwen3-Coder-30B-A3B-Instruct \
  --served-model-name team-coder \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --generation-config vllm \
  --enable-prefix-caching \
  --api-key "$VLLM_API_KEY"

Use latest only to get the first answer out of the server. Before anybody depends on it, pin a reviewed vLLM image release or digest and lock the model revision in your deployment configuration. The Hugging Face cache mount saves repeated weight downloads; adding a persistent vLLM cache volume later can also avoid redoing compilation work on every container replacement. --ipc=host is not decorative: vLLM’s Docker documentation calls out PyTorch shared memory, especially for tensor-parallel inference.

The other settings are where the operational behavior starts. --tensor-parallel-size 2 splits the model across two GPUs. --gpu-memory-utilization 0.90 deliberately leaves headroom rather than trying to reserve every last byte. --generation-config vllm prevents model-repository generation defaults from quietly becoming your production defaults. And --served-model-name team-coder means clients never need to know the Hugging Face repository name; you can change a model underneath only after a compatibility test.

Make the client setup boring

The useful property of vLLM for a developer-tools team is its OpenAI-compatible API. Put an internal TLS-authenticated gateway or reverse proxy in front of the container, then point tools at a stable base URL and model name. Keep the vLLM port loopback-only on a single host, as in the command above, or private to its Kubernetes network. Do not make a GPU pod’s port 8000 an internet-facing product.

curl --fail --silent --show-error \
  --max-time 10 \
  http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "team-coder",
    "temperature": 0.1,
    "messages": [
      {"role": "user", "content": "Explain why this Go test can race, then propose the smallest fix."}
    ]
  }'

Run that request from the same network path your editor extension, CLI, or agent will use. A green process log is not a working integration. Test the chat endpoint, streaming if your client uses it, tool-call formatting if your agent uses tools, and the largest prompt shape you expect from real repositories. Use a representative fixture repo and saved prompts instead of asking whether the model can solve a hand-picked toy bug.

Treat 32K context as a capacity decision, not a model judgment

A coding agent tends to send repository rules, diffs, test output, and tool transcripts repeatedly. That makes prefix caching worth enabling when requests share a long common prefix, but it does not make context free. The service has to hold cache state while users are waiting. Start at 32K, record how often requests hit that ceiling, then raise it only for a job that actually needs it—large migration plans or repository-wide analysis, for example.

Before opening access to the whole engineering organization, replay a small request corpus at concurrency 1, 4, and 8. Record time to first token, total completion time, error rate, GPU memory use, and how long requests wait before generation begins. vLLM exposes Prometheus-compatible metrics at /metrics; scrape it rather than guessing from GPU utilization alone. GPU utilization near 100% can mean useful batching, but it can also mean one oversized context has turned everybody else into a queue.

This is also where vLLM is bad at hiding mistakes. It will not tell you whether 32K is the right product limit, whether a model’s tool calling matches your agent harness, or whether your prompt construction is wasting half the context window. It gives you a fast server and knobs. You still need an acceptance suite that contains your languages, frameworks, repository layout, test commands, and failure modes.

Put security outside the model server

--api-key is useful for basic client authentication, but it is not a complete production security boundary. vLLM documents that the key applies to selected API path prefixes, while other endpoints on the same server can remain unauthenticated. Put authentication, TLS termination, rate limits, audit logs, and network policy in a gateway or proxy; use firewall rules so distributed-worker and cache-transfer ports are reachable only on trusted internal networks.

This matters more for coding models than for a casual chat bot. Prompts can contain proprietary source, stack traces, credentials pasted by mistake, and tool output. Decide up front whether you retain request bodies, how long you retain proxy logs, and who can access the Hugging Face token used to fetch weights. “Self-hosted” narrows where inference runs; it does not automatically create a safe data-handling policy.

Operate the upgrade path before you need it

Keep one production model alias, one candidate alias, and a small repeatable test set. Run the candidate against the same prompts before changing team-coder. Check compilation and test-fix tasks, structured output, refusal behavior, tool-call arguments, long-context requests, and the client’s retry behavior. Then drain or restart deliberately. A model upgrade that preserves the HTTP schema can still change formatting, stop behavior, or how eagerly it edits files.

Self-hosting pays off when the endpoint becomes infrastructure your tools can rely on: one address, one client contract, known telemetry, and a model change process. It is a poor fit when the team wants frontier-model capability with no GPU operator, no evaluation harness, and no appetite for debugging CUDA, model compatibility, or a full disk cache at 9 a.m. Start with the narrow service above. If it handles real developer traffic cleanly for a few weeks, add replicas and routing—not another dashboard.

Sources & citations

  1. [1]vLLM Docker deployment documentation
  2. [2]vLLM serving CLI reference
  3. [3]vLLM OpenAI-compatible server documentation
  4. [4]vLLM security documentation
  5. [5]vLLM production metrics documentation
  6. [6]Qwen3-Coder-30B-A3B-Instruct model card