Dev Tool Experiences
All articles

· 7 min read

Context Management Tips for Long Coding Sessions With Open-Weight Models

By C. Volkov

  • tools

Long coding sessions with open-weight models work better when you stop trying to preserve every message. Keep a small, inspectable working set and hand the model a fresh, task-specific brief whenever the conversation starts to accumulate dead branches, stale plans, and pasted output.

A bigger context window helps, but it doesn’t turn an eight-hour agent session into a durable engineering memory. It also competes with KV-cache memory and makes repeated prefills expensive if your serving stack cannot reuse them. The practical setup is a stable prefix, a short state file, and regular clean restarts—not one immortal chat.

Use a handoff file the model can read in under a minute

Put the current task state in the repository, not in an opaque chat transcript. I use a file such as .agent/now.md, generated or edited at the end of a meaningful chunk of work. It should be short enough that you’ll actually maintain it: aim for 300–800 words, not a project diary.

The file needs decisions and evidence, not a narrative. Record the goal, files changed, commands already run and their result, constraints discovered, the next smallest action, and any question that still needs a human answer. Include exact symbols and paths: packages/api/src/auth/session.ts, not “the auth code.” A model can search from a concrete anchor; it tends to improvise from a vague recap.

# .agent/now.md

## Goal
Make session renewal preserve the original `returnTo` path.

## Known facts
- `createSession()` is in `packages/api/src/auth/session.ts`.
- Callback validation rejects absolute URLs by design.
- Existing regression test: `apps/web/e2e/login.spec.ts`.

## Changes made
- Added `returnTo` to the signed session payload.
- No migration required; old payloads must default to `/`.

## Verified
- `pnpm --filter @acme/api test auth/session.test.ts` passes.
- `pnpm lint` has not been run yet.

## Next action
Add the old-payload regression test, then run the focused web E2E test.

## Do not change
Do not loosen callback URL validation.

Start the next run with the handoff file, the relevant repository instructions, and only the files needed for the next action. If the model needs history beyond that, have it retrieve it with git diff, git log -p, rg, or a targeted test command. This is slower than pretending the entire repository fits in a chat, but it is much faster than reviewing a confident change built on an abandoned hypothesis.

Make the reusable prefix truly stable

A coding session usually has a large repeated prefix: your system instructions, repository rules, tool definitions, architecture notes, and perhaps the task handoff. Keep that material in the same order and avoid injecting a timestamp, random run ID, or constantly rewritten status block ahead of it. This matters because vLLM’s automatic prefix caching reuses KV cache only when requests share a prefix, avoiding recomputation of that shared portion. Enable it explicitly when you operate the server:

vllm serve /models/your-model \
  --max-model-len 32768 \
  --enable-prefix-caching

32768 is a starting operating limit, not a quality claim. Set it below the model’s advertised maximum until you know what your GPU can hold alongside the KV cache and your normal parallelism. vLLM documents --max-model-len as the length it will serve, and its cache settings expose GPU-memory utilization separately. If requests begin queuing or failing as a long session grows, lower the served length before deciding the model needs more hardware.

Prefix caching is an inference-cost optimization, not a memory system. A cached prefix faithfully preserves bad instructions too. Keep policy and stable project rules near the front; put volatile discoveries in the handoff file after them. When the task changes from “fix the auth regression” to “redesign account linking,” start a new prefix instead of dragging the old task’s conclusions along.

Budget the context before it fills itself

Reserve room for the model to work. With a 32k served context, don’t allow your tool wrapper to stuff 31k tokens of logs and source into the request, then wonder why the final patch is rushed. Reserve 6k–8k tokens for tool calls, diffs, and the response; treat the remainder as your input budget. The exact split depends on your agent’s tool loop, but choosing one beats letting every retrieval call consume the remaining headroom.

The easiest win is to cap noisy command output. Give the model the failing test section, the stack trace, and the relevant git diff—not 4,000 lines of a monorepo test run. If it asks for more, retrieve more. For logs, preserve the first error, the last 100 lines, and a count of omitted lines. That is enough to guide the next command without donating half the context to progress bars and repeated warnings.

  1. At the 20-minute mark or after a failed approach, ask for a state handoff: decisions, evidence, changed files, open questions, next command.
  2. At roughly 60 minutes, or before switching subproblems, write .agent/now.md, commit or stash the work in progress, and start a fresh conversation.
  3. When a tool returns more than a screenful, summarize it into the handoff and retain the raw artifact on disk, where the model can retrieve a narrow slice later.
  4. Before implementation, make the model restate the next action and the test that will falsify it. If it cannot do that from the handoff, the context is already too muddy.

Don’t hand-roll the chat format

Open-weight models are unusually sensitive to the boring plumbing around messages. Use the chat template shipped with the tokenizer or model integration rather than converting your history into a generic User:/Assistant: string. Hugging Face’s documentation is blunt on this point: chat models are trained on specific control tokens and formats, and the wrong template can materially hurt performance. Inspect the template before you blame the model for a session that gets strange after several turns.

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("your-org/your-model")
print(tokenizer.chat_template)

This is especially important when you restart from a handoff. A clean restart with the correct template is useful. A clean restart with duplicated BOS/EOS tokens, a missing assistant-generation prompt, or a mismatched tool-use template is just a fresh source of confusing behavior.

Treat context shifting as an emergency valve, not continuity

Some local stacks can keep generating by dropping older tokens. llama.cpp exposes context shifting, but its server documentation lists it as disabled by default, and the implementation has cases where it cannot be used, including shared prompts and multimodal contexts. That should tell you how much semantic continuity to assign it: it is a way to prevent an abrupt stop, not proof that the model still has the facts that mattered at the beginning of the session.

If you use it, preserve the handoff separately and test behavior after a shift. Ask a fact with a known answer from the task state—what file owns the migration, what invariant must remain true—and compare the answer before and after. If the model starts inventing old decisions, restart from the handoff. A 15-second reset is cheaper than a clean-looking patch that quietly undoes work from two hours ago.

The operating rule

Keep durable knowledge in files, not chat; retrieve source and logs on demand; and restart when the task boundary changes. Open-weight models give you control over context length, serving, templates, and caching. They do not remove the need to decide what deserves to be remembered. That decision is the part that makes a long coding session stay useful.

Sources & citations

  1. [1]vLLM serve CLI reference
  2. [2]vLLM Automatic Prefix Caching documentation
  3. [3]vLLM model configuration reference
  4. [4]Hugging Face Transformers chat templates documentation
  5. [5]llama.cpp server README