Dev Tool Experiences
All articles

· 6 min read

Cut Your AI Coding Bill by Limiting Agent Loops, Not Code Quality

By B. Bautista

  • tools

You can usually cut an AI coding bill without making the agent noticeably worse: use autonomous mode only for work that needs it, put a hard ceiling on retries, and stop paying to re-send the same repository instructions. Downgrading every task to a cheaper model is the blunt instrument; reducing pointless agent turns is the part that preserves quality.

The useful unit to manage is not “tokens per developer.” It’s a task that produced a reviewable diff, passed the relevant check, and did not need to be re-run from scratch. A five-minute agent session that fixes a flaky test can be cheap. A 45-minute session that repeatedly reads the same monorepo, tries three broken commands, and ends with “I need clarification” is not.

Stop using agent mode for bounded edits

The easiest savings are hiding in tasks where you already know the files and the expected change. Don’t send an autonomous agent to discover a two-file rename, add a missing null check, or update a known API call. Use an edit-oriented workflow, supply the files, and run the test yourself. GitHub describes Copilot Edit mode as the option for quick, specific updates to a defined set of files, specifically noting that it gives you control over how many LLM requests are used. Its Agent mode is for multi-step work that needs tool use and iteration—and each prompt in that mode consumes AI credits.

Write the request so the agent has no reason to search broadly. This is the difference between a small edit and an exploratory session:

In packages/api/src/auth.ts, replace the deprecated verifyToken call with verifySession.
Only change auth.ts and auth.test.ts.
Run: pnpm --filter api test auth
Do not change dependencies or formatting outside those files.
Stop after the test passes; return the diff summary in three bullets.

That last line matters. “Find every related use and modernize the pattern” sounds productive, but it turns a bounded edit into repository archaeology. Ask for that broader cleanup only when you would genuinely assign the same open-ended task to a teammate.

Put a turn limit on unattended work

An agent that can run commands will use extra turns when a test fails, a command is unavailable, or its first diagnosis is wrong. Those recovery loops are where spend becomes hard to predict. For unattended Claude Code jobs, the CLI has a concrete guardrail: --max-turns limits agentic turns in non-interactive mode.

claude -p --max-turns 3 \
  "Fix the failing parser test in packages/parser. Run pnpm test -- parser. If the first repair does not make the test pass, stop and report the blocker."

Three turns is not a universal magic number. It is a useful default for a tightly scoped fix: inspect, edit, verify. If the task needs schema migration design, unfamiliar infrastructure investigation, or a change across services, give it a larger budget deliberately after a plan review. The important part is that “keep trying until it works” is no longer the default billing policy.

This is bad at problems with genuinely coupled failures. If one failing integration test requires a local emulator, seeded data, and a contract update in another repository, a low turn cap just creates another run. In that case, spend one cheap interaction producing a plan and acceptance criteria, then approve a larger execution budget once you know what the agent is being asked to do.

Escalate models after evidence, not by habit

Use your normal/default model for well-specified edits and mechanical test repairs. Move to the expensive model when the first run shows a real reasoning problem: ambiguous ownership, a risky migration, a design tradeoff, or a failure that survives a focused attempt. Don’t pay premium-model rates merely because the task sounds important in a ticket title.

Make escalation reproducible. A useful team rule is: one bounded run with an exact file set and test command; if it fails, capture the failing command, output, and current diff; then escalate with that evidence. The second model gets a cleaner problem and doesn’t have to spend tokens rediscovering it.

Also watch output, not just context. On OpenAI’s standard API rate card, gpt-5.3-codex is listed at $3.50 per million input tokens and $28 per million output tokens, while cached input is $0.35 per million. That doesn’t mean “never ask for an explanation.” It means don’t routinely ask an agent to narrate every file it read, every rejected hypothesis, and every command it ran after the logs and diff already exist.

Make repeat context stable enough to cache

If you operate an internal coding harness rather than only an IDE subscription, prompt caching is often the least controversial cost reduction because it doesn’t change the model’s capability. OpenAI’s caching documentation says cache reuse depends on an unchanged prompt prefix. Put stable material first: repository rules, tool definitions, coding conventions, and the durable portion of your system instructions. Put branch names, timestamps, ticket text, and the current task later.

For supported newer OpenAI models, the documented minimum cacheable visible prefix is 1,024 tokens. Cache reads can be much cheaper than uncached input, but writes have a cost too, so this is not a reason to pad every prompt. It pays when multiple requests really share the same beginning. It is a poor fit for one-off, highly variable requests where every run starts from different instructions and different tools.

The practical check is simple: log cached-input tokens beside total input tokens for agent runs. If cache-hit rate is near zero, inspect what changes near the beginning of the rendered request. Tool schemas reordered by a feature flag, a timestamp in developer instructions, or injecting a whole new tool list per task can erase the savings.

Budget the workflow where it can actually be enforced

A spreadsheet review at month end tells you who spent money; it does not stop the next runaway run. Put limits at the point where calls leave your network: provider project budgets, a gateway, or the agent runner. Anthropic’s Claude Code gateway guidance explicitly calls out centralized usage tracking, budgets, and rate limits as gateway functions. If you use a third-party proxy, treat it as production infrastructure: test fail-closed behavior, protect its credentials, and remember Anthropic does not audit or maintain those tools.

Track four fields per run: repository, task type, model, and terminal outcome. “Terminal outcome” can be as plain as merged, useful-diff-not-merged, no-diff, or human-took-over. After two weeks, you can find the expensive pattern worth fixing: perhaps autonomous review is repeatedly timing out, or test repair succeeds with the cheaper model, or a particular MCP tool triggers long detours.

Start with one policy tomorrow: no unbounded autonomous runs. Give every agent task a file or directory boundary, one verification command, a turn budget, and a stop condition. You’ll remove a lot of waste before you have to debate models, negotiate another seat tier, or tell engineers to use AI less.

Sources & citations

  1. [1]Anthropic Claude Code CLI reference
  2. [2]GitHub Docs: Asking GitHub Copilot questions in your IDE
  3. [3]OpenAI API pricing
  4. [4]OpenAI API prompt caching guide
  5. [5]Anthropic Claude Code LLM gateway configuration
Cut Your AI Coding Bill by Limiting Agent Loops, Not Code Quality | Dev Tool Experiences