Dev Tool Experiences
All articles

· 7 min read

What Terminal-Bench Actually Tells You About a Coding Agent

By T. Wijaya

  • tools

Terminal-Bench is a useful signal if you’re choosing an agent that must work independently in a terminal: inspect files, install or use tools, edit code or configuration, run checks, and leave a machine in a state that passes tests. It does not tell you that the agent will be pleasant in your editor, good at ambiguous product work, economical on your codebase, or trustworthy with credentials outside a sandbox.

That sounds obvious, but leaderboard screenshots routinely collapse all of those questions into one percentage. Don’t do that. Treat a Terminal-Bench score as a test of an agent-and-harness configuration under a particular dataset release, model, budget, and execution environment—not as a property stamped permanently on a product.

It measures completed work, not convincing chat

The good part is the verifier. A Terminal-Bench task packages an instruction, a containerized environment, a human-written solution, and tests. The agent gets a terminal, makes whatever sequence of moves its harness permits, and passes only when the resulting environment satisfies the test suite. That is substantially closer to “can it actually finish the task?” than evaluating a patch as text or asking a model to explain a command.

This catches the unglamorous failure modes that matter once an agent has a shell: assuming an undocumented CLI flag, editing the wrong file, stopping after a command exits zero, failing to inspect generated output, or making a numerically plausible but invalid result. The original task examples include an agent that completed a scientific fit yet failed to notice nonsensical parameters. Passing tests is not perfect truth, but it is a much better stopping condition than the agent saying it is done.

For a developer, the practical reading is: a high score is evidence of operational competence. The agent can likely loop through inspect → change → run → diagnose → retry without requiring a human to translate every observation into the next command. That is the capability you want for contained maintenance, migration prep, dependency investigation, and CI-style repair work.

The score belongs to the whole setup

Terminal-Bench evaluates an agent, not merely a foundation model. Planning prompts, tool schemas, shell interaction, context management, retry policy, time limits, model parameters, and permission rules can all change the result. Two entries using the same underlying model can therefore be meaningfully different products. Conversely, a great result from an agent configured with a generous budget says little about the default configuration you installed locally.

Before you use a score to make a buying decision, write down the comparison tuple: dataset version, agent version, model identifier, number of attempts, timeout or token budget, allowed network access, sandbox, and pass metric. If a vendor cannot provide that, you have marketing, not an evaluation. A result from Terminal-Bench 2.0 is also not directly comparable with one from the current line: Terminal-Bench is explicitly maintained as a continuous benchmark, and the project announced Terminal-Bench 4.0 on August 28, 2026 after calibrating resources, fixing tasks, and removing saturated tasks.

That churn is not a flaw. Old static benchmarks become training targets and product roadmaps. It does mean you should record the exact dataset tag in an internal spreadsheet or CI artifact. “We scored 61% on Terminal-Bench” is nearly useless six months later; “61% on this release, with five independent runs and this model budget” can still guide a decision.

What it is bad at measuring

Terminal-Bench is deliberately outcome-driven: agents may solve tasks through different approaches so long as the verifier accepts the final state. That is appropriate for measuring task completion, but it leaves important engineering questions outside the number. A passing agent may have produced a hard-to-review diff, used too many tokens, deleted and regenerated files you wanted preserved, or taken an approach your organization would reject in a real repository.

It also does not represent the full social environment of software work. Your actual task may begin with a vague ticket, require recognizing that the ticket is wrong, involve undocumented conventions, wait for a reviewer’s answer, or be constrained by secrets, service ownership, release policy, and blast radius. A container task with a verifier cannot establish that an agent has judgment about any of those things.

Security is the other missing inference. Terminal-Bench’s container setup is useful precisely because terminal agents are not yet reliable enough for unrestricted access. A good benchmark result is evidence that the agent is effective when permitted to act; it is not evidence that your approval model, egress controls, credential scope, or audit trail are adequate. If anything, a more capable terminal agent raises the value of testing those controls.

Use it as a filter, then run your own small gate

If you’re comparing agents for terminal-heavy work, use Terminal-Bench to narrow the field, not select a winner. First, require a recent result with complete run settings. Then test the finalists on five to ten scrubbed tasks from your own workflow: a failing CI job, a dependency upgrade, a multi-file refactor with a non-obvious invariant, a configuration repair, and a task where the correct action is to stop and ask a question.

  1. Pin the benchmark release and configuration before comparing public scores. Don’t compare a current run against an undated chart or a result from an older release.
  2. Check reliability, not only pass rate. Run each internal task more than once and save transcripts, diffs, commands, elapsed time, and model cost.
  3. Score the behavior you actually need: tests pass, diff is reviewable, no prohibited files or network calls occur, and the agent asks for approval at your chosen boundary.
  4. Include one deliberately misleading README comment or irrelevant instruction in a safe test task. You need to see whether the agent can use relevant environmental information without blindly executing everything it reads.
  5. Keep the benchmark sandboxed. Promotion to a repository with real credentials should require a separate permissions and audit review.

Run the harness before you trust a published number

If you build agents or operate a serious internal evaluation loop, run the benchmark’s oracle first. The project recommends running the oracle solution five times to confirm that the tasks are stable in your chosen sandbox before attributing failures to your agent. Its current repository documents the following Modal-backed smoke test:

uv tool install 'harbor[modal]'
uv run harbor run -d terminal-bench/terminal-bench@latest \
  -k 5 \
  --agent oracle \
  --n-concurrent 500 \
  --env modal

Don’t copy the concurrency value blindly into a personal account. The actionable point is the five oracle runs: if the known solution flakes in your environment, your agent comparison is already contaminated. Once the oracle is stable, change only one dimension at a time—agent harness, model, or policy—and preserve the run artifacts.

The best use of Terminal-Bench is therefore modest and valuable. It tells you whether an agent configuration has demonstrated end-to-end terminal task completion under real verification pressure. Use that to avoid tools that cannot finish contained work; use your own tasks to decide whether the survivors belong in your day.

Sources & citations

  1. [1]Terminal-Bench repository: current run commands, oracle validation guidance, and continuous-release model
  2. [2]Terminal-Bench paper: Terminal-Bench 2.0’s 89-task design, unique environments, human solutions, and comprehensive verification
  3. [3]Terminal-Bench launch post: task structure, terminal-agent failure examples, sandboxing rationale, and harness integrations
  4. [4]Terminal-Bench news index: Terminal-Bench 4.0 announcement dated August 28, 2026