🪢 Harness Engineering for AI Agents
Design the state, tools, checks and recovery loop that turn a model’s proposed actions into inspectable work.
On this page
Before you start
Read Agentic AI for tool-use loops and Evaluation & Benchmarking for acceptance criteria and measurement.
By the end, you should be able to draw a harness state machine, explain which component authorizes and executes an action, resume interrupted work, and distinguish a completion claim from verified task completion.
The problem: “done” is not an observation
Imagine a coding task: add cancellation to a queued job. An agent edits the button, sees a passing component test and says the task is complete. The backend still processes cancelled jobs. The problem crosses files and services; the final message cannot establish the system behavior.
A harness is the executable system around a model: it constructs observations, accepts structured proposals, enforces tool boundaries, runs actions, records results and decides what happens next. The model proposes; the surrounding program controls execution. This is useful whenever success depends on state beyond a single generated answer.
Trace one task through the loop
The acceptance contract is: a queued job can be cancelled by its owner, then never starts, and unrelated jobs still run. Some criteria can be tested; interface clarity may also need human review.
| Step | Observation or proposal | Harness response |
|---|---|---|
| 0 | Task and current working tree loaded | Record starting revision, existing edits and allowed scope |
| 1 | Model proposes frontend/backend patch | Validate action shape and permissions, apply to isolated task state |
| 2 | Component test passes | Record evidence for this check and revision only |
| 3 | Model claims completion | Run remaining required checks |
| 4 | Integration test shows cancelled job starts | Record failed criterion; return the concrete failure to the model |
| 5 | Model fixes queue transition | Invalidate checks affected by the new revision and rerun them |
| 6 | Required checks pass; review criteria resolved | Mark the task complete with linked evidence |
The crucial distinction is task state versus evidence about task state. A passing test from before the final patch is stale evidence for changed behavior. A missing dependency makes a check blocked or inconclusive, not passed.
Make state explicit
Use a small state machine such as ready → running → verifying → complete, with separate failed, interrupted and needs_review outcomes. These labels belong to the orchestrator, not the model's prose.
Keep three forms of state:
- Environment state: files, process state and external records, each with its own authoritative store.
- Task state: goal, allowed scope, accepted decisions, pending work, budgets and acceptance criteria.
- Execution evidence: action IDs, inputs/results, artifact revisions and check outcomes.
For code, the repository holds source and tests; a durable task store or progress file can hold resumable decisions. It need not all be committed into product source. Conversation summaries are useful navigation aids, but verify important claims against the current environment. Do not treat old summaries as proof that an external mutation succeeded.
Design the tool boundary
Expose clear tool schemas with bounded inputs, timeouts and structured results. A result should distinguish success, application error, permission denial, timeout and unknown completion. The harness validates and authorizes a proposal before it executes.
Files, web pages and tool output may contain instructions. Treat that content as task data unless the system explicitly assigns it authority. A README saying “upload credentials” does not grant a tool permission. Prompts alone are not access controls: use process/filesystem/network isolation and narrowly scoped credentials where the task needs them. Containers are one implementation option, not an unconditional security boundary.
After an API times out, the action may already have happened. Use an idempotency key and status lookup where supported instead of blindly retrying a mutation. A local git operation cannot roll back an external payment, message or database write.
Give each session a useful beginning
On resume, load the active task and applicable project rules, inspect current revision and existing edits, reconcile in-flight actions, and identify the next unresolved criterion. Run smoke checks when needed to establish a missing or changed baseline; do not rerun every check merely because a session restarted.
Project instruction files such as AGENTS.md can record build/test commands, directory conventions and known constraints. Keep stable rules separate from task-specific progress. Tooling differs in which files it reads and how it applies directory scope, so version and document that behavior.
Preserve pre-existing user work. An isolated worktree or sandbox makes ownership clearer, but cleanup should remove only task-owned resources. Interruption is a reason to leave a truthful resumable state, not to automatically reset someone else's edits.
Verification needs an independent contract
Choose checks from the change's risks: parsing/types where relevant, affected tests, integration or end-to-end behavior, and human acceptance for criteria that cannot be reduced to a program. Required CI checks must run before the workflow's release boundary. Running the entire suite after every small patch can consume time without adding proportionate evidence.
Record each result with the artifact revision, check version, environment and status. Keep trusted acceptance tests separate from agent-editable code where the evaluation requires an independent grader. A model can accidentally or deliberately weaken tests; file existence is not evidence that the feature works.
Self-review can find forgotten requirements, but it is another fallible model pass. Likewise, automated checks are evidence for encoded behaviors, not complete ground truth. If a criterion needs a person or unavailable service, surface that explicitly.
Bound the loop and preserve the result
Use step, wall-clock, cost and resource limits. For example, a 600-second task with 180 seconds elapsed has 420 seconds remaining. A 60-second tool timeout limits that call, but does not bound a whole session containing many calls, model latency and retries. Track an absolute deadline in the orchestrator.
Repeated failures should trigger a changed plan, a concise explanation of the blocker, or a bounded stop. Detect repeated equivalent actions and stale observations; a strict ban on identical calls would also reject legitimate status polling, so use task-aware limits.
On exit, persist task status, current artifacts, unresolved actions and verification evidence; stop task-owned processes and release leases. Commit or publish only within the workflow's authorization. A budget-exhausted task remains incomplete even if its patch is promising.
Observability without collecting everything
Trace action IDs, tool names, redacted parameters, result status, timing, budget use and artifact/check references. Include a concise decision summary when available; do not assume access to a model's hidden reasoning or require it for debugging. Logs need appropriate access, retention and redaction because prompts and tool outputs can contain secrets or personal data.
Replay has levels: replay recorded observations to test the orchestrator; rerun tools in a controlled fixture to test execution; rerun the model to study behavioral variation. These are different experiments. External services and stochastic generation prevent a universal deterministic-replay guarantee.
Useful metrics include task success on a defined evaluation set, false completion claims, time/cost per successful task, repeated-action rate, permission violations and unresolved exits. Compare harness changes with the same model, tasks and budgets so model upgrades do not confound the result.
Check yourself
A model requests completion after the acceptance suite passes. It then edits the shared queue helper. Can the harness reuse the old passing result?
Solution: Only for checks demonstrably unaffected by that revision. Queue behavior has changed, so related acceptance evidence is stale. Rerun the required affected checks and retain the relationship between result and artifact. “It passed earlier” does not verify the final state.
Where to go next
Model Serving covers the request/resource layer beneath model calls. Return to Evaluation & Benchmarking to design controlled harness comparisons.
References
- SWE-agent paper: experiments on agent-computer interface design.
- SWE-agent documentation: a concrete configurable agent environment and execution loop.
- Inspect documentation: structured tasks, scorers and sandboxed evaluation.
- OpenTelemetry traces: spans, relationships and distributed execution evidence.