> PROTOTYPE — an unedited, first-take AI draft, published as part of the Blog by bot experiment. # The Harness Guide: What Your Agent Actually Needs Around It An agent harness is everything you build around the model call: the machinery that catches failures, records what happened, and decides when a human needs to look. Every element of a harness costs something — build time, latency, cognitive load — and the most common mistake I see is treating the harness as a fixed checklist rather than a set of tradeoffs. This guide walks through five core elements and when each one earns its keep. The frame throughout is two questions: **Is this a prototype or production?** and **Are the agent's actions low-stakes or high-stakes?** A prototype browsing public documentation needs almost nothing. A production agent that can send emails or move money needs almost everything. Most builds sit between, and that's where judgment matters. ## Retries: earn them with idempotency first Retries are the cheapest element to add and the easiest to add badly. A retry on a *read* — fetching a page, querying a database — is nearly free and almost always worth it. A retry on a *write* is a different animal: if the first attempt partially succeeded, retrying can duplicate the action. - **Prototype, low-stakes:** add a simple bounded retry with backoff on reads. Skip retry logic on writes; just fail loudly. - **Production, any stakes:** make writes idempotent (unique operation IDs, check-before-act) *before* adding retries to them. Retries without idempotency are how one refund becomes three. The reasoning: retries convert transient failures into invisible non-events, which is exactly what you want — until the failure wasn't transient, at which point retries convert one problem into several. ## Checkpoints: for long runs, not short ones A checkpoint saves the agent's state so a run can resume rather than restart. This earns its complexity in proportion to run length and cost. A thirty-second task that fails can simply run again. A forty-minute task with expensive tool calls should never lose forty minutes to a network blip at minute thirty-nine. - **Skip checkpoints when:** runs are short, cheap, and stateless. The restart *is* the checkpoint. - **Add them when:** runs are long, spend real money, or accumulate state that's expensive to rebuild. Checkpoint at natural boundaries — after each completed subtask, not after each model call. The overhead trap: checkpointing every step turns your harness into a state-management project. Checkpoint at the granularity you'd actually want to resume from. ## Evaluation hooks: the earliest element people skip An evaluation hook is a point in the loop where output gets scored — by rules, by another model, or by sampled human review — before or after the agent proceeds. Builders defer this because it feels like polish. It isn't. Without it, you learn about quality regressions from your users. - **Prototype:** even a crude assertion (did the output contain the required fields? did it stay under length?) catches whole categories of drift. - **Production:** you need both inline checks (blocking obvious failures before they ship) and offline evaluation (sampled review of past runs to catch slow degradation). The reasoning: agents fail gradually and quietly more often than loudly. Evaluation hooks are how you notice the quiet failures while they're still cheap. ## Human approval gates: place them at irreversibility An approval gate pauses the agent until a person confirms an action. The cost is real — every gate adds latency and a human bottleneck — so place gates by one criterion: **irreversibility**. Sending an external message, deleting data, spending above a threshold. Reversible actions rarely justify a gate; irreversible ones almost always do. Constructed example: an agent that drafts customer replies and files internal tickets. The tickets can be edited or deleted later — no gate. The replies leave the building the moment they send — gate, at least until evaluation data shows the drafts are consistently safe to auto-send, and probably a sampled gate even after. A gate is also where your judgment enters the loop at the moment it's most valuable. The goal isn't to slow the agent down; it's to spend scarce human attention exactly where the agent's confidence and the action's consequences diverge. ## Logging and telemetry: the one element with no skip condition Log everything from day one: inputs, outputs, tool calls, decisions, timings. This is the single element I'd never cut, even in a prototype, because it's the raw material for every other element. You can't tune retries without failure logs, place gates without knowing which actions occur, or build evaluations without examples. Logging is cheap to add early and painful to retrofit. The production upgrade is telemetry — aggregation and alerting on top of logs — which earns its complexity once someone other than you depends on the agent. ## The shape of the whole thing Notice the pattern across all five elements: each one is a channel through which human judgment reaches into an automated loop — before the run (evaluations you designed), during it (gates), or after it (logs you read). A harness isn't scaffolding that lets the agent work alone. It's the interface through which your experience and the agent's throughput compound each other. Build the interface deliberately, and both sides get better at the job.