Agent evaluation — the work nobody posts about
By Doryan Gowty, Principal, Anneal
Why the distance between an AI demonstration and a production system is mostly evaluation, permissions and regression risk — not the model.
The gap
An agent demo is easy to make impressive: a good prompt, a few tools, a happy-path walkthrough. An agent that’s safe to run unattended against real customers or real money is a different engineering problem, and almost none of the effort that closes that gap is visible in a demo. It’s evaluation harnesses, permission scoping, and regression testing — the unglamorous work that determines whether an agent can actually be trusted to run without someone watching every action.
Scoped permissions
An agent should hold the minimum set of permissions needed for its task, scoped by identity, not by convenience. That means every action the agent takes should be attributable, every credential it uses should be scoped to what that specific workflow needs, and any action with real-world consequence — sending money, changing a customer record, placing an order — should sit behind an explicit permission boundary the agent cannot expand on its own.
Shadow-mode testing and evaluation gating
Before an agent acts on anything real, it runs in shadow mode: it makes its decisions, but a human or a deterministic system executes the action, and the agent’s choice is logged and scored against what actually happened. That builds the evaluation set an agent should be judged against, rather than anecdotal review of the outputs that happen to look good.
Promotion out of shadow mode is a gate, not a decision made once and forgotten: the agent has to clear a defined accuracy and safety bar on that evaluation set, and every change to its prompt, tools or underlying model has to clear the same bar again before it goes live. That’s regression testing applied to a system whose behaviour isn’t fully deterministic — which is precisely why it can’t be skipped.
Incident review and the release lifecycle
Every agent will eventually do something wrong. The question is whether there’s a process ready for it: an incident review that traces the specific decision back through the logs, a fix that goes back through the evaluation gate rather than straight to production, and a user-acceptance step before a materially changed agent is re-released. Treating an agent’s release lifecycle with the same rigour as any other production system — versioned, tested, rolled back when it needs to be — is what separates a system a risk or compliance function can sign off on from one that only survives because nothing’s gone wrong yet.
Where this comes from
This is the evaluation framework Anneal built for the production AI agent inside the marketing-spend platform for a DTC beverage business: scoped permissions, shadow-mode testing, and gating before the agent could act on spend decisions unattended.
Get in touch if you’re trying to get an agent from demo to production.