← All articles

Article · 4 August 2026 · 5 min read

Agent evaluation: the work nobody posts about

By Doryan Gowty, Principal, Anneal

Making claims about what you've built is one thing. When you have code or agents making consequential decisions that affect customers or your bottom line, that's where the rubber hits the road. So of the deep-dives I promised, this is the one I wanted to write first.

Unit testing provides a systematic way of testing what you're about to ship but it can't test ambiguous scenarios or customer interactions effectively. This is where a new evaluation framework is required. The unit tests I had built said nothing about whether the agent would call the right tool, with the right arguments, at the right moment, on a message it hadn't seen before. Which is the only thing that actually matters once a model is making the decisions.

Separately, around the same time, a support policy had to be rewritten twice — because live customer interactions kept surfacing edge cases no one had thought to test for.

The key message here is you can't treat an agent like ordinary software, because it isn't.

Why an agent breaks the normal testing model

Ordinary code is deterministic. Same input, same output, forever — so a test you wrote last year still means something today.

An agent isn't that. It has real permissions — it can issue a refund, send a message, write to a database — and its behaviour changes every time the underlying model changes, every time the prompt changes, every time the policy it reasons over changes. A refactor that touches none of your code can still change what the agent does, because the judgment lives in the model, not the code.

So the update you ship confidently on a green test suite is a regression risk with no safety net. Your tests verified the plumbing. They never checked the judgment.

That's the gap. Everything I build now exists to close it.

Testing judgment, not plumbing

The core move is simple to say and annoying to build: run the actual agent loop against a fixed set of real-shaped scenarios, on every change, and block the deploy if it regresses — exactly the way a failing unit test blocks a normal merge. No mocks on the part that matters. The real model, making real decisions, and getting marked on its homework.

Terminal output from an agent evaluation run: groundedness, policy accuracy, tool arguments, tone and escalation pass; identity scoping fails because the agent returned an order outside the verified customer boundary.
The evaluation framework needs to consider a range of criteria that matter to what the agent is trying to achieve

The interesting part isn't that idea, it's what "graded" means, because an agent fails in more than one way and no single check catches all of them. A few that turned out to matter:

Some failures are exact, and you check them exactly. Did the agent call the refund tool with the right amount? Did the order lookup get scoped to the customer who was actually authenticated — or could it be pointed at someone else's data? That last one isn't hypothetical. Pre-launch testing found the agent's lookup tools weren't constrained to the verified customer — in principle, a conversation could be steered into returning another customer's order. The fix was architectural (identity resolved once, every downstream call constrained to it), but the evaluation for it is a blunt, deterministic assertion: this tool call must be scoped to this identity, every time, no exceptions. When the correct answer is exact, you check it exactly.

Some failures have no exact answer, so you have another model judge them. "Did the agent apply the refund policy correctly and explain it in a reasonable tone" has no single right string to match against. For those, a second model scores the reply against a rubric — groundedness, policy accuracy, tone. It catches the replies that are technically tool-correct but wrong in substance: an ungrounded claim (hallucination), a misapplied exception, an answer that's accurate and unhelpful at the same time. It's not perfect — a model grading a model deserves its own skepticism — but it catches a class of error that exact-match checks structurally can't see.

And some failures only show up in the wild. A fixed test set can only contain the cases you thought of. So a sample of real production interactions gets reviewed too, and the genuinely novel ones get folded back into the fixed set — so the suite gets harder over time, taught by real traffic rather than my imagination.

The honest limits

I don't want you to walk away thinking that this issue is solved. It's a work in progress and will continue to be. The important thing is to design the framework to evolve with the system.

A model judging another model is a real dependency with real failure modes; it reduces the manual review load, it doesn't remove it. A golden dataset is only ever as good as the cases in it — which is exactly why the live-traffic loop matters. And none of this makes the agent correct; it makes regressions visible before a customer finds them, which is a much more modest and much more achievable goal.

Other strategies to reduce the gap

Putting a human in the loop provides another review layer that makes the final adjudication on whether a case was dealt with correctly. It's also a useful audit exercise.

Using a different model as the judge helps to eliminate blind spots. Selecting a reasonable model from another provider is a suitable strategy. Evaluation runs are also costly exercises, it can make sense to put in place a caching strategy and submit the evaluation as batch if this is available from the model API.

That modest goal is the whole difference between "it worked when I tried it" and something I'd let touch real customers and real money. The distance between those two is most of the actual work — and almost none of it is model quality.

The evaluation framework, and the five agent workflows it covers, are written up at anneal.io. This is the first of a few deep-dives into what I built for a small drinks brand over my career break; the platform and the marketing-mix model are next.

This is a continuation on the series from I took a career break. I ended up shipping a production AI system.