← All articles

Article · 21 September 2026 · 6 min read

Where the LLM actually belongs in an AI agent

By Doryan Gowty, Principal, Anneal

If you've been thinking about how to leverage AI at work you've probably thought about a situation where the model was in the driver's seat. It reads the situation, decides what to do, and gets on with it.

The agents worth running aren't built that way. In reality, the model is one component, and usually not the one making the decisions.

I have published a series of articles on work I did for a small DTC eCommerce business over the last couple of months. This is the final article on the stack I built, and it's for anyone weighing up how they could deploy agents in their business.

What an agent actually is

Agent architecture: an inbound message is routed by code to the language model, which reads a team-edited playbook and calls a fixed menu of tools. Tools send the reply, read and write business systems, or hold money and customer sends for team approval.
The model reads and chooses. Everything else is code, documents and people.

An LLM's job is narrower than most people assume. It reads something messy, like a customer email or a note from the team, and picks from a menu of actions someone else wrote. Ordinary software carries those actions out. Engineers call that surrounding software the harness, and in a well-built agent it's most of the system.

So the question to ask is: which parts should the model be doing at all?

When to use an LLM

The rule I worked to is simple: if a step has a deterministic outcome, it's code. If there's a right answer, a program should give it, the same way every time. Cost points the same way: every model call takes time and money, and a script checking a status field doesn't.

What's left for the model is what code is bad at: understanding people and writing back. Try handling "my order still hasn't turned up and it's been ages, also can I change the address for next time" with rules. A model does it easily.

Example refund conversation: a customer reports dented cans; the agent classifies the email, authenticates the customer, scopes the order lookup to that customer, replies, and issues a $6.40 partial refund.
An example model/tool flow from the DTC case study on anneal.io

So the model interprets and drafts, and code decides, checks and acts. Plenty of what gets called "agentic" barely needs a model. In the above example the model is doing very little: classify a request, respond to customer, that's really it.

Put the limits in the tools

Start with who can talk to the agent. Anyone can send it an email, so every email is text a stranger wrote, and some strangers will write instructions aimed at the model. Security people call this prompt injection. The risk is highest when one agent can see private data, read outside content and send messages out, and a customer-service agent does all three.

The instinctive way to control a model is to tell it what not to do. That isn't a control, it's a hope.

The real limits go in the tools. If the agent can create an order, the tool fixes the product and price; the model supplies only a name and address. It can't offer a price it was never given. Identity works the same way: code does the check, and the model never sees the stored details. A well-crafted email can talk a model into almost anything, but it can't talk a tool into doing something it wasn't built to do.

Most of the work is plumbing

An agent is only as good as the data it can look up. My most important reliability fix had nothing to do with AI. An integration silently stopped sending orders, and it took weeks to notice. Its replacement is a boring job that checks for new orders every 15 minutes. Nobody demos that, but without it the agent confidently tells customers wrong things.

Test it, and test the tester

I've already written about testing for this platform, unit tests for the code and an evaluation process for the agent. Running the agent evaluation on this system revealed a lot about where, and how, the handoff happened between the LLM and the harness.

Before any change reaches a customer, the agent runs through test scenarios scored by a second model from a different vendor. That caught a situation where the agent was guessing between two suppliers when it should have asked.

Testing needs to be part of your operating rhythm. Before going live, this particular agent ran in shadow mode, drafting replies into the team chat without sending anything. Now the same judge scores a sample of live conversations every day, and its scores are being checked against human reviews.

Where the flexibility lives

Building LLMs into the workflow provides a flexible way to parse plain text instructions and feed them into a structured workflow. This is where their power lies. The rules the agent works to live in plain documents, and the team never has to open them.

This gives the system the flexibility to quickly adapt as edge cases and exceptions appear. I use an escalation workflow to help manage this. The model is told 'when in doubt, ask' and it can present a suggestion on how the playbook can be updated to address this. This gives the team the option to either consent to the change or suggest a more appropriate rule. The team teaches the agent in the chat they already use. It's also where the model's output comes closest to changing how the business runs, so the change is narrow: only approved team members can trigger it, it touches one entry, and it can be rolled back.

This is also the weakest point in the system. Nobody reviews a new rule between a team member's reply and it going live. The chain starts with a customer's email, so in principle a stranger's words could end up shaping a rule the agent then applies to everyone. However, controls are in place, the agent drafts a reply someone needs to confirm before being sent.

Software used to sit behind the counter. Now it reads customer emails and writes back, which for a small team means faster answers and fewer things slipping through. It also means a mistake can reach a customer before anyone notices. So for every step I asked what the worst likely error looked like, and who would catch it. Models get things wrong, and so do people, but people usually know when they're unsure. A model has to be built to ask.

If you're thinking about introducing agents in your business

Don't start with which model it uses. Ask:

If the answers are vague, the model is probably doing more than it should.

That wraps up the series. The whole engagement is still there to walk through at anneal.io. If you're working out where an LLM belongs in your own business, I'd like to hear how.