The missing layer in agentic systems
Agents are rarely tested the way other software is. A case for evaluation suites that treat an agent's state machine with the rigour of a compiler test suite.
We have spent decades learning how to make software reliable. We write unit tests, integration tests and regression suites. We gate deployments on them. We do not ship a compiler change without running it against a large body of programs whose correct output we already know.
Then we build AI agents, systems that take multi-step actions with real side effects, and test them by running a few examples in a notebook and deciding they look good. The missing layer in most agentic systems is not a better model or a cleverer prompting framework. It is the testing discipline that every other kind of production software takes for granted.
Why agents are harder to test
Agents resist conventional testing for three reasons. Their outputs are non-deterministic, so the same input can produce different but equally valid results. Their behaviour is path-dependent, so a mistake in step two may only surface as a wrong answer in step seven. And their failures are often silent: an agent that took an unnecessary action, burned ten times the expected tokens or reached the right answer for the wrong reason still looks like it succeeded.
The usual response is to give up on rigour and fall back to vibes: run some examples, eyeball the output, ship. That is not good enough when the agent can send emails, change infrastructure or tell a customer something that is not true.
Treat the state machine as the unit under test
The first step is architectural. If your agent is a free-form loop where a model picks tools until it decides to stop, there is very little structure to test. If it is an explicit state machine, with named nodes, typed state and conditional edges, it becomes testable in the ordinary sense.
Each node is a function from state to state. You can test it in isolation with fixed inputs: does the planning node produce a valid plan for this goal, does the retrieval node return the right documents for this query, does the validation node reject this malformed output. Each edge is a condition. You can test that too: given this state, does the graph route to the node it should.
This is why I prefer frameworks like LangGraph for anything that matters. Not because the graph abstraction is magic, but because it forces the decisions an agent makes into places where you can observe and assert on them.
Assertions on behaviour, not just answers
Testing the final answer is necessary but not sufficient. Two agent runs can produce the same correct answer while one took four steps and the other took forty, or while one stayed inside its permissions and the other did not. So I assert on the trajectory as well as the result.
Useful trajectory assertions are usually simple. The agent called the search tool before answering. It never called the delete tool. It finished within its step budget. It asked for human approval before any action on the irreversible list. It did not repeat the same failing tool call more than twice. Each of these catches a whole family of bugs that answer-only evaluation will never see.
For the answer itself, I split the checks by type. Some properties are exact: the output parses as the expected schema, the cited document IDs exist, the numbers in the answer appear in the sources. Those are ordinary assertions and should be. Other properties are judgement calls, such as whether the answer is complete or whether the tone is right, and for those a model-based grader is reasonable, as long as the grader itself has been checked against human judgement on a sample.
Golden sets and regression suites
The core artefact is a golden dataset: a set of inputs with known-good outcomes and known-bad behaviours to avoid, built from real usage wherever possible. Every bug you find in production becomes a new case. Every edge case a domain expert points out becomes a new case. Over time this becomes the most valuable asset in the project, more durable than any prompt or model choice.
That suite runs on every change: new model version, prompt edit, tool change, retrieval tweak. Because outputs are non-deterministic, each case may run several times and the suite tracks pass rates rather than single pass or fail results. A change that drops a case from always passing to mostly passing is a regression, even if a single run happens to succeed.
Crucially, the suite gates deployment. If evaluation is a dashboard someone looks at occasionally, it will be ignored the week a deadline arrives. If it is a check in the pipeline, it gets respected.
Traces are the debugger
When a case fails, you need to know why. That means tracing every run at the level of individual steps: the state going into each node, the model call and its inputs, every tool call and its result, retries, token counts and latency. With traces, an agent failure becomes an ordinary debugging session. Without them, it is guesswork.
The same traces feed back into the test suite. A production trace that ends in a user complaint is a test case waiting to be written. Over time the loop closes: production behaviour shapes the golden set, and the golden set constrains what can reach production.
The discipline is the product
Model capability will keep improving, and the temptation will be to assume the next model fixes whatever is broken. Sometimes it will. But a better model with no evaluation layer is still a system you cannot reason about, and you will not know whether the upgrade helped or quietly broke the cases you care about.
The teams that ship reliable agents are not the ones with access to secret models. They are the ones who treat their agent like any other critical system: explicit structure, assertions on behaviour, a regression suite that grows with every bug, and traces that make failures explainable. That layer is unglamorous, and it is the part that decides whether an agent belongs in production.