How Do You Test an AI Agent? I Open-Sourced a Library for It
agenteval — a zero-dependency LLM-as-judge library for Go tests · built, then open-sourced
The decision in one line: an AI agent’s output is non-deterministic prose, so you can’t
assert.Equalit — but the fix isn’t to LLM-judge everything. Most of what matters (which tools it called, in what order) is deterministic and a plain Go test pins it better. I built a library that spends tokens only on the remainder — claims expressible only in prose — and open-sourced it: agenteval.
A note on this write-up: the decisions and engineering are mine; I used AI to turn my notes into readable prose. I’m a software engineer, not a writer, and I’d rather be upfront about that.
The problem
I had a production conversational-AI agent (my Go SaaS, T&A Flow) and no good way to test it. The reply is worded differently every single run, so assert.Equal(expected, reply) is meaningless. The instinct is to reach for an eval framework that runs everything through an LLM judge — but that’s slow, costs tokens on every check, and is itself non-deterministic. It felt like using a fuzzy tool to test for fuzzy bugs.
The insight: most of it isn’t fuzzy at all
When I looked at what I actually wanted to assert, most of it was deterministic:
- Did the agent call
check_availabilitybeforebook_appointment? - Did it call the pricing tool at all, or invent a number?
- What language / stage did the classifier return?
All of that is a plain Go table test + slices.Contains — faster, free, and more precise than any LLM judge. Don’t LLM-judge what an if can test.
Only a thin slice genuinely needs a judge — claims that live purely in prose:
“Tells the customer a person will call shortly.” · “Recovered gracefully after the first booking attempt failed.”
No equality check can pin those. That’s the only place agenteval spends a token.
What I built
A deliberately narrow Go library: LLM-as-judge assertions for go test. Zero dependencies (enforced — CI fails if go.sum is non-empty). No bundled provider — the only seam is a one-method Model interface, so you drop in a ~15-line adapter over whatever LLM client you already have. You hand it a Turn (the agent’s reply + its ordered tool-call trace) and a prose criterion; it returns a pass/fail verdict with the judge’s reasoning.
The interesting decisions are the ones that make an LLM judge trustworthy:
1 · Errors are not failures. This is the one I care about most. If the judge is unreachable, rate-limited, or returns garbage, agenteval returns a Go error — reported as “judge unavailable”, a distinct colour — never as Pass: false. A flaky provider must never read as an agent regression, and a suite that silently passes when its judge is down is worse than no suite at all. (A missing pass field in the response is an error too, not a false.)
2 · Refute over negate. LLM judges are notably bad at negation — ask “does it not state a price?” and they reason about prices and then flip the verdict. So the API makes you write the positive criterion (“states a specific price”) and call Refute to invert it. A known model weakness, encoded into the API shape.
3 · Pin exactly what the judge sees — a golden-file firewall. The bytes sent to the judge (system prompt + rendered turn) are snapshotted in a golden file. Any change to the prompt or the rendering — invisible, API-compatible, SemVer-silent — fails CI. It turns “we quietly reworded the judge prompt” into an acknowledged, reviewable event.
4 · Nothing pins the model’s weights. go.mod pins the prompt (module versions are immutable); it pins nothing about the provider — aliases get re-pointed, model IDs get re-served. So the docs prescribe dated snapshot model IDs and a calibration set of hand-labelled turns you run against the judge (not the agent) to catch silent model drift.
Where it earns its keep
I use it to test T&A Flow under fault injection — not happy paths. The eval suite kills the agent’s own DB pool mid-conversation, renames a table so a real tool query fails, and adds a constraint so one write fails while everything else is healthy — all against the live engine, real Postgres, real tools, no mocks. Then the judge checks the agent recovered and said the right thing. That’s the class of behaviour you can’t test any other way.
The honest costs
- A judge call costs tokens and adds latency, so judged tests sit behind a
//go:build evaltag —go test ./...stays free and fast; you runmake evaldeliberately. - The judge is non-deterministic too — mitigated with majority voting over samples and the calibration set, but it’s a real caveat, not magic.
- It’s narrow on purpose. It does not replay scenarios, mock your agent’s LLM, or enforce guardrails at runtime. Earlier versions had a bundled provider and a deterministic-assertion layer; I deleted both — plain Go already tests the deterministic part, and bundling a provider is a dependency I refuse to own.