Generative AI Architecture · Part 8 of 9
Evaluating AI Systems
"It looked fine when I tried it" is not a measurement.
Testing has come up in passing all along: an eval set for a prompt, validation at the end of a RAG pipeline, an audit trail for an agent's actions. Evaluation earns treatment as its own discipline, because a system built from a model, retrieved context, and tool calls fails in ways a traditional test suite isn't built to catch. That measurement belongs in the system from the start, before something goes wrong in front of a user.
Offline evaluation vs. online evaluation
Offline evaluation runs a fixed set of test cases against a system before it ships, the same idea as a prompt's eval set scaled up to a whole application: known inputs, checked automatically, run again every time something changes. Online evaluation measures the system against live traffic: sampling interactions, tracking outcomes over time, catching regressions and drift that a fixed offline set, however well constructed, wasn't built to anticipate. Offline evaluation catches a regression before a user ever sees it; online evaluation catches the failure mode nobody thought to write a test case for.
What gets measured
- Task success. Did the system accomplish what the user needed, judged against the task's own definition of done, not just whether it produced a well-formed response.
- Groundedness. For a RAG system specifically, whether the answer is supported by the retrieved sources, rather than by something the model recalled from training or invented outright.
- Hallucination rate. How often the system states something as fact that isn't true or isn't supported by anything it was given, measured explicitly rather than assumed to be rare.
- Retrieval quality. Whether retrieval returned the right chunks in the first place, independent of whether the model used them well — a system can have perfect generation and still fail because retrieval handed it the wrong documents.
- Tool correctness. For an agent, whether it selected the right tool, with the right parameters, and interpreted the result correctly.
- Safety. Whether the system avoids producing harmful, biased, or policy-violating output, including under adversarial input designed to provoke exactly that.
- Cost and latency. What each interaction costs to run and how long a user waits for it, tracked as a first-class metric rather than an afterthought once a bill arrives.
Cost per successful task
Cost per token is an easy number to reach for and a poor one to optimize against on its own. A cheaper model or a shorter prompt lowers cost per call, but if it also lowers the rate at which the task succeeds, the cheaper call ends up costing more once the failures are accounted for: a follow-up call to fix a wrong answer, a person manually correcting an incorrectly filed ticket, a customer who has to ask twice.
Cost per successful task is the better framing: total cost across every attempt at a task, divided by the number of times the task was completed correctly. Optimizing cost-per-token in isolation can make a system look cheaper on a dashboard while making it more expensive to operate.
| Model A | Model B | |
|---|---|---|
| Cost per call | $0.002 | $0.006 |
| Task success rate | 70% | 96% |
| Calls needed per successful task (on average) | 1.43 | 1.04 |
| Cost per successful task | $0.0029 | $0.0062 |
In this illustrative example Model A still comes out cheaper per successful task, which is the point: the framework doesn't automatically favor the pricier model, it just makes the underlying comparison visible instead of stopping at the sticker price. If Model A's success rate were 30% instead of 70%, the ranking would flip: cost per successful task would rise to $0.002 ÷ 0.30 ≈ $0.0067, more expensive than Model B's $0.0062, and cost-per-token alone would never have shown that.
This is the same evaluation discipline the previous article's audit trails feed into and the next article's infrastructure choices get measured against: a system is only as good as what you can show it does, repeatedly, at a cost you've deliberately chosen rather than discovered after the fact.