Tutorials › Generative AI Architecture › Evaluating AI Systems

Generative AI Architecture · Part 8 of 9

Evaluating AI Systems

"It looked fine when I tried it" is not a measurement.

Testing has come up in passing all along: an eval set for a prompt, validation at the end of a RAG pipeline, an audit trail for an agent's actions. Evaluation earns treatment as its own discipline, because a system built from a model, retrieved context, and tool calls fails in ways a traditional test suite isn't built to catch. That measurement belongs in the system from the start, before something goes wrong in front of a user.

Offline evaluation vs. online evaluation

Offline evaluation runs a fixed set of test cases against a system before it ships, the same idea as a prompt's eval set scaled up to a whole application: known inputs, checked automatically, run again every time something changes. Online evaluation measures the system against live traffic: sampling interactions, tracking outcomes over time, catching regressions and drift that a fixed offline set, however well constructed, wasn't built to anticipate. Offline evaluation catches a regression before a user ever sees it; online evaluation catches the failure mode nobody thought to write a test case for.

What gets measured

Cost per successful task

Cost per token is an easy number to reach for and a poor one to optimize against on its own. A cheaper model or a shorter prompt lowers cost per call, but if it also lowers the rate at which the task succeeds, the cheaper call ends up costing more once the failures are accounted for: a follow-up call to fix a wrong answer, a person manually correcting an incorrectly filed ticket, a customer who has to ask twice.

Cost per successful task is the better framing: total cost across every attempt at a task, divided by the number of times the task was completed correctly. Optimizing cost-per-token in isolation can make a system look cheaper on a dashboard while making it more expensive to operate.

Model AModel B
Cost per call$0.002$0.006
Task success rate70%96%
Calls needed per successful task (on average)1.431.04
Cost per successful task$0.0029$0.0062

In this illustrative example Model A still comes out cheaper per successful task, which is the point: the framework doesn't automatically favor the pricier model, it just makes the underlying comparison visible instead of stopping at the sticker price. If Model A's success rate were 30% instead of 70%, the ranking would flip: cost per successful task would rise to $0.002 ÷ 0.30 ≈ $0.0067, more expensive than Model B's $0.0062, and cost-per-token alone would never have shown that.

This is the same evaluation discipline the previous article's audit trails feed into and the next article's infrastructure choices get measured against: a system is only as good as what you can show it does, repeatedly, at a cost you've deliberately chosen rather than discovered after the fact.