Tutorials › Generative AI Architecture › AI Agents in Production

Generative AI Architecture · Part 7 of 9

AI Agents in Production

The reasoning loop is the easy part. Deciding what it's allowed to do is the rest of the job.

This module assumes the loop already exists. Agentic AI builds the reasoning loop from scratch: planning, tool calls, observations feeding back in, all the way to a final answer. What follows starts from a working loop and covers what has to surround it before it's safe to point at production: authorization, auditing, cost, and scale.

Planning, tools, actions, memory, orchestration

An agent's planning is the model deciding, from the current state of a task, what to do next. Tools are the external functions it can request be run on its behalf; actions are those requests actually being carried out. Memory is whatever state persists across steps or across sessions, from the running transcript of the current task to longer-term storage the agent can write to and later retrieve. Orchestration is the surrounding system that ties all of this together: routing between multiple agents, retrying failed tool calls, and enforcing the rules that follow.

The loop, with the piece that isn't the model

flowchart TD
  A[User request] --> B[Model reasoning]
  B --> C[Tool selection]
  C --> D{Authorization}
  D -- allowed --> E[Execution]
  D -- denied --> B
  E --> F[Observation]
  F --> B
  B --> G[Result]
  

Every box in that loop except one is the model, or code that mechanically feeds the model's output back into itself. Authorization is the one box that must not be the model. A model can request that an action happen; whether it's allowed to happen is a decision made by code outside the model, using rules the model doesn't get to set or override. A model that decides for itself whether its own request is acceptable is describing what it already did, which is not authorization.

The prompt asks; code decides. Instructing a model in its system prompt not to delete production data, or not to send an email without confirmation, is a request the model can misread, be talked out of by a cleverly worded input, or generalize incorrectly under an unusual case it wasn't trained on. A rule enforced by code that runs regardless of what the model outputs has none of those failure modes: it either allows the action or it doesn't, independent of how convincingly the model has argued for it.

Designing the authorization boundary

A few concrete practices make that boundary enforceable:

All six practices hold regardless of which orchestration framework or which model is in use. They're the same authorization boundary Identity and Access Management covers for human and service accounts, applied to a caller whose next request nobody can know in advance, because a model is the one deciding what to ask for.