Generative AI Architecture · Part 2 of 9
Prompt Engineering as Application Logic
A prompt is code your application depends on.
It's tempting to treat a prompt as a one-off piece of text you type until the output looks right, then paste into the codebase. That habit works for a demo and fails in production, for the same reason a magic number scattered through code fails: nobody wrote down what it's supposed to do, so nobody can tell when a change breaks it. A prompt that decides how your application behaves is application logic. It should be reviewed, versioned, and tested against representative cases before it ships.
The parts of a prompt
Most production prompts are assembled from a few distinct pieces, each doing a different job:
- System instructions set the model's role and constraints for the whole conversation: what it is, what it should never do, what format its answers should take. These are usually set once by the application, not by the user.
- The user prompt is the specific request or question for this turn.
- Few-shot examples are sample input/output pairs included in the prompt to demonstrate the format or style wanted, instead of describing it in the abstract. A model told “extract the date as YYYY-MM-DD” will sometimes still drift; the same instruction backed by two or three worked examples drifts far less.
- Retrieved or injected context is additional information supplied for this specific call (a document, a database row, a previous message) that the model needs but wasn't trained on. Retrieval-augmented generation is built entirely around this piece.
Structured output and prompt templates
When an application needs to parse a model's response programmatically, asking nicely for a particular format and hoping isn't a strategy. Most model APIs support a structured output mode, where you supply a schema (a JSON schema, for instance) and the model's response is constrained to match it. This turns "please respond in JSON" from a request the model might ignore into a guarantee enforced by the serving infrastructure, and it removes a category of bugs where a parser breaks because the model added a stray sentence before the JSON.
A prompt template is the reusable skeleton behind all of this: fixed instructions with placeholders filled in per call, kept in one place in the codebase instead of rebuilt inline at every call site. Small changes to wording can move output quality a lot, and one template is far easier to test and improve than the same instructions copy-pasted a dozen times with minor drift between the copies.
# One template, defined once and filled in per call
SUPPORT_TRIAGE_PROMPT = """
You are a support ticket triager. Given a ticket, output JSON with
a "category" string and an "urgency" field of "low", "medium", or "high".
Ticket: %s
"""
def build_prompt(ticket_text: str) -> str:
return SUPPORT_TRIAGE_PROMPT % ticket_text
Context management
Everything above shares one limited context window. A long conversation, a large retrieved document, and a growing list of few-shot examples all compete for the same token budget, and a prompt that grows unbounded over a session eventually gets truncated, slows down, and costs more per call. Managing context means deciding what stays in the prompt as a conversation grows: summarizing older turns instead of keeping them verbatim, dropping retrieved content once it's no longer relevant to the current question, and trimming few-shot examples to the smallest set that reliably produces the right format.
Testing a prompt like you'd test code
A prompt change that looks like an improvement on the one example you tried it on can just as easily be a regression on ten others you didn't. The fix is a small, concrete eval set: a fixed list of representative inputs, each paired with the properties a correct output must have, checked automatically every time the prompt changes.
For the support-triage prompt above, the eval set checks properties instead of matching expected outputs word for word:
| Input ticket | Property checked |
|---|---|
| “Site is down for all users” | urgency must be "high" |
| “Button color looks slightly off on dark mode” | urgency must not be "high" |
| “Can't reset my password, get an error” | category must be "account", output must be valid JSON matching the schema |
| A ticket with no clear category | output must still be valid JSON with some category value, not an error or an empty response |
Running this set against every prompt or model change turns "does this still work" into a question with a yes-or-no answer, and it catches regressions on the edge cases (the ambiguous ticket, the one with no clear category) that a quick manual check skips, because they're not what someone happens to type in while iterating.