Tutorials › Generative AI Architecture › Foundation Models and How They're Served

Generative AI Architecture · Part 1 of 9

Foundation Models and How They're Served

The same model, three different ways to run it.

A foundation model is a model trained once, on a broad enough slice of text, code, images, or audio that it becomes useful for many downstream tasks without retraining from scratch for each one. Understanding Transformers covers the mechanics of how one of these models learns and generates text. The question here starts from a working model: how do you put one into an application, and which way of doing it fits the system you're building?

Tokens and context windows

A model doesn't read raw characters. Text is first broken into tokens (roughly word pieces), and the model works entirely in terms of token sequences. Its context window is the maximum number of tokens it can consider at once: the system prompt, the conversation so far, any retrieved documents, and the new response all have to fit inside that same budget. A model that handles a 200-word question comfortably can still fail on a 40-page document stuffed into the same window, either by truncating it, losing track of details buried in the middle, or running past the limit and erroring out.

Context windows have grown from a few thousand tokens to a million or more in several models, but a larger window carries a price. Cost and latency both scale with how much text the model has to process on every call, so an application that reflexively stuffs everything available into the prompt pays for it twice: in the bill, and in how long the user waits.

Inference, temperature, and sampling

Inference is running a trained model to produce output, as distinct from training, which is how it learned in the first place. Every API call to a model is an inference request. At each step the model computes a probability across the entire vocabulary, and something has to decide which token gets chosen from that distribution.

Temperature is the knob that controls that choice. Near zero, the model almost always takes the single most likely next token, giving consistent, repeatable output. Turned up, it samples more freely from less likely tokens, producing output that varies more from run to run and reads as more exploratory. Low temperature suits tasks with a single correct answer, like extracting a date from a document. Higher temperature suits open-ended generation where variety is a feature, like brainstorming names.

Multimodality and reasoning models

Modern foundation models are increasingly multimodal: the same model can accept and sometimes produce images, audio, or video alongside text, by representing all of them as sequences the same underlying architecture can process. A model that can look at a screenshot and describe a bug, or listen to an audio clip and transcribe it, is doing that through the same next-token mechanism as a text-only model, with a wider range of inputs mapped into the same numerical space.

A newer category, often called reasoning models, is trained and prompted to generate an extended internal chain of intermediate steps before producing its final answer. That extra step trades latency and cost for improved accuracy on tasks that need several steps of reasoning, like debugging a piece of logic or planning a multi-step task. A lookup or classification task doesn't need it, and paying for the extra reasoning tokens on every call adds cost without adding correctness.

Where embeddings fit here. The same architecture that predicts text also produces embeddings: numerical representations of meaning. A single foundation model family often ships a generation endpoint and a separate embedding endpoint. One produces text; the other produces something you search over.

Three ways to put a model into an application

Once a model exists, there are three ways to make it part of a running system. They trade off cost, latency, control, privacy, customization, and operational burden against each other:

Hosted model APIManaged model platformSelf-hosted model
What it isCall a vendor's public endpoint for a model you don't manageDeploy, version, and sometimes fine-tune models on a cloud vendor's infrastructure, with more control over the deploymentRun open-weight model files on compute you provision and operate yourself
CostPay per token, no idle costPay per token or per provisioned capacity, depending on tierPay for the hardware whether or not it's busy
LatencyDepends on vendor's shared infrastructure and your network hop to itSimilar, with options like provisioned throughput to reduce varianceFully under your control, including running it close to the rest of your system
Control & customizationLimited to what the API exposes (system prompts, some parameters)Model selection, fine-tuning pipelines, private networking, more configurationFull control over weights, serving stack, and hardware
PrivacyData leaves your network to the vendor's endpointSame, though often with contractual and networking controlsData never has to leave your own infrastructure
Operational burdenEffectively noneSome: deployments, versions, quotasSubstantial: capacity planning, scaling, upgrades, GPU failures

Most teams start with a hosted API because it has no infrastructure to build. The move to a managed platform or a self-hosted model usually comes from a specific pressure: a privacy or data-residency requirement the hosted API can't satisfy, a cost curve that stops making sense at volume, a latency requirement the shared endpoint can't reliably hit, or a need to fine-tune on proprietary data in a way the vendor's API doesn't expose. It's rarely worth adopting the operational burden of self-hosting before one of those pressures shows up; Why AI-First Startups Attract Capital applies the same rule to architecture generally: build for the milestone in front of you. If self-hosting is the right call, Compute Options covers the infrastructure decisions (VMs, containers, Kubernetes) a self-hosted model deployment sits on top of.

The same three choices, per cloud

Each major cloud vendor offers something in both the hosted-API and managed-platform categories, under its own naming:

CloudHosted API / managed platform
AWSAmazon Bedrock (hosted access to multiple model providers through one API) and Amazon SageMaker AI (a broader platform for training, fine-tuning, and deploying models, including self-hosted ones)
GCPThe Gemini API (direct hosted access to Google's models) and Gemini Enterprise Agent Platform, formerly Vertex AI (the broader platform layer: fine-tuning, evaluation, deployment). Its Model Garden catalogs Google, partner, and open models you can call as a managed API or deploy yourself
AzureMicrosoft Foundry (formerly Azure AI Foundry), which brings together what used to be marketed separately as Azure OpenAI Service, alongside other model providers, under one platform