Generative AI Architecture · Part 9 of 9
GPUs, TPUs, and AI Infrastructure
The hardware every model call ultimately runs on, and how to compare what it costs.
Every model call, a hosted API request or a self-hosted fine-tuned model alike, runs on hardware built for one kind of math: the large matrix multiplications a transformer is made of. The obvious way to compare that hardware's cost, price per hour, is the wrong number to optimize.
Why not a regular CPU
A CPU executes a small number of instruction streams flexibly, which makes it good at the kind of varied, branching logic most software is built from. A transformer's core operation is the opposite: the same simple operation (multiply and add) repeated an enormous number of times, in parallel, across huge grids of numbers. GPUs (graphics processing units, originally built for rendering) and TPUs (tensor processing units, built by Google specifically for this kind of workload) are both designed around doing that one kind of math at massive scale, at the cost of being far less flexible than a CPU for general-purpose code. Newer accelerators, like AWS's Trainium and Inferentia, take specialization further, with Inferentia built for inference and Trainium for training, though Trainium is now used for inference as well.
Training vs. inference
Training is the process that produces a model's weights in the first place: enormous amounts of data pushed through the model repeatedly, with every pass also computing how to adjust every parameter. It's the most hardware-intensive workload in a generative AI system, often run across thousands of accelerators for weeks. Inference is running the already-trained model to produce a single response, several orders of magnitude cheaper per run but performed constantly, for every request, for as long as the system is live. A production AI system spends most of its hardware budget on inference over the model's lifetime, even though training gets most of the public attention.
Memory, batching, quantization, utilization
A few factors determine how efficiently that hardware is used:
- Memory. A model's weights, plus the working memory needed to process a request (the KV cache covered in Understanding Transformers is a major part of this at inference time), all have to fit on the accelerator. A model too large for a single chip's memory has to be split across several, adding communication overhead between them.
- Batching. Processing several requests together, instead of one at a time, uses the hardware's parallelism far more efficiently, since the same matrix operations run across a bigger batch at almost the same cost as running them on one request. The trade-off is latency: a request sometimes waits briefly for a batch to fill before it starts.
- Quantization. Storing and computing with lower-precision numbers (fewer bits per weight) shrinks a model's memory footprint and speeds up computation, usually at a small, carefully measured cost to accuracy. It's one of the main ways a model that wouldn't otherwise fit on affordable hardware becomes practical to run.
- Utilization. The percentage of an accelerator's capacity doing useful work at any given moment. Idle or underused accelerator time is pure waste, since the hardware is billed, or owned, whether or not it's busy.
Why price per GPU-hour is the wrong metric
Comparing hardware options on the advertised price per GPU-hour ignores how differently the same task uses that hour. A cheaper accelerator that takes twice as long to finish a training run, or handles half as many requests per second, can be the more expensive choice once the job is accounted for. Better metrics measure the outcome of the work:
- Cost per completed training run, the total spend across every hour of hardware time a full training run took.
- Tokens per second, a direct throughput measurement of how much useful inference work a given setup produces.
- Cost per million successful requests, the "cost per successful task" idea applied to infrastructure.
| Accelerator A | Accelerator B | |
|---|---|---|
| Advertised price | $2.00 / hour | $5.00 / hour |
| Hours to complete the training run | 500 | 150 |
| Cost per completed training run | $1,000 | $750 |
Accelerator B looks two and a half times as expensive by the hour. It's the cheaper choice by the time the training run finishes, because it gets there in less than a third of the time. Tokens per second and cost per million successful requests can reverse a comparison the same way.
Options by cloud
| Cloud | Options |
|---|---|
| AWS | EC2 GPU instances (the P and G families); Trainium (training, and increasingly inference); Inferentia for inference |
| GCP | GPUs on Compute Engine and GKE; Cloud TPU; managed training and serving on top of both through Gemini Enterprise Agent Platform (formerly Vertex AI) |
| Azure | GPU VM families (the NC, ND, and NV series); Azure Machine Learning as the managed layer on top |
The case studies that follow apply all of this, together with the cloud infrastructure fundamentals and the funding context, to three complete companies at three different stages.