Tutorials › Generative AI Architecture › GPUs, TPUs, and AI Infrastructure

Generative AI Architecture · Part 9 of 9

GPUs, TPUs, and AI Infrastructure

The hardware every model call ultimately runs on, and how to compare what it costs.

Every model call, a hosted API request or a self-hosted fine-tuned model alike, runs on hardware built for one kind of math: the large matrix multiplications a transformer is made of. The obvious way to compare that hardware's cost, price per hour, is the wrong number to optimize.

Why not a regular CPU

A CPU executes a small number of instruction streams flexibly, which makes it good at the kind of varied, branching logic most software is built from. A transformer's core operation is the opposite: the same simple operation (multiply and add) repeated an enormous number of times, in parallel, across huge grids of numbers. GPUs (graphics processing units, originally built for rendering) and TPUs (tensor processing units, built by Google specifically for this kind of workload) are both designed around doing that one kind of math at massive scale, at the cost of being far less flexible than a CPU for general-purpose code. Newer accelerators, like AWS's Trainium and Inferentia, take specialization further, with Inferentia built for inference and Trainium for training, though Trainium is now used for inference as well.

Training vs. inference

Training is the process that produces a model's weights in the first place: enormous amounts of data pushed through the model repeatedly, with every pass also computing how to adjust every parameter. It's the most hardware-intensive workload in a generative AI system, often run across thousands of accelerators for weeks. Inference is running the already-trained model to produce a single response, several orders of magnitude cheaper per run but performed constantly, for every request, for as long as the system is live. A production AI system spends most of its hardware budget on inference over the model's lifetime, even though training gets most of the public attention.

Memory, batching, quantization, utilization

A few factors determine how efficiently that hardware is used:

Why price per GPU-hour is the wrong metric

Comparing hardware options on the advertised price per GPU-hour ignores how differently the same task uses that hour. A cheaper accelerator that takes twice as long to finish a training run, or handles half as many requests per second, can be the more expensive choice once the job is accounted for. Better metrics measure the outcome of the work:

Accelerator AAccelerator B
Advertised price$2.00 / hour$5.00 / hour
Hours to complete the training run500150
Cost per completed training run$1,000$750

Accelerator B looks two and a half times as expensive by the hour. It's the cheaper choice by the time the training run finishes, because it gets there in less than a third of the time. Tokens per second and cost per million successful requests can reverse a comparison the same way.

Options by cloud

CloudOptions
AWSEC2 GPU instances (the P and G families); Trainium (training, and increasingly inference); Inferentia for inference
GCPGPUs on Compute Engine and GKE; Cloud TPU; managed training and serving on top of both through Gemini Enterprise Agent Platform (formerly Vertex AI)
AzureGPU VM families (the NC, ND, and NV series); Azure Machine Learning as the managed layer on top

The case studies that follow apply all of this, together with the cloud infrastructure fundamentals and the funding context, to three complete companies at three different stages.