Core Cloud Architecture · Part 13 of 13
Infrastructure as Code and Delivery
How changes to infrastructure get made, safely and repeatably.
Every piece of architecture (networks, compute, databases, caches, queues) eventually needs to change: a new environment stood up, a resource resized, a service added. The goal is making those changes without each one being a manual, error-prone, unrepeatable one-off.
Infrastructure as code
Infrastructure as code (IaC) means defining infrastructure (networks, compute, databases, permissions) in version-controlled configuration files, applied by a tool that reconciles the real infrastructure to match, rather than by hand-running commands or clicking through a console. This gives infrastructure the same properties good code already has: changes are reviewable before they happen, the current state is always visible in the repository rather than locked in one person's memory of what they clicked, and standing up a new environment means running the same configuration again instead of re-deriving it from scratch.
Terraform and OpenTofu (an open-source fork of Terraform) are cloud-agnostic IaC tools using their own configuration language, capable of managing AWS, GCP, and Azure resources side by side in the same codebase. AWS CloudFormation is AWS's provider-specific equivalent, tightly integrated with AWS but not portable to another cloud. Google Cloud's Infrastructure Manager runs Terraform configurations as a managed service, replacing its older Deployment Manager. Azure Bicep is Azure's own domain-specific language for the same purpose, a more readable layer over Azure's underlying deployment templates. Kubernetes manifests and Helm (a templating and packaging tool for those manifests) are the equivalent idea one layer up the stack, for what runs inside a Kubernetes cluster rather than the cloud infrastructure underneath it.
CI/CD
Continuous integration (CI) automatically builds and tests every code change as it's proposed, catching problems before they merge rather than after they're already live. Continuous delivery/deployment (CD) automatically ships a change that passes CI toward production: delivery stops short of an automatic production release and waits for a manual approval, while deployment goes all the way. Applied to infrastructure as code, the same pipeline that tests and deploys application code can plan and apply infrastructure changes, with the proposed change (a "plan," in Terraform's terms) reviewed before it's applied, the same way a code change is reviewed before it merges.
Immutable infrastructure
Immutable infrastructure means a running instance or container is never modified in place after it's deployed: a change means building a new image and replacing the old instance entirely, instead of patching a live one. This removes an entire category of problem: configuration drift, where a server that's been hand-tweaked over months no longer matches what its own deployment configuration says it should look like, making it unclear which one is true and impossible to reliably reproduce.
Blue/green and canary deployment
A blue/green deployment runs the new version (green) fully alongside the old one (blue) and switches traffic over all at once, keeping the old version standing by for an immediate rollback if something's wrong. A canary deployment routes a small percentage of traffic to the new version first, watching its error rate and latency before gradually increasing that percentage to 100%, catching a bad release while it's only affecting a small slice of users instead of everyone at once. Both answer the same question, whether a release is safe, at different points on a speed-versus-caution trade-off: blue/green is faster to fully switch over and simpler to reason about; canary catches problems earlier, with smaller blast radius, at the cost of a slower rollout and needing good enough metrics to trust the canary's signal.
Rollbacks
A rollback reverts to the previous known-good version once a release is found to be bad. It's only fast and reliable if it was planned for in advance: the old version's image or configuration still available and deployable, database migrations written so they can be reversed or are at least backward-compatible with the previous application version, and a deployment pipeline that can execute a rollback as a normal, tested operation instead of a frantic manual scramble the first time it's needed.
Automate the repeatable before building a platform
Every concept here is worth adopting early, in a small, unglamorous form: infrastructure defined in Terraform or OpenTofu from the start, a CI pipeline that runs tests and applies infrastructure changes after review, deployments that replace instances instead of patching them. That takes one straightforward tool applied consistently. It takes no dedicated platform team, no custom internal deployment tool, and no self-service provisioning portal of the kind a much larger engineering org might eventually justify.
The Generative AI Architecture series and the case studies build on this vocabulary, applying it to concrete systems at different company stages.