Core Cloud Architecture · Part 1 of 13
Regions, Availability Zones, and Edge Locations
The physical geography a request crosses before it ever reaches your code.
Every cloud provider's map is built from the same three pieces, nested inside each other: a region, made of availability zones, which are themselves collections of one or more physical data centers. Edge locations sit outside that nesting, scattered far more densely, closer to the people making requests. Each piece answers a different question: where data legally lives, what keeps a single failure from taking down the whole application, how to make a page load fast for someone on another continent. Mixing them up leads to designs that solve the wrong problem.
Region: a self-contained geographic footprint
A region is a large geographic area, usually named after a place like us-east-1 or europe-west1, that a provider treats as one independent deployment target. Each region has its own power, cooling, networking, and staff, and by design a failure in one region should have no effect on any other. Most cloud resources are created inside a specific region, which makes the region choice the first decision an architecture makes, before compute, storage, or anything else.
Regions matter for three reasons that rarely point the same direction. Latency: a region near your users means shorter round trips. Data residency and sovereignty: some laws (parts of the EU's GDPR regime, certain government contracts, some financial and healthcare regulation) require that specific categories of data physically stay within a country or economic bloc, which only a region choice can guarantee. Availability and disaster recovery: a second region gives you somewhere to fail over to if an entire region goes down, which does happen, rarely but not never.
Availability zone: the fault domain inside a region
A region is itself divided into availability zones (AZs), commonly three or more per region on AWS and GCP, each meant to be a physically separate facility (or cluster of facilities) with independent power, cooling, and networking, connected to the other AZs in the region by low-latency, high-bandwidth links. An AZ is the smallest unit a cloud provider guarantees as a fault domain: a boundary such that a failure inside it (a power outage, a cooling failure, a fire, a rack failing) is contained and shouldn't propagate to another AZ in the same region.
This is the practical reason multi-AZ deployment exists: running identical infrastructure in two or three AZs within the same region means a single facility going dark doesn't take the application down, while round trips between AZs stay fast enough (usually low single-digit milliseconds) that most applications don't notice the difference from running in one.
Edge locations: closer to the user
Providers also operate edge locations (sometimes called points of presence), numbering in the hundreds, far more numerous and far smaller than regions. An edge location doesn't run your application; it caches static content (images, scripts, video, API responses that don't change per request) and terminates network connections physically closer to the end user, so a request from Manila to an application hosted in Virginia doesn't have to make the entire round trip for content that hasn't changed. This is the mechanism behind a CDN.
flowchart TD R[Region — e.g. us-east-1] --> AZ1[Availability Zone A] R --> AZ2[Availability Zone B] R --> AZ3[Availability Zone C] AZ1 --> DC1[Data center] AZ2 --> DC2[Data center] AZ3 --> DC3[Data center] E[Edge locations — hundreds, worldwide] -.caches static content near.-> U((User)) U -.request.-> R
Regional vs. global services
Not every cloud service lives inside one region. Identity systems, DNS, and content delivery are usually global: one identity policy or one DNS record applies everywhere, with no region selection at all. Compute instances, most databases, and object storage buckets are usually regional: created in one region, replicated elsewhere only if you explicitly set that up. Which category a service falls into decides what a regional outage takes down: a regional database failure is contained to that region, while an identity outage, being global, can take down everything at once.
How much of this a startup needs
The options form a ladder, and the right rung depends on what's at stake if a piece of infrastructure disappears for an hour.
| Approach | What it protects against | Cost / complexity | Typical stage |
|---|---|---|---|
| Single AZ | Nothing beyond instance-level failure | Lowest | Prototype, internal tool, pre-PMF |
| Multi-AZ, single region | Data center-level failure (power, cooling, fire) | Modest: mostly configuration, some redundant infrastructure cost | Seed onward, once there are paying customers |
| Multi-region, active-passive | Entire region outage, with some recovery time | Meaningfully higher: a second environment to build, test, and keep in sync | Series A/B and up, or earlier if a specific compliance requirement demands it |
| Multi-region, active-active | Entire region outage, with near-zero recovery time; also lowest possible latency per region | Highest: conflict resolution, data replication design, doubled operational surface | Late-stage/enterprise, or a specific regulatory data-residency requirement |
Why Cloud Computing Transformed Startups makes the same point about managed services generally: the sophisticated-sounding option has to be paid for, in money and in the attention of a small team with better things to do early on. Multi-AZ within a single region is the usual early default; multi-region stays a later decision, made for a named reason.