Tutorials › Core Cloud Architecture › Regions, Availability Zones, and Edge Locations

Core Cloud Architecture · Part 1 of 13

Regions, Availability Zones, and Edge Locations

The physical geography a request crosses before it ever reaches your code.

Every cloud provider's map is built from the same three pieces, nested inside each other: a region, made of availability zones, which are themselves collections of one or more physical data centers. Edge locations sit outside that nesting, scattered far more densely, closer to the people making requests. Each piece answers a different question: where data legally lives, what keeps a single failure from taking down the whole application, how to make a page load fast for someone on another continent. Mixing them up leads to designs that solve the wrong problem.

Shared vocabulary, provider-specific guarantees. AWS, GCP, and Azure all describe their footprint using region/AZ/edge language, but what that language guarantees underneath differs by provider, and sometimes by specific region. Azure has regions with no availability zones at all, relying on a paired second region for disaster recovery instead of in-region AZs. GCP defines a zone more strictly than AWS's looser "one or more data centers" AZ definition, but GCP itself names current exceptions (Stockholm, Mexico, Osaka, and Montreal, among others) where multiple zones share a single building. Edge network architecture differs structurally too: CloudFront, Google's private backbone, and Azure Front Door are three different designs. Before an architecture assumes three independent AZs in a given region, check that the region provides them.

Region: a self-contained geographic footprint

A region is a large geographic area, usually named after a place like us-east-1 or europe-west1, that a provider treats as one independent deployment target. Each region has its own power, cooling, networking, and staff, and by design a failure in one region should have no effect on any other. Most cloud resources are created inside a specific region, which makes the region choice the first decision an architecture makes, before compute, storage, or anything else.

Regions matter for three reasons that rarely point the same direction. Latency: a region near your users means shorter round trips. Data residency and sovereignty: some laws (parts of the EU's GDPR regime, certain government contracts, some financial and healthcare regulation) require that specific categories of data physically stay within a country or economic bloc, which only a region choice can guarantee. Availability and disaster recovery: a second region gives you somewhere to fail over to if an entire region goes down, which does happen, rarely but not never.

Availability zone: the fault domain inside a region

A region is itself divided into availability zones (AZs), commonly three or more per region on AWS and GCP, each meant to be a physically separate facility (or cluster of facilities) with independent power, cooling, and networking, connected to the other AZs in the region by low-latency, high-bandwidth links. An AZ is the smallest unit a cloud provider guarantees as a fault domain: a boundary such that a failure inside it (a power outage, a cooling failure, a fire, a rack failing) is contained and shouldn't propagate to another AZ in the same region.

This is the practical reason multi-AZ deployment exists: running identical infrastructure in two or three AZs within the same region means a single facility going dark doesn't take the application down, while round trips between AZs stay fast enough (usually low single-digit milliseconds) that most applications don't notice the difference from running in one.

Edge locations: closer to the user

Providers also operate edge locations (sometimes called points of presence), numbering in the hundreds, far more numerous and far smaller than regions. An edge location doesn't run your application; it caches static content (images, scripts, video, API responses that don't change per request) and terminates network connections physically closer to the end user, so a request from Manila to an application hosted in Virginia doesn't have to make the entire round trip for content that hasn't changed. This is the mechanism behind a CDN.

flowchart TD
  R[Region — e.g. us-east-1] --> AZ1[Availability Zone A]
  R --> AZ2[Availability Zone B]
  R --> AZ3[Availability Zone C]
  AZ1 --> DC1[Data center]
  AZ2 --> DC2[Data center]
  AZ3 --> DC3[Data center]
  E[Edge locations — hundreds, worldwide] -.caches static content near.-> U((User))
  U -.request.-> R
  

Regional vs. global services

Not every cloud service lives inside one region. Identity systems, DNS, and content delivery are usually global: one identity policy or one DNS record applies everywhere, with no region selection at all. Compute instances, most databases, and object storage buckets are usually regional: created in one region, replicated elsewhere only if you explicitly set that up. Which category a service falls into decides what a regional outage takes down: a regional database failure is contained to that region, while an identity outage, being global, can take down everything at once.

How much of this a startup needs

The options form a ladder, and the right rung depends on what's at stake if a piece of infrastructure disappears for an hour.

ApproachWhat it protects againstCost / complexityTypical stage
Single AZNothing beyond instance-level failureLowestPrototype, internal tool, pre-PMF
Multi-AZ, single regionData center-level failure (power, cooling, fire)Modest: mostly configuration, some redundant infrastructure costSeed onward, once there are paying customers
Multi-region, active-passiveEntire region outage, with some recovery timeMeaningfully higher: a second environment to build, test, and keep in syncSeries A/B and up, or earlier if a specific compliance requirement demands it
Multi-region, active-activeEntire region outage, with near-zero recovery time; also lowest possible latency per regionHighest: conflict resolution, data replication design, doubled operational surfaceLate-stage/enterprise, or a specific regulatory data-residency requirement
Multi-region active-active answers a specific requirement. Multi-region reads as the more mature choice, which makes it tempting early. For most pre-seed and seed startups it isn't warranted: it multiplies operational surface area and cost against a rare failure mode (an entire cloud region disappearing), while the company still faces far larger existential risks (finding product-market fit, keeping the lights on) that a second region does nothing to reduce. Reach for it when a concrete business requirement demands it: a data residency law that requires EU customer data to stay in the EU while serving US customers from the US, a contractual uptime SLA a single-region design can't meet, or a customer base large enough that even a rare regional outage is unacceptable. "We might need it eventually" doesn't qualify.

Why Cloud Computing Transformed Startups makes the same point about managed services generally: the sophisticated-sounding option has to be paid for, in money and in the attention of a small team with better things to do early on. Multi-AZ within a single region is the usual early default; multi-region stays a later decision, made for a named reason.