Core Cloud Architecture · Part 2 of 13
Cloud Networking Fundamentals
Everything between a browser and your application, in the order a packet travels it.
Getting a request from a browser to a running application involves addressing, routing, a boundary between what's reachable from the internet and what isn't, and a handful of components almost every cloud architecture assembles the same way regardless of provider.
Addresses: IP, IPv4, IPv6, and CIDR
Every device on a network needs an address, and an IP address is that address. IPv4 addresses look like 10.0.4.12: four numbers from 0–255, giving about 4.3 billion possible addresses, a total the internet has already outgrown. IPv6 has a vastly larger address space, written in hexadecimal groups like 2001:db8::1. Most cloud infrastructure today is still primarily provisioned with IPv4 internally, with IPv6 support increasingly available alongside it.
CIDR notation (10.0.0.0/16) describes a range of addresses rather than one: the number after the slash is how many of the leading bits are fixed, so a smaller number after the slash means a larger range. A /16 covers about 65,000 addresses; a /24 covers 256. Cloud networks are carved up with this notation: a VPC (virtual private cloud, the isolated network you provision resources inside) gets a CIDR range, and subnets are smaller CIDR ranges carved out of it, typically one or more per availability zone.
Public vs. private, and how they talk to each other
A subnet is usually designated public or private. A public subnet has a route to an internet gateway, meaning resources in it can be reached directly from the internet if a firewall rule allows it. A private subnet has no such route: nothing outside the network can initiate a connection into it. The general pattern is that load balancers and anything that needs to be internet-facing live in public subnets, while application servers and, especially, databases live in private subnets, reachable only from inside the network.
A private subnet's resources still often need to reach out to the internet (to fetch a software update, call a third-party API) without being reachable from it. That's what NAT (network address translation) provides: a NAT gateway lets private-subnet traffic go out and come back on the same connection, without ever exposing those resources to unsolicited inbound traffic.
Routing, firewalls, and security groups
Route tables decide, for a given destination address, which path a packet takes next: to the internet gateway, to a NAT gateway, to another subnet, or to a peered network. Two more layers then decide whether a packet is allowed to continue at all. A firewall (or network ACL) is a coarse rule set typically applied to a subnet: allow or deny traffic on this port, from this range. A security group is a finer-grained rule set attached directly to a specific resource (a VM, a database, a load balancer), and is usually the layer that enforces "only the app servers can talk to the database on its port, and nothing else can."
DNS: translating names into addresses
DNS (Domain Name System) is what turns a domain name like yattas.com into an IP address a computer can connect to. A DNS record can point at a fixed address, or, more commonly for cloud architecture, at a load balancer or CDN endpoint whose underlying addresses can change without the domain name ever needing to. Managed DNS services also handle routing decisions beyond a plain lookup: latency-based or geographic routing to the nearest healthy region, for example.
Load balancers: spreading traffic across many instances
A load balancer sits in front of a fleet of application instances and distributes incoming requests across them, so no single instance is a bottleneck or a single point of failure. It also performs health checks, routing traffic away from an instance that stops responding correctly. Most cloud load balancers work at layer 7 (HTTP-aware, can route by path or header) as well as layer 4 (raw TCP/UDP, faster but blind to the request's contents).
Connecting networks: VPN, peering, private connectivity, transit
A VPN (virtual private network) creates an encrypted tunnel over the public internet between two networks, an office and a cloud VPC for instance. Peering connects two VPCs directly, so resources in each can reach the other privately without traversing the public internet at all. Private connectivity (AWS Direct Connect, GCP Cloud Interconnect, Azure ExpressRoute) is a dedicated physical link between an organization's own data center and the cloud provider's network, used when bandwidth, latency, or the enterprise's own compliance rules make even a VPN insufficient. A transit network (a transit gateway or hub) becomes worth it once there are enough separate VPCs that peering each pair directly would mean managing a connection for every pair. A hub lets every VPC connect once, to the hub.
CDN: caching at the edge
A content delivery network serves cacheable content from edge locations instead of routing every request back to the origin region. This cuts latency for anything that doesn't change per request or per user (images, scripts, video, sometimes entire API responses with a short cache lifetime) and reduces load on the origin infrastructure at the same time.
Equivalent services across the three major clouds
| Concept | AWS | GCP | Azure |
|---|---|---|---|
| Virtual network | VPC | VPC | Virtual Network |
| DNS | Route 53 | Cloud DNS | Azure DNS |
| Load balancing | Elastic Load Balancing (ALB/NLB) | Cloud Load Balancing | Azure Load Balancer / Application Gateway |
| CDN | CloudFront | Cloud CDN | Azure Front Door |
| Private connectivity | Direct Connect | Cloud Interconnect | ExpressRoute |
End to end: what happens when a browser loads your site
Put the pieces together and a single page load looks like this.
sequenceDiagram
participant B as Browser
participant D as DNS
participant E as Edge/CDN
participant L as LB
participant A as App
participant DB as DB
B->>D: Resolve domain
D-->>B: Edge IP
B->>E: HTTPS request
alt Cached at edge
E-->>B: Cached response
else Cache miss
E->>L: Forward to origin
L->>A: Route to healthy instance
A->>DB: Query
DB-->>A: Result
A-->>L: Response
L-->>E: Response
E-->>B: Response (cached if cacheable)
end
Every hop in that diagram is a place latency gets added and a place something can fail: a DNS misconfiguration, an unhealthy instance the load balancer should have routed around, a database slow enough to make the whole app feel slow. Autoscaling covers the load balancer/app-server boundary, and reliability patterns cover what to do when one of these hops is slow or failing outright.