Core Cloud Architecture · Part 4 of 13
Autoscaling and Elasticity
Adding capacity automatically is easy. Making every layer of the system able to absorb it is the hard part.
Elasticity is the promise that started the article on why cloud computing transformed startups: capacity that grows and shrinks with demand instead of being sized in advance for a guess about peak load. Autoscaling is the mechanism that delivers on it. It's also one of the easiest places to build a system that looks resilient in testing and falls over in production, because the part that scales automatically is rarely the part that breaks first.
Horizontal vs. vertical scaling
Vertical scaling means making an existing instance bigger: more CPU, more memory, on the same machine. It requires no changes to how the application is built, but it has a ceiling (there's a largest instance size available), and doing it live usually means downtime or a restart. Horizontal scaling means adding more instances running the same code, side by side behind a load balancer. It has no meaningful ceiling and each instance can be added or removed without touching the others, but the application has to support running many copies at once: no state stored only in one instance's memory, no assumption that "the" server is a single machine.
Reactive vs. predictive scaling
Reactive scaling responds to what's happening right now (CPU usage crossed a threshold, request queue length grew) and adds or removes capacity after the fact. Predictive scaling uses historical patterns (a known daily traffic curve, a scheduled marketing push) to add capacity ahead of an expected spike, rather than reacting once it's already underway. Reactive scaling is simpler to set up and is the default almost everywhere; predictive scaling earns its complexity once traffic has a predictable shape and the lag in reactive scaling is causing problems.
Request-based vs. queue-based scaling
Request-based scaling reacts to inbound HTTP traffic: requests per second, concurrent connections, CPU load caused by serving them. It's the natural fit for a web application or API. Queue-based scaling reacts to the depth of a work queue instead, how many jobs are waiting to be processed, and fits background or asynchronous work. A system commonly needs both: request-based scaling for the part serving live traffic, queue-based scaling for the part chewing through background jobs.
Cold starts, minimum instances, and scale-to-zero
A serverless function or serverless container that scales to zero instances when idle costs nothing while unused, but the next request has to wait for a fresh instance to start, initialize, and become ready. That delay is a cold start, and it ranges from tens of milliseconds to several seconds depending on the runtime and how much setup code runs before the first request is handled. Setting a minimum instance count above zero keeps that many instances warm at all times, trading a small ongoing cost for eliminating cold starts on the traffic that matters. That's a common choice for a user-facing API, and unnecessary for something that runs a few times a day and can tolerate the delay.
A related, often-overlooked limit is connection limits: every instance holds some number of open connections (to a database, to another service), and a fleet that scales out fast can hit a connection ceiling on the other end well before it hits any limit on itself.
The failure mode: autoscaling the frontend into a wall the backend can't move
Autoscaling adds capacity to whichever layer is configured to scale, and it's easy to configure only the layer that's visibly under load (usually the application servers) without noticing that a downstream dependency has a hard ceiling of its own.
A concrete version of this: a burst of traffic causes the load balancer to trigger autoscaling on the application tier, which happily goes from 4 instances to 40 in under a minute. Each instance opens its own pool of connections to the database. The database has a fixed maximum connection count, a usually modest limit that doesn't scale itself just because the application tier did. Once the connection ceiling is hit, new connection attempts start failing, application instances start timing out waiting for a database connection to free up, and those timeouts cause the load balancer to see slow, failing instances, which can trigger scaling out even further, adding still more instances fighting over the same fixed number of database connections.
flowchart TD
A[Traffic spike] --> B[App tier scales: 4→40 instances]
B --> C[Each instance opens a DB pool]
C --> D{DB max connections reached?}
D -- Yes --> E[Connections rejected]
E --> F[App times out on DB]
F --> G[LB sees failing instances]
D -- No --> H[System handles the spike]
The application tier scaled as designed. The system still went down, because the layer that became the bottleneck, the database's connection ceiling, wasn't part of the scaling story at all.
The same trap extends past databases: any downstream dependency with a fixed or slower-scaling capacity (a third-party API with a rate limit, a cache with fixed memory, another internal service that hasn't been scaled the same way) can become this kind of wall. Autoscaling only helps the layer it's applied to. Reliability in Distributed Systems covers the patterns (backpressure, load shedding, circuit breakers) that keep a bottleneck like this from taking the whole system down with it.