Tutorials › Core Cloud Architecture › Observability: Logs, Metrics, and Traces

Core Cloud Architecture · Part 12 of 13

Observability: Logs, Metrics, and Traces

Logs, metrics, and traces each answer a different question, and an incident usually needs all three.

Building a system is one problem; finding out what it's doing is another, especially once it's made of enough separate pieces (load balancer, several app instances, a cache, a database, a queue) that no single log file tells the whole story.

Logs, metrics, and traces

Logs are discrete, timestamped records of events (a request came in, an error occurred, a background job finished), usually free-text or structured as JSON, and the most granular of the three. Metrics are numeric measurements tracked over time (requests per second, error rate, CPU utilization), cheap to store and query at scale, and well suited to dashboards and alerting, but they tell you that something's wrong, rarely why. Traces follow a single request as it moves across every service it touches, recording how long each hop took, which is what lets you find where in a multi-service call chain the time went, rather than just knowing the total was too slow.

Dashboards, alerts, and SLO/SLA/SLI

A dashboard presents metrics visually, for humans watching or reviewing system health. An alert is the same data used differently: a rule that notifies someone automatically when a metric crosses a threshold, so catching a problem doesn't depend on someone happening to be looking at a dashboard at the right moment.

An SLI (service level indicator) is a specific measured metric, say the percentage of requests served in under 300 milliseconds. An SLO (service level objective) is an internal target for that indicator, say 99.5% of requests under 300ms over a rolling 30 days, used to decide when reliability work should take priority over new features. An SLA (service level agreement) is an external, usually contractual, promise to a customer, typically looser than the internal SLO backing it, so that normal operational variance doesn't routinely breach a customer-facing commitment.

Percentile latency: p50, p95, p99

An average latency number hides the requests you most need to know about. p50 (median) is the latency half of requests beat and half didn't, a reasonable read on the typical experience. p95 is the latency 95% of requests beat, the slowest one in twenty. p99 is the slowest one in a hundred. A system with a fine average and a p50 of 80ms can still have a p99 of 4 seconds, meaning 1% of users, which is a lot of users at scale, are having a bad experience the average never shows. Reliability work usually aims at the tail (p95 and p99) because that's where the average stops telling the truth.

Compare: observability tooling across the three major clouds

ConceptAWSGCPAzure
Logs & metricsCloudWatchCloud Observability (formerly Cloud Logging / Cloud Monitoring)Azure Monitor
Distributed tracingX-RayCloud TraceApplication Insights

Walkthrough: latency increased from 500 milliseconds to six seconds

A dashboard alert fires: p95 latency on the checkout API has climbed from a steady 500ms to over 6 seconds in the last hour. Here's a path to the cause, in the order worth checking.

  1. Check recent deploys. The single highest-probability cause of a sudden regression is a recent change. If a deploy to the checkout service or something it depends on lines up with when latency started climbing, that's the first thing to suspect. Once confirmed, rolling it back and investigating afterward usually beats debugging forward under pressure with the regression still live.
  2. If no recent deploy lines up, check dependency latency via traces. Pull a handful of slow request traces and find which hop grew. Say the traces show the application's own processing time unchanged at 40ms, while a call to the database now takes 5.8 seconds where it used to take 100ms. That isolates the problem to the database without guessing.
  3. Check downstream DB/cache metrics. With the database implicated, its own metrics are next: connection pool utilization, query latency, CPU, active connection count. Suppose connection pool utilization is pinned at 100% and query latency for a specific query shape has spiked, while overall query volume hasn't grown.
  4. Find the bottleneck. A slow query log or the database's own query-level metrics narrow it to one query: one that used to use an index and, after a recent schema change (possibly unrelated to the deploy checked in step 1, and easy to miss because it happened a day earlier), now falls back to a full table scan on a table that's grown substantially since the query was written. That query, now taking seconds instead of milliseconds, holds a connection for the whole time it runs. With a fixed connection pool size, enough slow queries running concurrently exhaust the pool, the same wall Autoscaling and Elasticity describes, and every other request waiting on a connection inherits the delay.

The alert said "checkout is slow." The trace said "the database call is slow." The database's own metrics said "one query is slow and pool utilization is maxed." Each signal narrowed what the previous one left open, until a system-wide symptom came down to one fixable cause: an index that stopped being used.

Observability tooling turns "something is wrong" into "this specific thing, in this specific place, is wrong." That narrowing is most of the work in an incident.