Tutorials › Core Cloud Architecture › Identity and Access Management

Core Cloud Architecture · Part 10 of 13

Identity and Access Management

Two separate questions: who are you, and what are you allowed to do.

Networking decides what can reach what: an app server can reach a database, a function can read from a queue. Identity and access management (IAM) decides whether a given caller is allowed to, and it applies as much to one piece of software calling another as it does to a person logging in.

Authentication vs. authorization

Authentication answers "who are you" by proving an identity, usually via a password, a certificate, a token, or a signed request. Authorization answers "what are you allowed to do," a separate question evaluated after authentication succeeds. A correctly authenticated identity can still be denied a specific action; the two checks are independent, and conflating them is a common source of access bugs, where "logged in" is quietly treated as equivalent to "allowed to do anything."

Identities: human and application

A human identity belongs to a person: an engineer with console access, a customer with an account. An application identity (a service account on GCP, an IAM role on AWS, a managed identity on Azure) belongs to a piece of software: an application server calling a database, a function calling another service. Application identities vastly outnumber human ones in most systems, and deserve at least as much attention to what they're allowed to do, since a compromised application identity can be exploited automatically, at machine speed, with no human in the loop to notice something looks wrong.

Identities: AI agents, a third category

An AI agent identity sits closer to an application identity than a human one, since it's software calling an API, but it raises questions a service account's defaults don't answer. Audit needs to capture more than the API call the agent made: it needs the goal or reasoning context that led to the call, because "the agent deleted this record" is a much less useful log line than "the agent deleted this record because it interpreted this instruction as authorizing cleanup of stale entries." Scoping needs to be tighter and shorter-lived than a standing service-account role, because an agent's job may exist only for the duration of one task; a session- or task-scoped credential fits that shape. Revocation needs to isolate one agent's authority: killing a single misbehaving agent's ability to act shouldn't require disabling a service account that other, unrelated things share. AI Agents in the Generative AI Architecture series covers the authorization side in more depth.

Roles, policies, and least privilege

A role is a named bundle of permissions, assigned to an identity, rather than permissions being granted one at a time to each identity individually. A policy is the document (usually JSON or a provider-specific equivalent) that spells out which actions are allowed on which resources, under which conditions. Role-based access control (RBAC) is the pattern of defining roles around job functions or application responsibilities ("read-only analyst," "checkout service") and assigning identities to roles, rather than managing permissions per-identity, which becomes unmanageable past a handful of people or services.

Least privilege is the governing principle underneath all of it: grant the permissions an identity needs to do its job, and no more. A service account that only ever reads from one storage bucket should have permission to read that one bucket, not read-write access to every bucket in the project. Least privilege sets how much damage a single compromised credential can do: a stolen credential with narrow permissions can only misuse those narrow permissions, while a stolen credential with broad ones can misuse all of them.

Workload identity

Workload identity lets a piece of software running in the cloud (a container, a function, a VM) obtain short-lived credentials automatically from the platform it's running on, tied to its own identity, instead of being handed a long-lived secret to authenticate with. This is what a Cloud Run service, a Lambda function, or an Azure Function's managed identity is doing under the hood: the platform vouches for the workload directly, based on where it's running, rather than the workload needing to prove itself with a static credential someone had to create and distribute.

Compare: identity services across the three major clouds

ConceptAWSGCPAzure
Core IAMAWS IAMGoogle Cloud IAMAzure RBAC
Human identity / directoryIAM Identity CenterCloud IdentityMicrosoft Entra ID
Application identityIAM rolesService accountsManaged identities

The problem with long-lived credentials

A long-lived credential (a static API key or password embedded in a config file, an environment variable, or worse, checked directly into source control) is a liability however carefully it was created. It doesn't expire on its own, so if it leaks (a misconfigured repo, a log line that shouldn't have printed it, a laptop that gets compromised), it keeps working until someone notices and manually revokes it, which can be a long time after the leak happened. It's also usually shared across every place that needs the same access, so revoking it to contain one leak breaks everything else legitimately using it, forcing an uncomfortable choice between leaving a known-compromised credential live or breaking production.

Short-lived credentials and managed/workload identity mechanisms fix both problems at once: a credential that expires in minutes limits how long a leaked copy is useful, and because it's issued per-workload rather than shared, revoking or rotating one doesn't require touching every other consumer. Manual rotation was tedious and easy to put off, which is what made long-lived keys tempting in the first place. Workload identity removes that work, and has become the default recommendation instead of an advanced hardening step.

A static key checked into a repository is among the most common credential leaks. It costs nothing to check for: automated secret-scanning tools exist to catch this pattern before or shortly after a commit lands, and running one is far cheaper than the incident that follows a leak.