Tutorials › Core Cloud Architecture › Storage Fundamentals: Object, Block, and File

Core Cloud Architecture · Part 5 of 13

Storage Fundamentals: Object, Block, and File

Three different answers to "where does the data live," each suited to a different access pattern.

Compute and networking are stateless by design: an instance can disappear and be replaced without losing anything important, as long as the data it was working with lives somewhere durable. That somewhere is storage, and cloud storage comes in three distinct shapes, each matching a different way of reading and writing data.

Object storage

Object storage holds whole files, called objects, each with a unique key, accessed over HTTP rather than through a filesystem. There's no folder hierarchy in the traditional sense (the "folders" you see in a console are a display convention over keys that just happen to contain slashes), and no in-place partial edits: an object is written whole and read whole, or replaced whole. This makes object storage the natural home for images, documents, exported datasets, trained model artifacts, and backups: anything read and written as a complete unit, at large scale, cheaply. Amazon S3, Google Cloud Storage, and Azure Blob Storage are the three major offerings, and they're similar enough in shape that this is one of the more portable pieces of cloud architecture between providers.

Block storage

Block storage presents raw, fixed-size blocks of storage to a single virtual machine. It behaves like an attached hard drive because functionally that's what it is. It's what a VM's operating system and a database's data files sit on, formatted with a normal filesystem, supporting random access and in-place writes the way a local disk does. Block storage is fast, but it's normally attached to one instance at a time, and it's the storage layer underneath most managed databases even when the database service hides that detail from you.

File storage

File storage provides a shared filesystem that multiple instances can mount and access concurrently over a standard network filesystem protocol (NFS is the common one), with the same nested-folder structure a local filesystem has. It's the right fit when several instances need to read and write the same files at once: a shared upload directory, or a legacy application built assuming a POSIX filesystem it can't easily be rewritten around. It's less commonly reached for in a newly built cloud-native application, where object storage plus a database usually covers the same needs more cheaply, but it remains the correct choice for the shared-filesystem and legacy cases it's built for.

ObjectBlockFile
UnitWhole objects, addressed by keyFixed-size blocksFiles in folders
AccessHTTP APIAttached to one VM, as a diskNetwork filesystem, mountable by many
Concurrent accessMany readers/writers, whole-object replaceTypically one VM at a timeMany readers/writers, file-level
Typical useFiles, backups, datasets, model artifactsVM disks, database data filesShared uploads, legacy POSIX workloads
AWSS3EBSEFS
GCPCloud StoragePersistent Disk / HyperdiskFilestore
AzureBlob StorageManaged DisksAzure Files

Durability, availability, and what the difference means in practice

Durability is the probability that data, once stored, is never lost, commonly quoted at "eleven nines" (99.999999999%) for object storage, achieved by transparently replicating every object across multiple devices and often multiple facilities. AWS S3 and Google Cloud Storage both publish this figure as their default behavior. Azure Blob Storage's durability depends on which redundancy tier you choose, from LRS (locally redundant, replicated within a single facility only) up through RA-GZRS (geo-zone-redundant with read access, replicated across separate regions and readable from the secondary). The cheapest tier does not replicate across facilities at all, so check which tier an existing setup is on before assuming the eleven-nines figure applies to it. Availability is a separate number: the probability the data can be reached and read right now, typically several nines lower than durability. Data can be perfectly durable, neither lost nor corrupted, and still be unreachable for the minutes a service disruption lasts.

Lifecycle policies and archival tiers

Object storage in particular supports lifecycle policies: rules that automatically move objects between storage tiers, or delete them, based on age or access pattern. A typical policy moves an object from a standard tier to a cheaper infrequent-access tier after 30 days, then to a much cheaper archival tier (retrieval takes minutes to hours, but storage cost drops sharply) after 90, then deletes it entirely after some retention period if that's appropriate for the data. This turns storage cost management from a manual, easily-forgotten task into a rule that runs itself.

Replication: regional vs. multi-region

By default, most object storage replicates data across multiple availability zones within one region, enough to survive a single data center failure. Multi-region replication copies data into a second, geographically separate region as well, protecting against an entire region becoming unavailable and often reducing read latency for users near the second region. It costs more, in storage and often in data transfer, and, like the multi-region compute question, it's worth reaching for once a concrete requirement calls for it: disaster recovery time objectives, data residency across a specific set of regions, latency for a global user base.

Across all three storage types, object storage is the right default for anything that isn't being queried or transactionally updated. Data that needs to be queried, joined, indexed, and updated transactionally belongs in a database instead.