Tutorials › Cloud and AI Architecture: Case Studies › Sentrio

Cloud and AI Architecture: Case Studies · Part 3 of 4

Sentrio

The product works. Now it has to survive an enterprise security review.

The company

Sentrio helps security operations center (SOC) teams triage and investigate alerts faster. It ingests signals from a customer's existing logs, security tools, and alert streams, uses AI to correlate related signals into a single incident, summarizes what happened in plain language, and suggests next investigative steps. Analysis that used to take an analyst twenty minutes of manual log correlation now takes a few. The SOC team stays in place; what shrinks is the time each analyst spends on the repetitive parts of triage.

The company has grown to fifty to a hundred employees, twenty-five to forty of them engineers. It has meaningful annual recurring revenue and hundreds to low thousands of customers, and it's past the product-market-fit question that dominated its first two years: the product works, teams use it, and it's raised a Series A on the strength of that traction. What's changed is who's buying. The security teams that adopted Sentrio first were small, willing to try something new, and empowered to buy it themselves. The enterprise teams asking to buy in now are large, risk-averse, and several approvals away from a purchase order. They arrive with a security questionnaire, a procurement process, and a list of things the product has to do before anyone signs: single sign-on, audit logging, and a commitment about which country their data sits in.

Two constraints bind at once, and both can end the company. A pre-seed company optimizes for how fast it can learn. A Series A company selling to enterprises optimizes for whether a large customer's security team will approve the vendor at all, and whether the system holds up under higher, less predictable traffic than it saw a year ago. Passing security review with a product that falls over during an incident wins nothing, and neither does flawless uptime nobody is allowed to buy.

Usage pattern

Two kinds of traffic, with different shapes.

Alert ingestion is continuous and uneven. Every customer streams logs and alerts in around the clock, at a baseline that scales with the size of their environment. On top of that baseline sit sharp, unpredictable spikes: an active intrusion, a misconfigured logging rule, a noisy deployment. A single customer's volume can jump fiftyfold for a few hours with no warning. Those spikes are usually independent of each other, but not always. When a widely exploited vulnerability lands, every customer spikes on the same afternoon, which is the load case worth designing for.

Analyst-facing queries follow shift patterns. SOC teams work around the clock, so there's no quiet window in aggregate, though any single customer's dashboard traffic concentrates during their own daytime hours. Query load is modest next to ingestion volume, and it's the half that a human is waiting on, so it gets priority when the two compete for the same database.

The asymmetry drives much of the design below. Ingestion has to absorb a burst without dropping data and without slowing the queries an analyst is running mid-investigation.

What's new since pre-seed and seed

The technical work at pre-seed is building something that works. The technical work at Series A is making it survivable by someone else's standards. Three things change.

First, the buyer changes, and the buyer brings requirements that have nothing to do with whether the product is good: federated identity, retained audit trails, a documented answer to where data lives. Second, load stops being predictable. Early customers generate traffic in proportion to their size; enterprise customers generate traffic in proportion to whatever is happening in their environment that day. Third, the team gets large enough that coordination becomes a real cost, which is what turns infrastructure-as-code and a deployment pipeline from good hygiene into a hard requirement.

Requirements

Functional

Non-functional

Compliance

Enterprise customers, particularly in regulated industries, will ask about SOC 2 (a report on Sentrio's own security controls), GDPR (if any customer or their end users are in the EU), and potentially HIPAA, since a healthcare organization's security logs can carry health-adjacent signals from its own environment.

Major risks

Enterprise trust risk is now dominant. A single visible security incident, or failing one large customer's security review, costs more at this stage than a slow month of usage would. The system has to hold up technically and in front of a security team looking for reasons to say no.

Reliability risk has grown alongside the customer base. A SOC team's dependence on this product during active incidents means an outage has a sharper cost than it did when usage was lighter and less time-sensitive. Reliability engineering, covered in Reliability and Distributed Systems, stops being optional at this stage.

Concentration risk appears for the first time. At hundreds of customers with enterprise pricing, a handful of accounts can represent a large share of revenue, and each of them has a security team with standing to demand an architectural change. That gives a single customer leverage over the roadmap that no customer had a year ago.

Architecture: AWS

The requirements above pull in two directions. Enterprise buyers want SSO, audit trails, regional deployment, and an SLA. The traffic pattern wants an ingestion path that absorbs a correlated spike. And the company is now large enough that platform investment pays for itself. "Reach for the most managed option in every category" is still the right instinct, with more places where it needs an argument.

flowchart TB
  U[Analyst's
browser] --> CF[CloudFront + WAF,
S3 static frontend] U --> ALB[ALB] ALB --> ECS[ECS on Fargate:
API] ECS --> BR[Bedrock:
model routing] ECS --> PLAT[Cognito +
enterprise SSO
Secrets Manager
CloudWatch + X-Ray] ECS ~~~ ALERTS[Customer log /
alert feeds] ALERTS --> KDS[Kinesis
Data Streams] KDS --> ECSW[ECS on Fargate:
correlation workers] ECSW --> REDIS[(ElastiCache
Redis)] ECSW --> AUR[(Aurora Postgres
multi-AZ +
read replicas)] ECS --> AUR ECSW --> OS[(OpenSearch:
log search)] ECSW --> S3L[S3: raw
log archive] OS ~~~ S3L AUR -->|export| RS[(Redshift Serverless:
analytics)]

Frontend and API

The frontend is still static, served from S3 behind CloudFront, now with AWS WAF attached, an addition earned once the customer base includes enterprises whose security reviews ask about edge protection directly. The API layer moves from Lambda to ECS on Fargate behind an Application Load Balancer.

Why a managed container platform instead of Kubernetes?

Twenty-five to forty engineers is enough people to start feeling Lambda's per-function fragmentation, and still not enough workload diversity to justify Kubernetes' scheduling flexibility and the operational cost of running EKS well. ECS on Fargate sits in between: full containers and long-lived services, without a Kubernetes control plane and the tooling around it for a team that doesn't need it yet. Compute Options covers this decision framework in general terms. The ADRs below set the condition for revisiting it.

SSO and audit logging

Cognito federates to each enterprise customer's own identity provider over SAML or OIDC, so a customer's security team can provision and de-provision Sentrio access through the same system they already use for every other vendor, instead of Sentrio maintaining a second, separate set of credentials. Every access to customer data is written to an application-level audit log table in Aurora, distinct from CloudTrail's infrastructure-level audit trail: CloudTrail answers "who called which AWS API," the audit log table answers "which analyst viewed which customer's incident," which is the question both Sentrio's own SOC 2 evidence and a customer's own investigation need answered.

Alert ingestion (event-driven)

Customer log and alert feeds arrive continuously, around the clock, with the bursts described in the usage pattern above. Kinesis Data Streams ingests that continuous flow, and a fleet of ECS on Fargate correlation workers consumes from it: enriching each alert, correlating it against related signals, and writing the result to Aurora and to OpenSearch. Kinesis's ordered, replayable stream model fits this better than a plain queue would, since a burst of alerts during an active incident needs to be processed in order and without silently dropping anything under load. See Messaging and Event-Driven Architecture for the general trade-offs between a queue and a stream.

Database and caching

Aurora PostgreSQL is now multi-AZ by default, with one or more read replicas absorbing the read-heavy load from customer dashboards and internal analytics queries, keeping that traffic off the primary that the ingestion pipeline is writing to. ElastiCache for Redis caches hot-path lookups the correlation workers hit repeatedly (recent alert context, rate-limit counters, session data), work that would otherwise mean repeated round trips to the database for the same small set of records. This is the caching pattern covered generally in Caching.

Object storage and log search

Raw ingested logs are archived to S3, which is cheap enough per terabyte to keep years of them. Full-text search across alert and log data now runs on OpenSearch, separate from the RAG vector store: OpenSearch is built for this kind of high-volume, full-text and structured log search, where Aurora's pgvector extension remains the right tool for the narrower job of semantic search over a knowledge base of playbooks and threat-intelligence documents. Two different search problems, kept in two engines.

LLM and model routing

Sentrio calls Bedrock, but no longer against a single model for every request. A routing layer sends complex incident summarization and reasoning to a stronger, more expensive model, and simpler, high-volume tasks (formatting, basic classification) to a faster, cheaper one. Across millions of monthly requests, the price difference between tiers is large enough to pay for the routing logic several times over.

RAG evaluation

A scheduled evaluation pipeline runs a held-out set of real (anonymized) past incidents through the current retrieval-and-summarization pipeline and scores the output against a known-good answer, tracking groundedness and task success over time so a model or prompt change can be checked against a regression before it ships. Production traffic is sampled into that set continuously, which matters here more than a fixed suite would: the attack techniques the model summarizes change month to month, and a test set frozen a year ago stops representing the job. See Evaluating AI Systems.

Analytics, cost monitoring, and disaster recovery

Redshift Serverless now backs both customer-facing incident analytics and internal product and model-performance metrics, fed by a straightforward export pipeline from Aurora and the evaluation pipeline above. At this data volume, running those queries directly against the operational database would put reporting load in competition with the ingestion pipeline. Cost monitoring runs on AWS Cost Explorer and Budgets, with resources tagged by team and, where practical, by tenant, so a cost spike can be traced to its cause rather than showing up only as a bigger total bill.

Sentrio also deploys to more than one AWS region as separate regional deployments rather than one active-active system spanning regions, so a customer requiring EU data residency lands entirely within an EU region. That's regional choice for data residency, distinct from multi-region resilience for a single tenant (see Regions, AZs, and Edge). Disaster recovery within each region is a documented, tested plan: automated Aurora backups, a defined recovery time objective and recovery point objective, and a periodic restore drill, rather than an untested assumption that the backups would restore cleanly.

CI/CD and infrastructure as code

Every environment is defined in Terraform, and deploys run through a CI/CD pipeline (CodePipeline and CodeBuild, or GitHub Actions targeting ECS) rather than a hand-run script. With twenty-five to forty engineers touching this system, undocumented infrastructure and manual deploys stop being a shortcut and start being a source of drift and incidents. Infrastructure as Code and Delivery covers the general reasoning.

Identity, secrets, and monitoring

IAM roles remain scoped per service to least privilege, the granularity an enterprise security review checks for now that Sentrio's customers are SOC teams. Secrets Manager continues to hold credentials and API keys, so nothing sensitive sits in an environment variable an auditor could quote back as a finding. CloudWatch and X-Ray now cover distributed tracing across the ECS services and the Kinesis pipeline, answering "which of these five services is slow" where single-function logs could only say "something is slow."

Architecture: GCP

Same requirements. A couple of decisions land in a different place on GCP than they did on AWS, beyond a change of product name.

flowchart TB
  U[Analyst's
browser] --> CDN[Cloud CDN +
Cloud Armor, Cloud
Storage frontend] U --> RUN[Cloud Run: API] RUN --> VTX[Gemini Enterprise
Agent Platform
(formerly Vertex AI):
model routing] RUN --> PLAT[Identity Platform +
Workforce Identity
Federation
Secret Manager
Cloud Logging /
Monitoring / Trace] RUN ~~~ ALERTS[Customer log /
alert feeds] ALERTS --> PS[Pub/Sub] PS --> RUNW[Cloud Run:
correlation workers,
push subscription] RUNW --> MEM[(Memorystore
Redis)] RUNW --> ADB[(AlloyDB
primary +
read pool)] RUN --> ADB RUNW --> BQ[(BigQuery:
log search +
analytics)] RUNW --> GCS_L[Cloud Storage:
raw log archive] BQ ~~~ GCS_L

The frontend stays on Cloud Storage behind Cloud CDN, now with Cloud Armor in front for the same reason AWS added WAF: enterprise security reviews expect edge protection to be there.

Why Cloud Run still, instead of GKE, even at this scale?

GCP's service model diverges from the AWS design here. On AWS, the correlation-worker fleet runs on ECS on Fargate because Lambda's request-response model fits a continuously-consuming stream processor poorly. On GCP, Pub/Sub supports push subscriptions that invoke a Cloud Run service directly for each message, which means Cloud Run can absorb the correlation-worker role the same way it absorbs the API role, without introducing a second compute platform at all. Where AWS's answer to "more workload diversity" was "add ECS alongside Lambda," GCP's answer is "Cloud Run already covers this." GKE Autopilot remains the right call once a workload needs long-lived, stateful processing that doesn't fit a request- or push-driven model, which isn't the case here yet.

SSO and audit logging

Identity Platform handles customer-facing auth, and Workforce Identity Federation is the specific GCP product that lets an enterprise customer's own workforce identities (from their own SAML or OIDC provider) access scoped resources without Google having to create and manage a separate Google identity for every one of that customer's users. Cloud Audit Logs plus an application-level audit table in AlloyDB give the same infrastructure-level/application-level split as the AWS design.

Alert ingestion (event-driven)

Pub/Sub ingests the continuous stream of customer log and alert feeds, with the correlation workers subscribed via push subscription as described above. Same event-driven shape as Kinesis on AWS, with a different delivery model underneath: Pub/Sub is push-based publish-subscribe, where Kinesis is an ordered, replayable stream a consumer polls, a distinction covered generally in Messaging and Event-Driven Architecture. Sentrio's correlation workload processes each alert independently rather than strictly in order, so Pub/Sub's model costs nothing here.

Database and caching: AlloyDB

Query volume and vector-search performance justify AlloyDB over Cloud SQL at this stage. AlloyDB's columnar engine accelerates the read-heavy analytical queries now coming from customer dashboards and internal reporting, and its optimized vector index outperforms stock pgvector on Cloud SQL at Sentrio's document and query volume for the RAG knowledge base of playbooks and threat intelligence. A read pool handles the dashboard and analytics load, keeping it off the primary the ingestion pipeline writes to, the same role read replicas play in the AWS design. Memorystore for Redis caches the same hot-path lookups ElastiCache handles on AWS.

Log search and analytics: BigQuery does both jobs

A second divergence from the AWS design. On AWS, full-text log search (OpenSearch) and analytics (Redshift) are two separate services solving two separate problems. On GCP, BigQuery does both: it comfortably handles the high-volume, structured and semi-structured querying that log search needs, and it's already the natural home for analytics and reporting. Sentrio's GCP design routes ingested logs and alerts into BigQuery once, and both the log-search use case and the analytics use case query the same tables. One capability maps to two services on AWS and to one on GCP.

LLM and model routing

Gemini Enterprise Agent Platform (formerly Vertex AI) routes between Gemini model tiers, a faster, cheaper tier for high-volume simple tasks and a stronger tier for complex incident reasoning, the same routing instinct as Bedrock's on AWS, reasoned from the platform's own tiered model family.

RAG evaluation, cost monitoring, and disaster recovery

A scheduled evaluation pipeline built on the platform's evaluation tooling runs the same held-out real-incident test set as the AWS design and tracks the same metrics over time. Cloud Billing budgets and resource labels track spend by team and tenant, and Sentrio deploys to separate GCP regions for customers needing specific data residency, the same regional-choice pattern as AWS. Disaster recovery is built on AlloyDB's automated backups and cross-zone durability within each region, with the same recovery-objective and restore-drill discipline as the AWS design.

CI/CD and infrastructure as code

Every environment is defined in Terraform, deployed through Cloud Build or GitHub Actions targeting Cloud Run, matching the discipline the AWS design adopts at this stage for the same reason: a team this size can't safely rely on hand-run deploys.

Identity, secrets, and monitoring

IAM roles remain scoped per Cloud Run service to least privilege. Secret Manager continues to hold credentials. Cloud Logging, Cloud Monitoring, and Cloud Trace now cover distributed tracing across the Cloud Run services and Pub/Sub pipeline.

Architecture: Azure

Same requirements again. Azure shares GCP's instinct to avoid a second compute platform and reaches it through a different mechanism, and it carries one advantage specific to a security product's customer base.

flowchart TB
  U[Analyst's
browser] --> FD[Front Door + WAF,
Blob Storage
frontend] U --> CA[Container Apps:
API] CA --> AOAI[Azure OpenAI:
model routing] CA --> PLAT[Entra ID +
Entra External ID
Key Vault
Azure Monitor /
App Insights] CA ~~~ ALERTS[Customer log /
alert feeds] ALERTS --> EH[Event Hubs] EH --> CAW[Container Apps:
correlation workers,
KEDA-scaled] CAW --> REDIS[(Azure Managed
Redis)] CAW --> PG[(Azure DB for
PostgreSQL Flexible
Server + read replicas)] CA --> PG CAW --> ADX[(Azure Data Explorer:
log search +
analytics)] CAW --> BLOB_L[Blob Storage:
raw log archive] ADX ~~~ BLOB_L

The frontend stays on Blob Storage behind Front Door, whose WAF is a built-in configuration rather than a separate attached resource, one fewer thing to wire together than on the other two clouds.

Why Container Apps still, instead of AKS, even at this scale?

Both AWS and GCP faced this same fork, and each cloud's own service model produced a different answer. AWS added a second compute platform (ECS) because Lambda's request-response shape doesn't fit a continuously-consuming stream processor. GCP kept a single platform because Cloud Run can be invoked directly via Pub/Sub push subscriptions. Azure also keeps a single platform, but through a third mechanism: Container Apps is built on KEDA, which can scale the number of running replicas up and down based on the backlog depth of an Event Hubs consumer group, not just HTTP request volume. The correlation workers still run as a polling consumer process, unlike Cloud Run's push-invoked model, but KEDA handles the scaling decision automatically based on how far behind that consumer is. Three clouds, three mechanisms. GCP and Azure both reach "no second platform needed"; AWS's answer required one. The primitives compose differently.

SSO and audit logging

Entra ID handles workforce-style federation and Entra External ID handles customer-facing identity, the same split as the other two clouds. A meaningful share of enterprise security teams already run Entra ID (or its predecessor, Azure AD) as their own organization's identity provider, and for those customers, federating a vendor's app into an identity system they already administer tends to be a smoother security review than integrating a SAML connector to an unfamiliar third-party identity product. That advantage is specific to this company's customer base and doesn't generalize to every product in every market. Azure Monitor Activity Logs plus an application-level audit table in the Postgres database give the same infrastructure/application split used on the other two clouds.

Alert ingestion (event-driven)

Event Hubs ingests the continuous stream of customer log and alert feeds, Azure's high-throughput streaming ingestion product, playing the same role as Kinesis on AWS. The KEDA-scaled Container Apps workers described above consume from it, correlate and enrich each alert, and write results onward.

Database and caching

Azure Database for PostgreSQL Flexible Server now runs with read replicas absorbing dashboard and analytics load, keeping it off the primary. GCP and Azure diverge here. GCP's answer to needing more throughput at this stage was switching engines entirely, to AlloyDB. Azure has no equivalent tier; Flexible Server's answer to the same pressure is scaling up compute size and adding read replicas on the same engine. A different lever reaching the same place. Azure Managed Redis (which replaces Azure Cache for Redis) caches the same hot-path lookups ElastiCache and Memorystore handle on the other two clouds.

Log search and analytics: Azure Data Explorer

Azure Data Explorer (Kusto/ADX) is the deliberate choice here over Synapse or Fabric. ADX is purpose-built for this workload: high-volume, time-series log and event data, queried with KQL. It's the same engine behind Microsoft's own Sentinel security product, a strong signal for a security-log-heavy workload. Ingested alerts and logs land in ADX once, and both log search and analytics reporting query the same tables, echoing the "one service, two jobs" shape BigQuery plays on GCP, reached from a security-workload starting point. Synapse or Fabric remain the right choice for broader business intelligence reporting if that need grows separately.

LLM and model routing

Azure OpenAI, accessed through Microsoft Foundry (formerly Azure AI Foundry), routes between model tiers the same way Bedrock and Gemini Enterprise Agent Platform do on the other two clouds: a stronger tier for complex incident reasoning, a faster and cheaper tier for high-volume simple tasks.

RAG evaluation, cost monitoring, and disaster recovery

Microsoft Foundry's evaluation tooling runs the same held-out real-incident test set as the other two clouds. Azure Cost Management with budgets and resource tags tracks spend by team and tenant, and Sentrio deploys to separate Azure regions for customers needing specific data residency, the same regional-choice pattern as AWS and GCP. Disaster recovery relies on Flexible Server's automated and geo-redundant backup options, with the same recovery-objective and restore-drill discipline as the other two designs.

CI/CD and infrastructure as code

Environments are defined in Terraform (or Bicep, where a team prefers Azure's native IaC language), deployed through Azure DevOps or GitHub Actions targeting Container Apps, matching the same discipline adopted on the other two clouds at this stage.

Identity, secrets, and monitoring

Managed identities on each Container Apps service remain scoped to least privilege. Key Vault continues to hold secrets. Azure Monitor and Application Insights now cover distributed tracing across the Container Apps services and Event Hubs pipeline.

Capability comparison

CapabilityAWSGCPAzureDecision rationale
DNSRoute 53Cloud DNSAzure DNSStill a thin, interchangeable layer, even at this scale.
CDN + edge protectionCloudFront + WAFCloud CDN + Cloud ArmorFront Door + built-in WAFAll three now attach edge protection, a direct response to enterprise security reviews expecting it.
Load balancingApplication Load Balancer, provisioned directlyBuilt into Cloud Run's ingressBuilt into Container Apps' ingressAWS's design provisions an explicit load-balancer resource; GCP and Azure abstract it into the compute platform. Same job, different amount of visible infrastructure.
Compute (API + workers)ECS on FargateCloud RunContainer AppsAll three still avoid Kubernetes; see the ADR below on why, and how each cloud gets there differently.
KubernetesEKSGKEAKSStill deliberately unused; see the ADR below.
Relational databaseAurora PostgreSQL, multi-AZ + read replicasAlloyDB, primary + read poolAzure DB for PostgreSQL Flexible Server + read replicasGCP switches to a higher-throughput engine at this stage; AWS and Azure scale the same engine instead. Different available levers, both workable.
NoSQL databaseDynamoDBFirestoreCosmos DBNot used. Sentrio's access patterns are covered by relational, vector, and log-native engines.
CacheElastiCache for RedisMemorystore for RedisAzure Managed Redis (replaces Azure Cache for Redis)Newly adopted on all three: the correlation workers re-read the same small set of records often enough that caching them removes a meaningful share of database round trips.
Object storageS3Cloud StorageBlob StorageRaw log and alert archive on all three; close to a true equivalence.
Streaming ingestionKinesis Data StreamsPub/SubEvent HubsThree delivery models: an ordered, replayable stream a consumer polls (Kinesis), push-based pub/sub (Pub/Sub), and a partitioned event stream a KEDA-scaled consumer polls (Event Hubs).
Event bus / fan-outEventBridgePub/Sub (same product as above)Event GridStill not separately needed; each ingestion pipeline has one consumer type, not several independent ones reacting to the same event.
SecretsSecrets ManagerSecret ManagerKey VaultFunctionally equivalent across all three.
IAM and SSOIAM + Cognito with SAML/OIDC federationCloud IAM + Identity Platform + Workforce Identity FederationEntra ID + Entra External IDAll three now federate to a customer's own identity provider; Azure has a specific edge with customers who already run Entra ID themselves.
ObservabilityCloudWatch + X-RayCloud Logging / Monitoring / TraceAzure Monitor + Application InsightsAll three now include distributed tracing across services, not just per-service logs.
Log search + analyticsOpenSearch (search) + Redshift Serverless (analytics), two servicesBigQuery, one service for bothAzure Data Explorer, one service purpose-built for log-shaped data specificallyThree different shapes for the same underlying need.
LLM platformBedrock, with model routing across tiersGemini Enterprise Agent Platform (formerly Vertex AI), with model routing across Gemini tiersAzure OpenAI via Microsoft Foundry (formerly Azure AI Foundry), with model routing across tiersAll three now route by task complexity rather than calling one model for everything.
Vector searchpgvector on Aurorapgvector with ScaNN indexing on AlloyDBpgvector on Azure DB for PostgreSQLA vector extension on the existing relational database still covers the RAG knowledge base; AlloyDB's optimized index is an incremental GCP-specific advantage at this volume.
RAG evaluationCustom scheduled evaluation pipelineGemini Enterprise Agent Platform evaluation toolingMicrosoft Foundry evaluation toolingAll three score against held-out incidents, with production traffic sampled in continuously so the test set tracks what the model actually sees.
GPU / accelerator infrastructureNone provisionedNone provisionedNone provisionedStill fully abstracted by the managed LLM platforms above.

Architecture decision records

Decision

Use a managed container platform (ECS on Fargate, Cloud Run, or Container Apps) instead of Kubernetes, for now.

Why

Each cloud's own primitives (ECS task definitions, Cloud Run's Pub/Sub push integration, Container Apps' KEDA scaling) already cover the workload diversity Sentrio has today, and no workload on any of the three needs Kubernetes-level scheduling control.

Alternatives considered

Trade-off

Less control over fine-grained scheduling and networking than Kubernetes offers. That flexibility has no current use, and the revisit condition below keeps the decision open.

Revisit when

Workload diversity grows enough (more independent services, more varied resource profiles) that a managed container platform's simpler model starts constraining rather than helping.

Decision

Continue with a shared database and row-level, tenant-ID-scoped isolation instead of separate databases per customer.

Why

Hundreds to low thousands of customers make per-tenant databases an operational multiplier the team can't absorb. Row-level isolation enforced at the identity layer, the query layer, and the search engine's own access controls has held up under every enterprise security review so far.

Alternatives considered

Trade-off

A single largest enterprise customer occasionally asks, during due diligence, whether their data is "physically separated." The current answer is no, and that's occasionally a harder sales conversation than a per-tenant model would be.

Revisit when

A large enough contract specifically requires physical isolation as a condition of signing, or the shared-schema model becomes an operational bottleneck independent of any single customer's request.

Decision

Deploy multi-AZ within each region instead of multi-region active-active.

Why

Regional deployment for data residency is now a firm requirement. Resilience against a full regional outage is a different problem, and multi-AZ within a region already covers the far more common failure mode (a single zone or host going down) at a fraction of the operational cost of active-active replication across regions.

Alternatives considered

Trade-off

A full regional outage still takes that region's customers offline until it resolves. That's an accepted risk at this stage, weighed against the reliability gains from multi-AZ, which already address the more likely failure.

Revisit when

An enterprise contract's SLA specifically requires surviving a full regional outage with no downtime.

Decision

Buy a managed model-evaluation tool (each cloud's own AI evaluation offering) instead of building an in-house evaluation framework.

Why

Building a correctness-and-groundedness evaluation harness from scratch is ongoing engineering work with no direct customer-facing payoff; each cloud's managed evaluation tooling already covers the core need (scoring model output against held-out examples, tracking regressions over time) well enough to start from.

Alternatives considered

Trade-off

Less control over exactly how evaluation scoring works than a fully custom pipeline would give. For the current need, that control isn't yet worth its engineering cost.

Revisit when

Evaluation needs grow specific enough (a scoring dimension unique to security-incident correctness, say) that the managed tooling's built-in criteria stop being an adequate fit.

Decision

Ingest customer alerts through a streaming platform (Kinesis, Pub/Sub, or Event Hubs) instead of a simple point-to-point queue.

Why

Alert volume is now continuous rather than triggered by a single user action, and a streaming platform preserves ordering and replay in a way a simple queue doesn't, which matters when a burst of related alerts during an active incident needs to be processed coherently.

Alternatives considered

Trade-off

A streaming platform is a heavier abstraction to operate correctly (consumer groups, checkpointing, partition management) than a plain queue.

Revisit when

Per-customer alert volume becomes uneven enough that noisy tenants starve quieter ones on shared shards. That forces a partitioning or per-tenant stream decision rather than a return to queues.

Decision

Federate authentication to each enterprise customer's own identity provider instead of maintaining Sentrio-managed credentials for every user.

Why

Enterprise customers expect to provision and de-provision vendor access through their own identity system, and a security review specifically checks for this; a Sentrio-managed password store is both a worse customer experience and a larger attack surface to defend.

Alternatives considered

Trade-off

Federation adds integration work per customer's identity provider quirks, and a customer's own IdP outage now affects Sentrio login too.

Revisit when

A customer's identity provider falls outside the SAML and OIDC support already built, or contracts start requiring automated user provisioning (SCIM) on top of federation.

Decision

Route analytics and dashboard queries to read replicas instead of querying the primary database directly.

Why

The ingestion pipeline now writes to the primary continuously; letting dashboard and reporting queries compete with that write load risks slowing down alert correlation during active incidents, the moments when speed matters most.

Alternatives considered

Trade-off

Read replicas introduce replication lag; a dashboard can be seconds behind the true current state. Acceptable for reporting, not used for anything that needs the current write.

Revisit when

Replication lag itself becomes noticeable enough to affect a customer-facing dashboard's usefulness.

Decision

Adopt a dedicated log-search and analytics engine (OpenSearch and Redshift, BigQuery, or Azure Data Explorer) instead of continuing to query the operational database directly.

Why

Log and alert volume at Sentrio's scale means scanning far more data than the operational database is built to search efficiently. Moving the workload out protects the operational database from analytical query load entirely, where read replicas only partly do for structured reporting.

Alternatives considered

Trade-off

A second data store to keep in sync with the source of truth, and a cost line item of its own.

Revisit when

Retention costs in the search cluster outgrow the value of keeping everything hot, which turns into a tiering decision (recent data searchable, older data archived) rather than a return to querying the operational database.

Due diligence questions

Why a managed container platform instead of Kubernetes, given you're already at 25–40 engineers?
Workload shape is the trigger, and headcount isn't. Kubernetes-level scheduling control has no use in anything currently running, and each cloud's own primitives already cover the workload diversity that exists today.

How do you know one customer's alerts never leak into another customer's dashboard?
Layered enforcement, so no single check is trusted alone: tenant ID at the database row level, tenant scoping at the identity-federation layer, and tenant filtering in the log-search engine's own query layer. Enterprise customers' security reviews have tested it from the outside.

What's your recovery time objective if the primary database goes down?
Multi-AZ failover handles most single-instance failures in minutes; a documented, tested recovery point and recovery time objective covers the rarer case of needing a full restore from backup, with a periodic drill to confirm that number stays true rather than becoming stale.

How do you handle a customer whose alert volume suddenly spikes 50x during an active incident?
The streaming ingestion layer (Kinesis, Pub/Sub, or Event Hubs) absorbs the burst without dropping data, and the correlation-worker fleet scales out automatically behind it; the read replicas and cache layer keep that burst from also slowing down every other customer's dashboard at the same time.

Why didn't you build your own model-evaluation pipeline?
Because the managed evaluation tooling on each cloud already answers the core question well enough, and engineering time spent building a custom harness would come directly out of time spent on the product.

What happens if an enterprise customer requires HIPAA-level guarantees you can't currently make?
It becomes a sales conversation, a compliance review, and likely a contractual and architectural change together. A service configuration alone doesn't solve it, and nobody should represent compliance posture to a customer without a qualified compliance professional involved.

How do you keep LLM inference cost from growing faster than revenue?
Model routing sends the bulk of high-volume, simple tasks to a cheaper model tier, reserving the expensive tier for complex reasoning, and cost monitoring is tied to tags so a spend spike can be traced to a specific cause rather than showing up only as a bigger total bill.

What would make you adopt Kubernetes?
A meaningful increase in workload diversity: services with resource or scheduling needs that don't fit the request-driven or event-driven models the current platforms are built around.

How do you know your model-routing layer is saving money without hurting answer quality?
The evaluation pipeline scores both tiers against the same held-out incident set, so a routing change that saves money but measurably drops task-success or groundedness scores gets caught before it ships broadly, not after a customer notices.

What's the biggest technical debt from the seed stage you're still carrying?
The tenant-isolation model is still logical, not physical, and the largest enterprise prospects are the ones most likely to eventually ask for something stronger. It hasn't been a blocker yet, and it's the item most likely to force a structural architectural change rather than an incremental one.

Architecture decisions we would revisit after Series B