Tutorials › Cloud and AI Architecture: Case Studies › Casewell

Cloud and AI Architecture: Case Studies · Part 2 of 4

Casewell

A three-person engineering team, fewer than a hundred customers, and one architecture built three times over.

Casewell is a pre-seed company with three engineers, and that constraint decides most of what follows. Below is the company, its requirements, and one architecture designed three times over, once per major cloud provider.

Casewell is fictional: a composite typical of its funding stage, modeled on no specific business.

The company

Casewell is an AI research assistant built for small and mid-size law firms. An associate can ask it a question in plain language, something like "find precedent for a landlord withholding a security deposit past the statutory deadline in this jurisdiction," and it returns relevant case law, cross-referenced against the firm's own past filings and internal memos, with citations back to both. The pitch to a firm is simple: the research that used to take a junior associate an afternoon now takes minutes, and it draws on the firm's own institutional memory instead of only public sources.

The company has five to ten employees, three of them engineers. It has fewer than a hundred paying law-firm customers, all small or mid-size firms rather than the largest national practices. It's priced as SaaS, either per seat (one price per associate with access) or per firm (a flat fee scaled loosely to headcount), and most current customers are still on whichever pricing model let them get in the door fastest; the company hasn't settled on one. It's raised a pre-seed or seed round and is expected to grow. Growth at this stage means firms adopting the product and continuing to use it; a revenue target comes later.

Time is the dominant constraint at this stage. Three engineers can't operate a complex platform on top of building the product. Every infrastructure decision below gets weighed against one question first: does this let a three-person team ship and learn faster, or does it just look more sophisticated?

Usage pattern

Traffic isn't steady. Litigation runs on deadlines: filing dates, discovery cutoffs, court calendars. Usage spikes around each one, then drops off. A firm might barely touch the product for two weeks and then have every associate hammering it the two days before a filing is due. Whatever gets built has to absorb that kind of burst without either falling over during it or paying for capacity that sits idle the rest of the time.

Requirements

Functional

Non-functional

Major risks

Product risk dominates. Current-generation models can produce plausible-looking answers. The open question is whether associates and partners trust an AI-sourced answer enough to rely on it, and whether that trust survives the first time the assistant gets something subtly wrong. Legal research has a low tolerance for confident, well-formatted mistakes, and adoption depends on the product earning trust one correct answer at a time.

This is why the two evaluation requirements above are requirements rather than a later investment. A fabricated case citation in legal work is not a quality issue to be improved on next quarter; lawyers have been sanctioned for filing one. A product that can produce one and has no mechanism to catch it is not shippable to a law firm at any stage, including this one. The bar at three engineers is a cheap, automated bar, not a mature evaluation practice, but it can't be zero.

Confidentiality risk comes second. Attorney-client privilege and work-product protection push a confidentiality failure well past a breach notification: it becomes a professional-responsibility problem for the firm and an existential one for Casewell. Multi-tenant isolation is a condition of the product being usable at all.

Architecture: AWS

These requirements point toward one instinct: reach for the most managed option in every category, even where it costs more per unit than running it yourself, because the team has three engineers and no time to operate infrastructure. Here's the resulting design.

flowchart TB
  U[Associate's
browser] --> CF[CloudFront +
S3 static frontend] U --> APIGW[API Gateway] APIGW --> LAM[Lambda: API] LAM --> BR[Bedrock: LLM] BR -->|answer| VER[Citation verifier] LAM --> SQS[SQS queue] SQS --> LAMW[Lambda:
document worker] LAM --> AUR[(Aurora Postgres
+ pgvector)] LAMW --> AUR LAM --> S3D[S3:
document storage] LAMW --> S3D EV[Lambda: evaluation
runner, scheduled] --> BR EV --> AUR LAM --> PLAT[Cognito
Secrets Manager
CloudWatch] EV --> PLAT

Frontend

A static single-page app, built once and served as static files from S3 behind CloudFront. There's no server-rendering requirement here and no reason to pay for one; the app calls the API layer directly for everything dynamic. CloudFront also gives free TLS and a CDN edge at close to no operational cost, so nobody on the team manages a certificate.

API and compute

The API layer is API Gateway in front of Lambda functions, one function (or small group of functions) per logical operation: authentication callbacks, document upload, question answering, analytics reads. Compute Options covers the general trade-offs; at this traffic pattern, serverless compute is close to a free win.

Why serverless?

Casewell's traffic is bursty and, most of the time, small: quiet stretches between filing deadlines, sharp spikes around them. A fleet of always-on servers sized for the spike sits mostly idle the rest of the time; sized for the average, it falls over during the spike. Lambda scales per request and bills per invocation, so the bursty part of the usage pattern stops being a capacity-planning exercise. The trade-off: cold starts add latency to an occasional request, and a sustained, steady-state high-traffic workload would eventually be cheaper on containers running around the clock. At fewer than a hundred customers with spiky usage, that crossover point isn't close.

LLM and RAG

Amazon Bedrock provides the LLM itself, called from the Lambda API layer with the retrieved context assembled into the prompt. Retrieval draws from two sources, matching the two corpora from the requirements: each firm's own uploaded documents, and a shared public case-law corpus common to every tenant. Both are embedded and stored in Aurora PostgreSQL Serverless v2 with the pgvector extension: one database, two logically separated sets of vectors, tenant ID carried on every row so a query can never cross into another firm's documents. Retrieval-Augmented Generation covers the general shape of this pipeline; the storage choice gets its own ADR below.

Why managed AI APIs instead of self-hosting?

Self-hosting an open-weight model means owning GPU provisioning, model updates, scaling under load, and every failure mode that comes with running inference infrastructure, work that has nothing to do with whether lawyers find the product useful. Bedrock trades a per-token cost premium for all of that disappearing. For a team still trying to find out whether the product works at all, the premium is worth paying. The question this stage needs answered is product-market fit; GPU/TPU AI Infrastructure is a problem worth having later.

Answer evaluation and citation verification

Two mechanisms, at two different points in the system.

The first runs inline, on every answer. Before a response is returned, the API layer parses the citations out of it and checks each one against the set of documents retrieval actually returned for that question. A citation that doesn't match anything in that set never came from the firm's documents or the case-law corpus; the model produced it. Those get stripped, and the answer is returned with a note that a source couldn't be confirmed. This is ordinary application code, not a service, and it costs one pass over the retrieved chunks. It catches the single failure mode with the worst consequences for a law firm.

The second runs on a schedule, in a separate Lambda evaluation runner. It holds a set of research questions with known-correct answers, built with a few cooperative customer firms, and runs the whole pipeline against them: retrieval, prompt assembly, generation, citation check. Each run scores whether the correct authority was retrieved at all, whether the answer's claims trace back to the retrieved text, and how often the citation verifier had to strip something. Results land in Aurora and the run-over-run deltas go to CloudWatch, so a prompt or model change that quietly degrades retrieval quality shows up as a number before a customer notices it. Three engineers can run this; it's a scheduled job and a spreadsheet's worth of test cases. Evaluating AI Systems covers what this practice grows into with more people and more usage behind it.

Database

Aurora PostgreSQL Serverless v2 does double duty: ordinary relational data (firms, users, documents, usage events) and vector search for RAG, in the same database. See Databases: Relational and NoSQL and Vector Databases for the general trade-offs; the specific reasoning for Casewell is one of the ADRs below.

Object storage

Uploaded documents (the original PDFs, not the extracted text) live in S3, one bucket with per-tenant key prefixes and bucket policies that scope access by tenant. This follows the general pattern in Storage Fundamentals.

Async processing and queue

A large PDF filing can run hundreds of pages, and parsing, chunking, and embedding it takes long enough that it can't happen inline during the HTTP request that uploaded it. The upload handler drops a message onto an SQS queue; a separate Lambda worker picks it up, does the processing, and writes the resulting chunks and embeddings into Aurora. If that worker fails partway through, the message becomes visible again after its visibility timeout and gets retried, instead of the upload silently vanishing. See Messaging and Event-Driven Architecture for the general pattern.

Identity, secrets, and monitoring

Cognito handles associate login, with each user scoped to a firm (tenant) at the identity level, so tenant isolation starts at authentication rather than being bolted on afterward in application code. Secrets Manager holds API keys and database credentials; nothing sensitive sits in an environment variable or, worse, in code. CloudWatch covers logs, metrics, and basic alarms: enough for three engineers to know something broke, without standing up a dedicated observability stack. IAM roles are scoped per Lambda function to exactly what that function touches, following the least-privilege pattern in Identity and Access Management.

Why single-region?

A single AWS region, one Aurora writer, no cross-region replication. Multi-region active-active buys protection against a full regional outage, at the cost of a data replication strategy, conflict resolution, and doubled operational surface area. For a company that doesn't yet know if firms will renew, that's cost with no matching benefit. A regional outage is survivable and rare; product risk and confidentiality risk are where the engineering time goes.

Architecture: GCP

This design is re-derived from GCP's own service model against the same requirements, rather than translated service by service from the AWS version above. Most of it lands in a similar place, because the requirements are the same. Where it doesn't, the difference is worth the paragraph that explains it.

flowchart TB
  U[Associate's
browser] --> CDN[Cloud CDN +
Cloud Storage
frontend] U --> RUN[Cloud Run: API] RUN --> VTX[Gemini Enterprise Agent Platform
formerly Vertex AI
Gemini models] VTX -->|answer| VER[Citation verifier] RUN --> TASKS[Cloud Tasks] TASKS --> RUNW[Cloud Run:
document worker] RUN --> SQL[(Cloud SQL Postgres
+ pgvector)] RUNW --> SQL RUN --> GCS_D[Cloud Storage:
documents] RUNW --> GCS_D EV[Cloud Run job:
evaluation, scheduled] --> VTX EV --> SQL RUN --> PLAT[Identity Platform
Secret Manager
Cloud Logging /
Monitoring] EV --> PLAT

The static single-page app is served from Cloud Storage behind Cloud CDN, the same shape as AWS's S3-plus-CloudFront pairing and for the same reason: the frontend is a bundle of static files, and serving it from object storage at the edge removes a server from the picture entirely. Uploaded documents live in Cloud Storage the same way, one bucket with per-tenant object prefixes and IAM conditions scoping access by tenant.

API and compute: Cloud Run

The API layer runs on Cloud Run: one container holding every route for the API. Where the AWS design splits handlers into separate Lambda functions, Cloud Run's unit is a single request-driven container that can serve an entire API's worth of routes, which fits a three-person team's mental model better: one deployable, one image, one thing to reason about, instead of a directory of individually versioned functions. It still scales to zero between the deadline-driven traffic spikes described in the requirements, and it still bills per request, so it gets the same bursty-traffic benefit as Lambda with less code fragmentation.

Why Cloud Run over GKE Autopilot?

GKE Autopilot removes node management from Kubernetes but keeps the Kubernetes API, its resource model, and its operational vocabulary (pods, deployments, ingress objects) in the picture. None of that buys anything here: there's one service, no need for custom scheduling, network policies, or cross-service traffic management that would justify carrying Kubernetes concepts around. Cloud Run gets the same "give it a container, it handles the rest" simplicity as the frontend's static hosting, applied to the API. Compute Options covers this decision generally; for Casewell, Kubernetes is a bet on future operational complexity this team doesn't have yet.

LLM and RAG

Gemini Enterprise Agent Platform (formerly Vertex AI) serves the Gemini model family to the Cloud Run API layer, with retrieved context assembled into the prompt the same way as the AWS design, and the same two-corpus retrieval shape, embedded and stored in Cloud SQL for PostgreSQL with the pgvector extension. The case for a managed model API over self-hosting is the same as Bedrock's, above. The platform also gives an upgrade path if the product later needs a fine-tuned model, without first building the serving infrastructure that would require.

Cloud SQL versus AlloyDB

GCP offers a second managed Postgres option, AlloyDB, built for higher-throughput and analytics-heavy workloads and priced accordingly. At fewer than a hundred customers with a modest, if bursty, query volume, Cloud SQL for PostgreSQL is the right-sized choice: cheaper, simpler to reason about, and adequate for the read and write volume this stage produces. AlloyDB becomes worth revisiting once query volume or vector-search latency under load becomes a bottleneck.

Async processing and queue

Document processing uses Cloud Tasks, a subtle but meaningful difference from the SQS pattern on AWS: Cloud Tasks is push-based, invoking an HTTP endpoint on the worker directly with built-in retry and rate-limiting controls, where SQS is pull-based and the worker polls the queue for messages. For one well-defined job type (process this uploaded document) with one consumer, Cloud Tasks' push model needs less plumbing than a Pub/Sub topic and subscription would for a point-to-point job queue. Pub/Sub earns its place once more than one kind of consumer reacts to the same event.

Answer evaluation and citation verification

The inline citation check is application code inside the Cloud Run API container, unchanged in substance from the AWS version. The scheduled half runs as a Cloud Run job triggered by Cloud Scheduler, which is a closer fit than a second always-on service: it's a batch task with a start and an end, and Cloud Run jobs are built for exactly that, where Cloud Run services are built to serve requests. Scores land in Cloud SQL and the run-over-run deltas in Cloud Monitoring.

Identity, secrets, and monitoring

Identity Platform handles associate login, with custom claims carrying the tenant (firm) association the same way Cognito's user pool groups do on AWS. Secret Manager holds credentials and API keys. Cloud Logging and Cloud Monitoring give the same coverage as CloudWatch, with IAM roles on each Cloud Run service scoped to only what that service needs.

Single-region, for the same reason as AWS: one Cloud SQL primary, no cross-region replication, a regional outage being a bad day rather than an existential one at this stage.

Architecture: Azure

Same requirements, same constraints. Here's what they produce on Azure.

flowchart TB
  U[Associate's
browser] --> FD[Front Door + CDN,
Blob Storage
frontend] U --> CA[Container Apps: API] CA --> AOAI[Azure OpenAI] AOAI -->|answer| VER[Citation verifier] CA --> SB[Service Bus queue] SB --> CAW[Container Apps:
document worker] CA --> PG[(Azure DB for PostgreSQL
+ pgvector)] CAW --> PG CA --> BLOB_D[Blob Storage:
documents] CAW --> BLOB_D EV[Container Apps job:
evaluation, scheduled] --> AOAI EV --> PG CA --> PLAT[Entra External ID
Key Vault
Azure Monitor /
App Insights] EV --> PLAT

The static single-page app is served from Blob Storage's static website hosting behind Azure Front Door, which handles both the CDN edge and TLS termination in one resource, landing in the same place as the pairings on the other two clouds. Uploaded documents live in Blob Storage, one container with per-tenant path prefixes and access scoped by tenant through Azure RBAC conditions.

API and compute: Container Apps

The API layer runs on Azure Container Apps: a single containerized service holding the API, scaling on request volume the same way Cloud Run does. Container Apps is built on KEDA underneath, so it can scale a service on triggers other than raw HTTP concurrency, including the depth of a Service Bus queue. The document-processing worker uses that directly.

Why Container Apps over AKS?

AKS gives full control over the Kubernetes API, custom scheduling, and networking, none of which this design needs: there's one API service and one background worker, not a large mesh of interdependent services that benefits from Kubernetes' flexibility. Container Apps gets the container-based deployment model without requiring the team to run or reason about a Kubernetes control plane, the same "managed platform over self-run cluster" instinct behind the Cloud Run choice above.

LLM and RAG

Azure OpenAI Service, accessed through Microsoft Foundry (formerly Azure AI Foundry), provides the LLM. Retrieval follows the same shape as the other two clouds, stored in Azure Database for PostgreSQL Flexible Server with the pgvector extension. The case for a managed API over self-hosting matches Bedrock and Gemini Enterprise Agent Platform above.

Azure OpenAI has one distinguishing wrinkle: it offers the same underlying OpenAI models as the public API, under Microsoft's enterprise contractual terms and data-handling commitments, including keeping request and response data inside the customer's own tenant boundary instead of sending it to OpenAI directly. For a product handling privileged legal documents, that posture is part of why it's the right choice on Azure.

Async processing and queue

Document processing uses a Service Bus queue. The upload handler enqueues a message; the Container Apps worker, scaled via KEDA on queue depth rather than HTTP traffic, picks it up, parses and embeds the document, and writes the result to Postgres. This is functionally closest to the SQS-based design on AWS, a pull-based queue with a separate worker, rather than Cloud Tasks' push-based model on GCP, and the retry-on-failure behavior works the same way.

Answer evaluation and citation verification

The inline citation check again lives in the API service's own code. The scheduled evaluation runs as a Container Apps job on a cron trigger, the same shape as a Cloud Run job on GCP: a container that starts, does the run, writes its scores to Postgres and its deltas to Application Insights, and exits. Microsoft Foundry ships evaluation tooling that covers groundedness scoring out of the box, and it's worth knowing about, though at Casewell's size the held-out question set and a few scoring rules are enough and carry no extra service to learn.

Identity, secrets, and monitoring

Entra External ID (Azure's customer-facing identity product, distinct from workforce Entra ID) handles associate login, with tenant (firm) association carried as a custom attribute on the identity, the same role Cognito and Identity Platform play in the other two designs. Key Vault holds secrets and connection strings. Azure Monitor and Application Insights cover logging, metrics, and basic alerting. Managed identities on each Container Apps service scope access to exactly the resources that service needs.

Single-region here too, for the same reason as the other two clouds.

Intentionally Not Implemented

Capability comparison

CapabilityAWSGCPAzureDecision rationale
DNSRoute 53Cloud DNSAzure DNSA thin, interchangeable layer at this stage.
CDNCloudFrontCloud CDNFront Door + CDNServes the static frontend at the edge; picked mainly for how it pairs with each cloud's static-hosting option.
Load balancingHandled by API GatewayHandled by Cloud Run's ingressHandled by Container Apps' ingressEach compute platform's built-in ingress absorbs the job, so there's no load balancer to configure by hand on any of the three.
Compute (API)Lambda, one function per operationCloud Run, one container for the whole APIContainer Apps, one container for the whole APIAWS's function-per-handler model and GCP/Azure's container-per-service model are different default units of deployment.
Container platformFargateCloud RunContainer AppsNot used on AWS; Lambda covers this workload there. On GCP and Azure the container platform is the API compute, covered in the row above. Compute Options covers the general choice between the two models.
KubernetesEKSGKEAKSDeliberately unused on all three; see Intentionally Not Implemented above.
Relational databaseAurora PostgreSQL Serverless v2Cloud SQL for PostgreSQLAzure Database for PostgreSQL Flexible ServerAll three are managed Postgres with pgvector enabled; close to a true equivalence.
NoSQL databaseDynamoDBFirestoreCosmos DBNot used. Casewell's access patterns are relational: firms, users, documents, and the relationships between them.
CacheElastiCacheMemorystoreAzure Managed Redis (replaces Azure Cache for Redis)Not used yet. Traffic volume doesn't justify a cache layer; the database handles current load directly.
Object storageS3Cloud StorageBlob StorageStores uploaded documents, tenant-scoped by key prefix and access policy on all three; effectively interchangeable.
Messaging (queue)SQSCloud TasksService BusSQS and Service Bus are pull-based queues a worker polls; Cloud Tasks is push-based, invoking the worker's endpoint directly.
Event bus / pub-subEventBridgePub/SubEvent GridNot used. One consumer per job type right now; an event bus solves fan-out to multiple consumers, a problem this stage doesn't have.
SecretsSecrets ManagerSecret ManagerKey VaultFunctionally equivalent across all three.
IAMIAM + CognitoCloud IAM + Identity PlatformEntra ID + Entra External IDEach pairs a workload-identity system with a customer-facing identity product; the split is consistent across clouds even though product names differ.
ObservabilityCloudWatchCloud Logging / Cloud MonitoringAzure Monitor / Application InsightsBasic logs, metrics, and alarms on all three, sized to what a three-person team needs.
Analytics warehouseRedshiftBigQuerySynapseNot used. Usage analytics run as direct queries against the operational database; data volume doesn't justify a separate warehouse.
LLM platformBedrockGemini Enterprise Agent Platform, formerly Vertex AI (Gemini)Azure OpenAI (via Microsoft Foundry, formerly Azure AI Foundry)All three are managed, pay-per-token LLM APIs, differing mainly in enterprise data-handling terms and model selection.
Vector searchpgvector on Aurorapgvector on Cloud SQLpgvector on Azure DB for PostgreSQLThe same architectural decision on all three: a vector-search extension on the existing relational database instead of a dedicated vector database. See the ADR below.
RAG orchestrationHand-rolled in the Lambda API layerHand-rolled in the Cloud Run API layerHand-rolled in the Container Apps API layerA packaged managed RAG product (Bedrock Knowledge Bases, for instance) was available on all three; all three chose application-code orchestration. See the ADR below.
Answer evaluationScheduled Lambda against a held-out question setCloud Run job on Cloud SchedulerContainer Apps job on a cron triggerA requirement, not a later investment: a fabricated citation is the failure mode a law firm can't absorb. Built in application code on all three rather than on each cloud's evaluation product, which is more tooling than a held-out set of questions needs at this size.
GPU / accelerator infrastructureNone provisionedNone provisionedNone provisionedAbstracted away by the managed LLM APIs above; relevant only if Casewell needs to fine-tune or self-host a model.

Architecture decision records

Decision

Use a managed vector-search extension (pgvector) on the existing relational database instead of standing up a dedicated vector database.

Why

Casewell already needs a relational database for firms, users, and documents. Adding pgvector to that same database means one system to operate, one connection pool, one backup policy, and transactional consistency between a document's metadata and its embeddings, instead of two systems that have to be kept in sync.

Alternatives considered

Trade-off

pgvector's approximate nearest-neighbor performance falls behind a purpose-built vector database at very large scale and very high query-per-second rates. Casewell's document volume per firm and overall query rate are nowhere near that threshold yet.

Revisit when

Vector search latency becomes measurably worse under real load, or overall embedding volume grows enough that dedicated vector infrastructure would clearly outperform the relational extension.

Decision

Use a managed LLM API (Bedrock, Gemini Enterprise Agent Platform, or Azure OpenAI) instead of self-hosting a model.

Why

Self-hosting means owning GPU provisioning, scaling, and model lifecycle management, none of which moves the needle on whether the product works. A managed API converts that entire problem into a per-token line item.

Alternatives considered

Trade-off

Per-token pricing on a managed API is more expensive per request than self-hosted inference at high, sustained volume. At current volume, that crossover point is far away.

Revisit when

Inference cost becomes a material share of total spend at meaningful, sustained query volume.

Decision

Deploy to a single region on each cloud, with no cross-region replication.

Why

A regional outage is rare and, for a company this size, survivable. Multi-region active-active is an ongoing engineering cost paid on every deploy going forward, not a one-time setup fee.

Alternatives considered

Trade-off

A regional outage takes the whole product down until it resolves. At current scale, that's an acceptable, rare risk against the alternative of ongoing operational overhead this team doesn't have the headcount to carry.

Revisit when

Customers start signing contracts with an uptime SLA that a single-region design can't credibly support.

Decision

Use shared infrastructure with logical multi-tenant isolation (tenant ID on every row, scoped IAM and object-storage policies) instead of separate infrastructure per law firm.

Why

Per-tenant infrastructure (a separate database or environment per firm) would multiply operational surface area by the number of customers, which doesn't scale for three engineers and is unnecessary given fewer than a hundred tenants.

Alternatives considered

Trade-off

Logical isolation depends on every code path correctly filtering by tenant ID, and a bug in that filtering is a confidentiality incident. Authentication-level tenant scoping and consistent query patterns manage that risk without eliminating it.

Revisit when

A customer's contract or a compliance requirement demands physically separate infrastructure, or the shared-schema model becomes an operational bottleneck.

Decision

Use managed serverless or container compute (Lambda, Cloud Run, Container Apps) instead of a self-managed Kubernetes cluster.

Why

There's no workload here, a handful of API routes and one background worker, that needs Kubernetes' scheduling flexibility or custom networking control, and no team with the spare capacity to operate a cluster.

Alternatives considered

Trade-off

Less flexibility over scheduling, networking, and deployment strategy than Kubernetes would offer. That flexibility has no current use.

Revisit when

The number of independently deployed services grows enough that they need coordinated scheduling and networking control that a managed serverless platform doesn't expose.

Decision

Hand-roll RAG orchestration in the application layer instead of adopting a packaged managed RAG product.

Why

Casewell's retrieval logic is specific: two distinct corpora per query (firm-private documents and shared public case law), with tenant-scoped access rules that a generic managed RAG product's default assumptions don't map onto cleanly. Writing the retrieval and prompt-assembly logic directly keeps that specificity visible and testable.

Alternatives considered

Trade-off

More application code to write and maintain than a packaged product would require, in exchange for retrieval logic the team fully understands and can audit for tenant-isolation correctness.

Revisit when

A managed RAG product's tenant-scoping and multi-corpus support matures enough to cover this exact access pattern without a workaround.

Decision

Process large document uploads asynchronously through a queue instead of synchronously within the upload request.

Why

Parsing, chunking, and embedding a large PDF filing takes longer than an HTTP request should reasonably block for, and a failed in-request attempt would mean re-uploading the whole file.

Alternatives considered

Trade-off

A document isn't searchable the instant it's uploaded; there's a short processing delay, communicated to the user rather than hidden.

Revisit when

Near-instant searchability of a freshly uploaded document becomes a hard product requirement.

Decision

Treat citation verification and scheduled answer evaluation as launch requirements, and build both in application code rather than adopting a managed evaluation product.

Why

A fabricated case citation is the one output failure a law firm cannot absorb, and it's cheap to catch: every citation in an answer either matches a document retrieval returned or it doesn't. Pairing that inline check with a scheduled run against a held-out question set gives the team a number that moves when retrieval or prompting degrades, which is the difference between knowing quality dropped and hearing about it from a customer.

Alternatives considered

Trade-off

The held-out question set is small and built with a handful of friendly firms, so it covers the research questions those firms happen to ask rather than the full space of legal research. It will miss failure modes nobody thought to write a test case for. It still catches regressions, which is the job it's there to do.

Revisit when

Usage grows enough to sample real production questions for evaluation instead of relying on a hand-built set, or a scoring dimension specific to legal correctness needs more than a groundedness check.

Decision

Run usage analytics as direct queries against the operational database instead of building a dedicated analytics warehouse.

Why

The current analytics need is basic: which firms and associates are using the product. That's answerable with a handful of SQL queries against the same database already storing the data, with no ETL pipeline to build or maintain.

Alternatives considered

Trade-off

Analytics queries run against the same database serving live traffic, which could eventually compete for resources with production queries.

Revisit when

Analytics queries start measurably affecting production latency, or the reporting needs grow past what direct SQL against the operational schema can reasonably answer.

Due diligence questions

Why did you choose a managed vector extension over a dedicated vector database?
Because the query volume and document count don't come close to the point where a dedicated vector database would outperform pgvector, and keeping vectors in the same database as their source metadata avoids a second system to keep in sync.

What happens if traffic increases 20x tomorrow?
The serverless and managed-container compute layers scale automatically with no code change. The database is the likely first bottleneck. Aurora Serverless v2, Cloud SQL, and Azure DB for PostgreSQL Flexible Server all support vertical scaling and read replicas, the next lever to pull.

Where is the single point of failure?
The relational database. It's a single writer in a single region on all three clouds; if it goes down, nothing that depends on it (auth checks, document metadata, vector search) works. That's an accepted risk at this stage.

How do you prevent one law firm's data from being visible to another?
Isolation is enforced at more than one layer: tenant ID on every database row and every retrieval query, tenant scoping baked into the identity token at login rather than checked only in application code, and object-storage access policies scoped by tenant prefix. No single layer is trusted alone.

How would you cut the cloud bill by 50%?
Almost everything here already bills per use rather than for idle capacity, so the biggest remaining lever is the LLM API itself: shorter prompts, more aggressive caching of repeated questions, and routing simpler queries to a cheaper model tier where full reasoning capability isn't needed.

Why aren't you using Kubernetes?
The scheduling and networking control it provides has no use in a system this small, and running it would put a three-person team on cluster operations.

What would make you change that decision?
A meaningful increase in the number of independently deployed services that need coordinated scheduling, or a workload with resource requirements the managed platforms can't express well.

What happens when the LLM provider has an outage?
Question answering fails for the duration of the outage; there's no fallback provider wired in at this stage. That's a known gap, traded deliberately against the complexity of running multi-provider failover at a stage that doesn't have to guarantee availability through someone else's outage.

How do you measure whether an answer is correct?
At two points. Every answer's citations are checked against the documents retrieval actually returned before the answer is displayed, so a case the model invented gets stripped rather than shown. Separately, a scheduled job runs a held-out set of research questions with known-correct answers through the full pipeline and scores whether the right authority was retrieved and whether the answer's claims trace back to it. That gives a number that moves when a prompt or model change degrades quality, which is what makes it safe to change either one. Usage analytics run alongside both, and they measure engagement, not correctness.

What's the first thing that breaks if you 10x your customer count?
The relational database's connection and query capacity, followed closely by the document-processing queue's throughput during a shared filing-deadline spike across many more firms at once.