Cloud and AI Architecture: Case Studies · Part 2 of 4
Casewell
A three-person engineering team, fewer than a hundred customers, and one architecture built three times over.
Casewell is a pre-seed company with three engineers, and that constraint decides most of what follows. Below is the company, its requirements, and one architecture designed three times over, once per major cloud provider.
Casewell is fictional: a composite typical of its funding stage, modeled on no specific business.
The company
Casewell is an AI research assistant built for small and mid-size law firms. An associate can ask it a question in plain language, something like "find precedent for a landlord withholding a security deposit past the statutory deadline in this jurisdiction," and it returns relevant case law, cross-referenced against the firm's own past filings and internal memos, with citations back to both. The pitch to a firm is simple: the research that used to take a junior associate an afternoon now takes minutes, and it draws on the firm's own institutional memory instead of only public sources.
The company has five to ten employees, three of them engineers. It has fewer than a hundred paying law-firm customers, all small or mid-size firms rather than the largest national practices. It's priced as SaaS, either per seat (one price per associate with access) or per firm (a flat fee scaled loosely to headcount), and most current customers are still on whichever pricing model let them get in the door fastest; the company hasn't settled on one. It's raised a pre-seed or seed round and is expected to grow. Growth at this stage means firms adopting the product and continuing to use it; a revenue target comes later.
Usage pattern
Traffic isn't steady. Litigation runs on deadlines: filing dates, discovery cutoffs, court calendars. Usage spikes around each one, then drops off. A firm might barely touch the product for two weeks and then have every associate hammering it the two days before a filing is due. Whatever gets built has to absorb that kind of burst without either falling over during it or paying for capacity that sits idle the rest of the time.
Requirements
Functional
- Web application: the only interface. No native mobile app.
- Authentication: individual associate logins, scoped to their firm.
- Document upload: firms upload their own past filings, memos, and case files for the assistant to draw on.
- LLM-powered question answering: natural-language questions, answered in natural language, with citations.
- Retrieval-augmented generation (RAG) over two distinct corpora: each firm's own uploaded documents, and a shared corpus of public case law common to every tenant.
- Citation verification: every citation in an answer is checked against the documents actually retrieved for that answer before the answer reaches the associate. A citation the model produced but the retrieval step never returned is a fabrication, and it gets stripped or flagged rather than displayed.
- Answer evaluation: a held-out set of research questions with known-correct answers, re-run against the full pipeline before any prompt, model, or retrieval change ships, scored on whether the right authority was found and whether the answer stayed grounded in it.
- Basic usage analytics: which firms and associates are using the product, and how often, to tell the company whether anyone's getting value out of it.
- Asynchronous document processing: large PDF filings need to be parsed, chunked, and embedded before they're searchable, and that can't happen inline while someone waits on an upload button.
- Multi-tenant isolation: one firm's documents, and any answers derived from them, must never be visible to another firm.
- Basic observability: enough logging and error visibility for a three-person team to know when something's broken, without a dedicated on-call rotation or a platform team to run one.
Non-functional
- Reasonable latency: a research question answered in a few seconds is fine; nobody expects sub-second responses from a system that's synthesizing case law, but a thirty-second wait will read as broken.
- Acceptable early-stage uptime: firms use this during business hours in their own time zone, not around the clock. High availability matters less here than at later stages; a brief outage is a bad afternoon, not a breached SLA with a Fortune 500 customer.
- Data confidentiality between tenants: the one requirement that can't slip. Legal documents routinely contain privileged and confidential client information, and a firm that discovers another firm's filings leaked through the platform doesn't file a bug report, they leave and tell every other firm they know.
Major risks
Product risk dominates. Current-generation models can produce plausible-looking answers. The open question is whether associates and partners trust an AI-sourced answer enough to rely on it, and whether that trust survives the first time the assistant gets something subtly wrong. Legal research has a low tolerance for confident, well-formatted mistakes, and adoption depends on the product earning trust one correct answer at a time.
This is why the two evaluation requirements above are requirements rather than a later investment. A fabricated case citation in legal work is not a quality issue to be improved on next quarter; lawyers have been sanctioned for filing one. A product that can produce one and has no mechanism to catch it is not shippable to a law firm at any stage, including this one. The bar at three engineers is a cheap, automated bar, not a mature evaluation practice, but it can't be zero.
Confidentiality risk comes second. Attorney-client privilege and work-product protection push a confidentiality failure well past a breach notification: it becomes a professional-responsibility problem for the firm and an existential one for Casewell. Multi-tenant isolation is a condition of the product being usable at all.
Architecture: AWS
These requirements point toward one instinct: reach for the most managed option in every category, even where it costs more per unit than running it yourself, because the team has three engineers and no time to operate infrastructure. Here's the resulting design.
flowchart TB U[Associate's
browser] --> CF[CloudFront +
S3 static frontend] U --> APIGW[API Gateway] APIGW --> LAM[Lambda: API] LAM --> BR[Bedrock: LLM] BR -->|answer| VER[Citation verifier] LAM --> SQS[SQS queue] SQS --> LAMW[Lambda:
document worker] LAM --> AUR[(Aurora Postgres
+ pgvector)] LAMW --> AUR LAM --> S3D[S3:
document storage] LAMW --> S3D EV[Lambda: evaluation
runner, scheduled] --> BR EV --> AUR LAM --> PLAT[Cognito
Secrets Manager
CloudWatch] EV --> PLAT
Frontend
A static single-page app, built once and served as static files from S3 behind CloudFront. There's no server-rendering requirement here and no reason to pay for one; the app calls the API layer directly for everything dynamic. CloudFront also gives free TLS and a CDN edge at close to no operational cost, so nobody on the team manages a certificate.
API and compute
The API layer is API Gateway in front of Lambda functions, one function (or small group of functions) per logical operation: authentication callbacks, document upload, question answering, analytics reads. Compute Options covers the general trade-offs; at this traffic pattern, serverless compute is close to a free win.
Why serverless?
Casewell's traffic is bursty and, most of the time, small: quiet stretches between filing deadlines, sharp spikes around them. A fleet of always-on servers sized for the spike sits mostly idle the rest of the time; sized for the average, it falls over during the spike. Lambda scales per request and bills per invocation, so the bursty part of the usage pattern stops being a capacity-planning exercise. The trade-off: cold starts add latency to an occasional request, and a sustained, steady-state high-traffic workload would eventually be cheaper on containers running around the clock. At fewer than a hundred customers with spiky usage, that crossover point isn't close.
LLM and RAG
Amazon Bedrock provides the LLM itself, called from the Lambda API layer with the retrieved context assembled into the prompt. Retrieval draws from two sources, matching the two corpora from the requirements: each firm's own uploaded documents, and a shared public case-law corpus common to every tenant. Both are embedded and stored in Aurora PostgreSQL Serverless v2 with the pgvector extension: one database, two logically separated sets of vectors, tenant ID carried on every row so a query can never cross into another firm's documents. Retrieval-Augmented Generation covers the general shape of this pipeline; the storage choice gets its own ADR below.
Why managed AI APIs instead of self-hosting?
Self-hosting an open-weight model means owning GPU provisioning, model updates, scaling under load, and every failure mode that comes with running inference infrastructure, work that has nothing to do with whether lawyers find the product useful. Bedrock trades a per-token cost premium for all of that disappearing. For a team still trying to find out whether the product works at all, the premium is worth paying. The question this stage needs answered is product-market fit; GPU/TPU AI Infrastructure is a problem worth having later.
Answer evaluation and citation verification
Two mechanisms, at two different points in the system.
The first runs inline, on every answer. Before a response is returned, the API layer parses the citations out of it and checks each one against the set of documents retrieval actually returned for that question. A citation that doesn't match anything in that set never came from the firm's documents or the case-law corpus; the model produced it. Those get stripped, and the answer is returned with a note that a source couldn't be confirmed. This is ordinary application code, not a service, and it costs one pass over the retrieved chunks. It catches the single failure mode with the worst consequences for a law firm.
The second runs on a schedule, in a separate Lambda evaluation runner. It holds a set of research questions with known-correct answers, built with a few cooperative customer firms, and runs the whole pipeline against them: retrieval, prompt assembly, generation, citation check. Each run scores whether the correct authority was retrieved at all, whether the answer's claims trace back to the retrieved text, and how often the citation verifier had to strip something. Results land in Aurora and the run-over-run deltas go to CloudWatch, so a prompt or model change that quietly degrades retrieval quality shows up as a number before a customer notices it. Three engineers can run this; it's a scheduled job and a spreadsheet's worth of test cases. Evaluating AI Systems covers what this practice grows into with more people and more usage behind it.
Database
Aurora PostgreSQL Serverless v2 does double duty: ordinary relational data (firms, users, documents, usage events) and vector search for RAG, in the same database. See Databases: Relational and NoSQL and Vector Databases for the general trade-offs; the specific reasoning for Casewell is one of the ADRs below.
Object storage
Uploaded documents (the original PDFs, not the extracted text) live in S3, one bucket with per-tenant key prefixes and bucket policies that scope access by tenant. This follows the general pattern in Storage Fundamentals.
Async processing and queue
A large PDF filing can run hundreds of pages, and parsing, chunking, and embedding it takes long enough that it can't happen inline during the HTTP request that uploaded it. The upload handler drops a message onto an SQS queue; a separate Lambda worker picks it up, does the processing, and writes the resulting chunks and embeddings into Aurora. If that worker fails partway through, the message becomes visible again after its visibility timeout and gets retried, instead of the upload silently vanishing. See Messaging and Event-Driven Architecture for the general pattern.
Identity, secrets, and monitoring
Cognito handles associate login, with each user scoped to a firm (tenant) at the identity level, so tenant isolation starts at authentication rather than being bolted on afterward in application code. Secrets Manager holds API keys and database credentials; nothing sensitive sits in an environment variable or, worse, in code. CloudWatch covers logs, metrics, and basic alarms: enough for three engineers to know something broke, without standing up a dedicated observability stack. IAM roles are scoped per Lambda function to exactly what that function touches, following the least-privilege pattern in Identity and Access Management.
Why single-region?
A single AWS region, one Aurora writer, no cross-region replication. Multi-region active-active buys protection against a full regional outage, at the cost of a data replication strategy, conflict resolution, and doubled operational surface area. For a company that doesn't yet know if firms will renew, that's cost with no matching benefit. A regional outage is survivable and rare; product risk and confidentiality risk are where the engineering time goes.
Architecture: GCP
This design is re-derived from GCP's own service model against the same requirements, rather than translated service by service from the AWS version above. Most of it lands in a similar place, because the requirements are the same. Where it doesn't, the difference is worth the paragraph that explains it.
flowchart TB U[Associate's
browser] --> CDN[Cloud CDN +
Cloud Storage
frontend] U --> RUN[Cloud Run: API] RUN --> VTX[Gemini Enterprise Agent Platform
formerly Vertex AI
Gemini models] VTX -->|answer| VER[Citation verifier] RUN --> TASKS[Cloud Tasks] TASKS --> RUNW[Cloud Run:
document worker] RUN --> SQL[(Cloud SQL Postgres
+ pgvector)] RUNW --> SQL RUN --> GCS_D[Cloud Storage:
documents] RUNW --> GCS_D EV[Cloud Run job:
evaluation, scheduled] --> VTX EV --> SQL RUN --> PLAT[Identity Platform
Secret Manager
Cloud Logging /
Monitoring] EV --> PLAT
The static single-page app is served from Cloud Storage behind Cloud CDN, the same shape as AWS's S3-plus-CloudFront pairing and for the same reason: the frontend is a bundle of static files, and serving it from object storage at the edge removes a server from the picture entirely. Uploaded documents live in Cloud Storage the same way, one bucket with per-tenant object prefixes and IAM conditions scoping access by tenant.
API and compute: Cloud Run
The API layer runs on Cloud Run: one container holding every route for the API. Where the AWS design splits handlers into separate Lambda functions, Cloud Run's unit is a single request-driven container that can serve an entire API's worth of routes, which fits a three-person team's mental model better: one deployable, one image, one thing to reason about, instead of a directory of individually versioned functions. It still scales to zero between the deadline-driven traffic spikes described in the requirements, and it still bills per request, so it gets the same bursty-traffic benefit as Lambda with less code fragmentation.
Why Cloud Run over GKE Autopilot?
GKE Autopilot removes node management from Kubernetes but keeps the Kubernetes API, its resource model, and its operational vocabulary (pods, deployments, ingress objects) in the picture. None of that buys anything here: there's one service, no need for custom scheduling, network policies, or cross-service traffic management that would justify carrying Kubernetes concepts around. Cloud Run gets the same "give it a container, it handles the rest" simplicity as the frontend's static hosting, applied to the API. Compute Options covers this decision generally; for Casewell, Kubernetes is a bet on future operational complexity this team doesn't have yet.
LLM and RAG
Gemini Enterprise Agent Platform (formerly Vertex AI) serves the Gemini model family to the Cloud Run API layer, with retrieved context assembled into the prompt the same way as the AWS design, and the same two-corpus retrieval shape, embedded and stored in Cloud SQL for PostgreSQL with the pgvector extension. The case for a managed model API over self-hosting is the same as Bedrock's, above. The platform also gives an upgrade path if the product later needs a fine-tuned model, without first building the serving infrastructure that would require.
Cloud SQL versus AlloyDB
GCP offers a second managed Postgres option, AlloyDB, built for higher-throughput and analytics-heavy workloads and priced accordingly. At fewer than a hundred customers with a modest, if bursty, query volume, Cloud SQL for PostgreSQL is the right-sized choice: cheaper, simpler to reason about, and adequate for the read and write volume this stage produces. AlloyDB becomes worth revisiting once query volume or vector-search latency under load becomes a bottleneck.
Async processing and queue
Document processing uses Cloud Tasks, a subtle but meaningful difference from the SQS pattern on AWS: Cloud Tasks is push-based, invoking an HTTP endpoint on the worker directly with built-in retry and rate-limiting controls, where SQS is pull-based and the worker polls the queue for messages. For one well-defined job type (process this uploaded document) with one consumer, Cloud Tasks' push model needs less plumbing than a Pub/Sub topic and subscription would for a point-to-point job queue. Pub/Sub earns its place once more than one kind of consumer reacts to the same event.
Answer evaluation and citation verification
The inline citation check is application code inside the Cloud Run API container, unchanged in substance from the AWS version. The scheduled half runs as a Cloud Run job triggered by Cloud Scheduler, which is a closer fit than a second always-on service: it's a batch task with a start and an end, and Cloud Run jobs are built for exactly that, where Cloud Run services are built to serve requests. Scores land in Cloud SQL and the run-over-run deltas in Cloud Monitoring.
Identity, secrets, and monitoring
Identity Platform handles associate login, with custom claims carrying the tenant (firm) association the same way Cognito's user pool groups do on AWS. Secret Manager holds credentials and API keys. Cloud Logging and Cloud Monitoring give the same coverage as CloudWatch, with IAM roles on each Cloud Run service scoped to only what that service needs.
Single-region, for the same reason as AWS: one Cloud SQL primary, no cross-region replication, a regional outage being a bad day rather than an existential one at this stage.
Architecture: Azure
Same requirements, same constraints. Here's what they produce on Azure.
flowchart TB U[Associate's
browser] --> FD[Front Door + CDN,
Blob Storage
frontend] U --> CA[Container Apps: API] CA --> AOAI[Azure OpenAI] AOAI -->|answer| VER[Citation verifier] CA --> SB[Service Bus queue] SB --> CAW[Container Apps:
document worker] CA --> PG[(Azure DB for PostgreSQL
+ pgvector)] CAW --> PG CA --> BLOB_D[Blob Storage:
documents] CAW --> BLOB_D EV[Container Apps job:
evaluation, scheduled] --> AOAI EV --> PG CA --> PLAT[Entra External ID
Key Vault
Azure Monitor /
App Insights] EV --> PLAT
The static single-page app is served from Blob Storage's static website hosting behind Azure Front Door, which handles both the CDN edge and TLS termination in one resource, landing in the same place as the pairings on the other two clouds. Uploaded documents live in Blob Storage, one container with per-tenant path prefixes and access scoped by tenant through Azure RBAC conditions.
API and compute: Container Apps
The API layer runs on Azure Container Apps: a single containerized service holding the API, scaling on request volume the same way Cloud Run does. Container Apps is built on KEDA underneath, so it can scale a service on triggers other than raw HTTP concurrency, including the depth of a Service Bus queue. The document-processing worker uses that directly.
Why Container Apps over AKS?
AKS gives full control over the Kubernetes API, custom scheduling, and networking, none of which this design needs: there's one API service and one background worker, not a large mesh of interdependent services that benefits from Kubernetes' flexibility. Container Apps gets the container-based deployment model without requiring the team to run or reason about a Kubernetes control plane, the same "managed platform over self-run cluster" instinct behind the Cloud Run choice above.
LLM and RAG
Azure OpenAI Service, accessed through Microsoft Foundry (formerly Azure AI Foundry), provides the LLM. Retrieval follows the same shape as the other two clouds, stored in Azure Database for PostgreSQL Flexible Server with the pgvector extension. The case for a managed API over self-hosting matches Bedrock and Gemini Enterprise Agent Platform above.
Azure OpenAI has one distinguishing wrinkle: it offers the same underlying OpenAI models as the public API, under Microsoft's enterprise contractual terms and data-handling commitments, including keeping request and response data inside the customer's own tenant boundary instead of sending it to OpenAI directly. For a product handling privileged legal documents, that posture is part of why it's the right choice on Azure.
Async processing and queue
Document processing uses a Service Bus queue. The upload handler enqueues a message; the Container Apps worker, scaled via KEDA on queue depth rather than HTTP traffic, picks it up, parses and embeds the document, and writes the result to Postgres. This is functionally closest to the SQS-based design on AWS, a pull-based queue with a separate worker, rather than Cloud Tasks' push-based model on GCP, and the retry-on-failure behavior works the same way.
Answer evaluation and citation verification
The inline citation check again lives in the API service's own code. The scheduled evaluation runs as a Container Apps job on a cron trigger, the same shape as a Cloud Run job on GCP: a container that starts, does the run, writes its scores to Postgres and its deltas to Application Insights, and exits. Microsoft Foundry ships evaluation tooling that covers groundedness scoring out of the box, and it's worth knowing about, though at Casewell's size the held-out question set and a few scoring rules are enough and carry no extra service to learn.
Identity, secrets, and monitoring
Entra External ID (Azure's customer-facing identity product, distinct from workforce Entra ID) handles associate login, with tenant (firm) association carried as a custom attribute on the identity, the same role Cognito and Identity Platform play in the other two designs. Key Vault holds secrets and connection strings. Azure Monitor and Application Insights cover logging, metrics, and basic alerting. Managed identities on each Container Apps service scope access to exactly the resources that service needs.
Single-region here too, for the same reason as the other two clouds.
Intentionally Not Implemented
- No multi-region active-active. One region, covered above.
- No custom Kubernetes platform. Managed serverless and container compute cover every workload in this design. EKS, GKE, and AKS would add a control plane to operate in exchange for orchestration control none of these workloads needs.
- No service mesh. There are a handful of services talking to a handful of managed products, not a large fleet of interdependent microservices that needs traffic shaping, mutual TLS, and distributed tracing between hops.
- No bespoke model hosting. A managed LLM API, covered above.
- No complex data lake. Usage analytics needs are basic enough for a handful of queries against the operational database and platform metrics; a dedicated warehouse and ETL pipeline would be built for data volume this company doesn't have yet.
- No internal developer platform. Three engineers can hold this whole architecture in their heads; a platform team's job (paved roads, self-service tooling, internal APIs for infrastructure) doesn't exist yet because there's no one it would be serving.
Capability comparison
| Capability | AWS | GCP | Azure | Decision rationale |
|---|---|---|---|---|
| DNS | Route 53 | Cloud DNS | Azure DNS | A thin, interchangeable layer at this stage. |
| CDN | CloudFront | Cloud CDN | Front Door + CDN | Serves the static frontend at the edge; picked mainly for how it pairs with each cloud's static-hosting option. |
| Load balancing | Handled by API Gateway | Handled by Cloud Run's ingress | Handled by Container Apps' ingress | Each compute platform's built-in ingress absorbs the job, so there's no load balancer to configure by hand on any of the three. |
| Compute (API) | Lambda, one function per operation | Cloud Run, one container for the whole API | Container Apps, one container for the whole API | AWS's function-per-handler model and GCP/Azure's container-per-service model are different default units of deployment. |
| Container platform | Fargate | Cloud Run | Container Apps | Not used on AWS; Lambda covers this workload there. On GCP and Azure the container platform is the API compute, covered in the row above. Compute Options covers the general choice between the two models. |
| Kubernetes | EKS | GKE | AKS | Deliberately unused on all three; see Intentionally Not Implemented above. |
| Relational database | Aurora PostgreSQL Serverless v2 | Cloud SQL for PostgreSQL | Azure Database for PostgreSQL Flexible Server | All three are managed Postgres with pgvector enabled; close to a true equivalence. |
| NoSQL database | DynamoDB | Firestore | Cosmos DB | Not used. Casewell's access patterns are relational: firms, users, documents, and the relationships between them. |
| Cache | ElastiCache | Memorystore | Azure Managed Redis (replaces Azure Cache for Redis) | Not used yet. Traffic volume doesn't justify a cache layer; the database handles current load directly. |
| Object storage | S3 | Cloud Storage | Blob Storage | Stores uploaded documents, tenant-scoped by key prefix and access policy on all three; effectively interchangeable. |
| Messaging (queue) | SQS | Cloud Tasks | Service Bus | SQS and Service Bus are pull-based queues a worker polls; Cloud Tasks is push-based, invoking the worker's endpoint directly. |
| Event bus / pub-sub | EventBridge | Pub/Sub | Event Grid | Not used. One consumer per job type right now; an event bus solves fan-out to multiple consumers, a problem this stage doesn't have. |
| Secrets | Secrets Manager | Secret Manager | Key Vault | Functionally equivalent across all three. |
| IAM | IAM + Cognito | Cloud IAM + Identity Platform | Entra ID + Entra External ID | Each pairs a workload-identity system with a customer-facing identity product; the split is consistent across clouds even though product names differ. |
| Observability | CloudWatch | Cloud Logging / Cloud Monitoring | Azure Monitor / Application Insights | Basic logs, metrics, and alarms on all three, sized to what a three-person team needs. |
| Analytics warehouse | Redshift | BigQuery | Synapse | Not used. Usage analytics run as direct queries against the operational database; data volume doesn't justify a separate warehouse. |
| LLM platform | Bedrock | Gemini Enterprise Agent Platform, formerly Vertex AI (Gemini) | Azure OpenAI (via Microsoft Foundry, formerly Azure AI Foundry) | All three are managed, pay-per-token LLM APIs, differing mainly in enterprise data-handling terms and model selection. |
| Vector search | pgvector on Aurora | pgvector on Cloud SQL | pgvector on Azure DB for PostgreSQL | The same architectural decision on all three: a vector-search extension on the existing relational database instead of a dedicated vector database. See the ADR below. |
| RAG orchestration | Hand-rolled in the Lambda API layer | Hand-rolled in the Cloud Run API layer | Hand-rolled in the Container Apps API layer | A packaged managed RAG product (Bedrock Knowledge Bases, for instance) was available on all three; all three chose application-code orchestration. See the ADR below. |
| Answer evaluation | Scheduled Lambda against a held-out question set | Cloud Run job on Cloud Scheduler | Container Apps job on a cron trigger | A requirement, not a later investment: a fabricated citation is the failure mode a law firm can't absorb. Built in application code on all three rather than on each cloud's evaluation product, which is more tooling than a held-out set of questions needs at this size. |
| GPU / accelerator infrastructure | None provisioned | None provisioned | None provisioned | Abstracted away by the managed LLM APIs above; relevant only if Casewell needs to fine-tune or self-host a model. |
Architecture decision records
Decision
Use a managed vector-search extension (pgvector) on the existing relational database instead of standing up a dedicated vector database.
Why
Casewell already needs a relational database for firms, users, and documents. Adding pgvector to that same database means one system to operate, one connection pool, one backup policy, and transactional consistency between a document's metadata and its embeddings, instead of two systems that have to be kept in sync.
Alternatives considered
- A dedicated vector database or search service (OpenSearch, Vector Search (formerly Vertex AI Vector Search), Azure AI Search): more headroom at very high query volume, more moving parts to operate.
- A vector database as a separate managed SaaS product: one more vendor relationship and billing relationship for a three-person team to manage.
Trade-off
pgvector's approximate nearest-neighbor performance falls behind a purpose-built vector database at very large scale and very high query-per-second rates. Casewell's document volume per firm and overall query rate are nowhere near that threshold yet.
Revisit when
Vector search latency becomes measurably worse under real load, or overall embedding volume grows enough that dedicated vector infrastructure would clearly outperform the relational extension.
Decision
Use a managed LLM API (Bedrock, Gemini Enterprise Agent Platform, or Azure OpenAI) instead of self-hosting a model.
Why
Self-hosting means owning GPU provisioning, scaling, and model lifecycle management, none of which moves the needle on whether the product works. A managed API converts that entire problem into a per-token line item.
Alternatives considered
- Self-hosting an open-weight model on GPU instances: full control over cost-per-token at scale, at the cost of owning inference infrastructure.
- A smaller, cheaper open-weight model run on CPU: likely too slow or too weak for the reasoning this product asks of it.
Trade-off
Per-token pricing on a managed API is more expensive per request than self-hosted inference at high, sustained volume. At current volume, that crossover point is far away.
Revisit when
Inference cost becomes a material share of total spend at meaningful, sustained query volume.
Decision
Deploy to a single region on each cloud, with no cross-region replication.
Why
A regional outage is rare and, for a company this size, survivable. Multi-region active-active is an ongoing engineering cost paid on every deploy going forward, not a one-time setup fee.
Alternatives considered
- Multi-region active-passive with a cold standby: lower ongoing cost than active-active, still meaningfully more complexity than this stage needs.
- Multi-region active-active: the strongest availability posture, and the most complexity to design, test, and operate correctly.
Trade-off
A regional outage takes the whole product down until it resolves. At current scale, that's an acceptable, rare risk against the alternative of ongoing operational overhead this team doesn't have the headcount to carry.
Revisit when
Customers start signing contracts with an uptime SLA that a single-region design can't credibly support.
Decision
Use shared infrastructure with logical multi-tenant isolation (tenant ID on every row, scoped IAM and object-storage policies) instead of separate infrastructure per law firm.
Why
Per-tenant infrastructure (a separate database or environment per firm) would multiply operational surface area by the number of customers, which doesn't scale for three engineers and is unnecessary given fewer than a hundred tenants.
Alternatives considered
- Per-tenant database instances: strong isolation guarantees, but operational cost that scales linearly with customer count.
- Per-tenant schemas within one database instance: a middle ground, still meaningfully more migration and connection-management overhead than a single shared schema.
Trade-off
Logical isolation depends on every code path correctly filtering by tenant ID, and a bug in that filtering is a confidentiality incident. Authentication-level tenant scoping and consistent query patterns manage that risk without eliminating it.
Revisit when
A customer's contract or a compliance requirement demands physically separate infrastructure, or the shared-schema model becomes an operational bottleneck.
Decision
Use managed serverless or container compute (Lambda, Cloud Run, Container Apps) instead of a self-managed Kubernetes cluster.
Why
There's no workload here, a handful of API routes and one background worker, that needs Kubernetes' scheduling flexibility or custom networking control, and no team with the spare capacity to operate a cluster.
Alternatives considered
- Managed Kubernetes (EKS/GKE/AKS): useful once there are many interdependent services needing fine-grained orchestration; premature here.
Trade-off
Less flexibility over scheduling, networking, and deployment strategy than Kubernetes would offer. That flexibility has no current use.
Revisit when
The number of independently deployed services grows enough that they need coordinated scheduling and networking control that a managed serverless platform doesn't expose.
Decision
Hand-roll RAG orchestration in the application layer instead of adopting a packaged managed RAG product.
Why
Casewell's retrieval logic is specific: two distinct corpora per query (firm-private documents and shared public case law), with tenant-scoped access rules that a generic managed RAG product's default assumptions don't map onto cleanly. Writing the retrieval and prompt-assembly logic directly keeps that specificity visible and testable.
Alternatives considered
- A managed, packaged RAG product (Bedrock Knowledge Bases and equivalents): less code to write and maintain, less control over the retrieval logic's tenant-scoping behavior.
Trade-off
More application code to write and maintain than a packaged product would require, in exchange for retrieval logic the team fully understands and can audit for tenant-isolation correctness.
Revisit when
A managed RAG product's tenant-scoping and multi-corpus support matures enough to cover this exact access pattern without a workaround.
Decision
Process large document uploads asynchronously through a queue instead of synchronously within the upload request.
Why
Parsing, chunking, and embedding a large PDF filing takes longer than an HTTP request should reasonably block for, and a failed in-request attempt would mean re-uploading the whole file.
Alternatives considered
- Synchronous processing with a long request timeout: simpler code, but a fragile user experience and no natural retry mechanism on failure.
Trade-off
A document isn't searchable the instant it's uploaded; there's a short processing delay, communicated to the user rather than hidden.
Revisit when
Near-instant searchability of a freshly uploaded document becomes a hard product requirement.
Decision
Treat citation verification and scheduled answer evaluation as launch requirements, and build both in application code rather than adopting a managed evaluation product.
Why
A fabricated case citation is the one output failure a law firm cannot absorb, and it's cheap to catch: every citation in an answer either matches a document retrieval returned or it doesn't. Pairing that inline check with a scheduled run against a held-out question set gives the team a number that moves when retrieval or prompting degrades, which is the difference between knowing quality dropped and hearing about it from a customer.
Alternatives considered
- Usage signals as a proxy (do associates come back, do they ask follow-up questions): free, and it measures engagement rather than correctness. It cannot detect a confident, well-formatted wrong answer, which is the specific risk here.
- Each cloud's managed evaluation tooling (Microsoft Foundry's evaluators, the Gemini Enterprise Agent Platform's evaluation service): more capable, and more product surface to learn than a held-out question set and a scoring script need at this size.
- Human review of a sample of answers: the highest-quality signal, and not something three engineers and no legal staff can run at any useful frequency.
Trade-off
The held-out question set is small and built with a handful of friendly firms, so it covers the research questions those firms happen to ask rather than the full space of legal research. It will miss failure modes nobody thought to write a test case for. It still catches regressions, which is the job it's there to do.
Revisit when
Usage grows enough to sample real production questions for evaluation instead of relying on a hand-built set, or a scoring dimension specific to legal correctness needs more than a groundedness check.
Decision
Run usage analytics as direct queries against the operational database instead of building a dedicated analytics warehouse.
Why
The current analytics need is basic: which firms and associates are using the product. That's answerable with a handful of SQL queries against the same database already storing the data, with no ETL pipeline to build or maintain.
Alternatives considered
- A dedicated analytics warehouse (Redshift/BigQuery/Synapse) with an ETL pipeline: the right tool at serious data volume, unnecessary overhead at this one.
Trade-off
Analytics queries run against the same database serving live traffic, which could eventually compete for resources with production queries.
Revisit when
Analytics queries start measurably affecting production latency, or the reporting needs grow past what direct SQL against the operational schema can reasonably answer.
Due diligence questions
Why did you choose a managed vector extension over a dedicated vector database?
Because the query volume and document count don't come close to the point where a dedicated vector database would outperform pgvector, and keeping vectors in the same database as their source metadata avoids a second system to keep in sync.
What happens if traffic increases 20x tomorrow?
The serverless and managed-container compute layers scale automatically with no code change. The database is the likely first bottleneck. Aurora Serverless v2, Cloud SQL, and Azure DB for PostgreSQL Flexible Server all support vertical scaling and read replicas, the next lever to pull.
Where is the single point of failure?
The relational database. It's a single writer in a single region on all three clouds; if it goes down, nothing that depends on it (auth checks, document metadata, vector search) works. That's an accepted risk at this stage.
How do you prevent one law firm's data from being visible to another?
Isolation is enforced at more than one layer: tenant ID on every database row and every retrieval query, tenant scoping baked into the identity token at login rather than checked only in application code, and object-storage access policies scoped by tenant prefix. No single layer is trusted alone.
How would you cut the cloud bill by 50%?
Almost everything here already bills per use rather than for idle capacity, so the biggest remaining lever is the LLM API itself: shorter prompts, more aggressive caching of repeated questions, and routing simpler queries to a cheaper model tier where full reasoning capability isn't needed.
Why aren't you using Kubernetes?
The scheduling and networking control it provides has no use in a system this small, and running it would put a three-person team on cluster operations.
What would make you change that decision?
A meaningful increase in the number of independently deployed services that need coordinated scheduling, or a workload with resource requirements the managed platforms can't express well.
What happens when the LLM provider has an outage?
Question answering fails for the duration of the outage; there's no fallback provider wired in at this stage. That's a known gap, traded deliberately against the complexity of running multi-provider failover at a stage that doesn't have to guarantee availability through someone else's outage.
How do you measure whether an answer is correct?
At two points. Every answer's citations are checked against the documents retrieval actually returned before the answer is displayed, so a case the model invented gets stripped rather than shown. Separately, a scheduled job runs a held-out set of research questions with known-correct answers through the full pipeline and scores whether the right authority was retrieved and whether the answer's claims trace back to it. That gives a number that moves when a prompt or model change degrades quality, which is what makes it safe to change either one. Usage analytics run alongside both, and they measure engagement, not correctness.
What's the first thing that breaks if you 10x your customer count?
The relational database's connection and query capacity, followed closely by the document-processing queue's throughput during a shared filing-deadline spike across many more firms at once.