Generative AI Architecture · Part 5 of 9
Vector Databases: When You Need One
A specialized index earns its place once scale and speed demand it.
The indexing step of a RAG pipeline needs somewhere to store embeddings and search them quickly. That somewhere can be a dedicated vector database, purpose-built for the job, or an extension added to a database the team already runs. The choice affects cost, latency, and how much there is to operate.
What a vector index does
Comparing a query vector against every stored vector one at a time, an exact nearest-neighbor search, is accurate but scales linearly with the number of vectors: doubling the collection roughly doubles the search time. A dedicated vector index instead builds a structure (commonly a graph-based method like HNSW) that finds vectors that are very likely, though not mathematically guaranteed, to be the true nearest neighbors, in a small fraction of the time. That trade, exactness for speed, is what makes semantic search practical at scale: it's the difference between a search that takes milliseconds over millions of vectors and one that takes seconds.
When it earns its place
A dedicated vector database is worth adopting when a collection is large enough, or a query volume high enough, that exact search becomes a measurable bottleneck, and when the team needs vector-specific features: filtering by metadata alongside similarity, hybrid keyword-plus-vector queries, or horizontal scaling built specifically for this workload.
It's unnecessary overhead in the common case where a team already runs a relational database and the collection of embeddings is small enough, tens of thousands to low millions of vectors, that a vector extension on that same database keeps up. A second specialized data store is a second thing to operate, back up, secure, and keep in sync with the primary one, and that cost only pays for itself once the scale or the feature need arrives.
Options by cloud
| Cloud | Dedicated vector search | Vector support added to an existing database |
|---|---|---|
| AWS | OpenSearch Service (vector engine); Amazon S3 Vectors for low-cost storage of very large collections; Amazon Bedrock Knowledge Bases as a managed RAG layer on top | Aurora PostgreSQL with the pgvector extension |
| GCP | Vector Search (formerly Vertex AI Vector Search) | AlloyDB or Cloud SQL for PostgreSQL, both with pgvector |
| Azure | Azure AI Search | Azure Database for PostgreSQL with pgvector; Cosmos DB's vector search capability |
Either path, dedicated service or extension on an existing database, plugs into the same pipeline and does semantic search. The decision comes down to operational cost and scale.