Generative AI Architecture · Part 3 of 9
Embeddings: Meaning as Geometry
Turning meaning into numbers, so that similar things end up close together.
An embedding is a list of numbers, a vector, that represents a piece of text, an image, or another piece of content, produced so that things with similar meaning end up as nearby vectors and things with different meaning end up far apart. Everything a transformer does internally already depends on this idea. Tokens, Embeddings, and Position covers how a model turns individual tokens into vectors as its first step; an embedding model stops at that representation and hands it back instead of continuing on to generate text.
Measuring "close together"
Two vectors are compared with the same tools covered in AI Core Math Review: a dot product, or the closely related cosine similarity, gives a single number that's larger when two vectors point in a similar direction and smaller (or negative) when they don't. Applied to embeddings, that number becomes a similarity score between two pieces of content. “How do I reset my password?” and “I forgot my login credentials” share almost no words in common, but a good embedding model places them close together in vector space, because they mean nearly the same thing.
Semantic search vs. exact keyword match
A traditional keyword search index matches on the literal words present in a query and a document. It's fast and precise when the vocabulary matches, and it fails silently when it doesn't: a search for “cancel my subscription” won't find a help article titled “how to end recurring billing”, because the two share no word for the index to match on.
Embedding-based search, often called semantic search, sidesteps that problem by comparing meaning instead of surface text. The query and every candidate document are each converted into an embedding once, and retrieval becomes a search for the nearest vectors to the query's vector, a nearest-neighbor search. For generative AI, that's the mechanism that lets a model be handed only the handful of documents relevant to a question, out of a much larger collection it was never trained on.
| Keyword search | Semantic (embedding) search | |
|---|---|---|
| Matches on | Literal words and phrases | Meaning, regardless of exact wording |
| Strength | Precise when vocabulary is known and consistent (product codes, exact names) | Finds relevant content phrased differently than the query |
| Weakness | Misses paraphrases and synonyms entirely | Can retrieve content that's topically similar but not what's needed |
Neither approach dominates the other, so production retrieval systems commonly run both and combine the results.
What makes an embedding good
Embedding models are themselves trained, usually by being shown pairs of text known to be related (a question and its correct answer, two paraphrases of the same sentence) and pairs known to be unrelated, then adjusted so related pairs land closer together and unrelated pairs land farther apart. The result is specific to what the model was trained on: an embedding model trained mostly on general web text won't necessarily place two pieces of legal jargon or two lines of source code as accurately as a model trained with that kind of content in mind. Which embedding model you pick is a retrieval-quality decision in its own right.
One practical detail shapes everything built on top: embeddings are computed once per document, ahead of time, and stored, so they don't have to be recomputed on every query. That stored collection of vectors needs somewhere to live and a way to be searched quickly.
Every vector in an index has to come from the same model
A model's output coordinates mean something only relative to that model. Dimension 412 of one embedding model and dimension 412 of another were learned independently and encode unrelated things, so a dot product between a vector from model A and a vector from model B produces a number with no meaning behind it. It will still be computed, and it will still rank results, which is what makes this failure quiet: retrieval keeps returning documents, they're just the wrong ones. Models also differ in how many dimensions they output, so in many cases the mismatch surfaces as an outright error instead.
Two consequences follow:
- The query has to be embedded by the same model as the documents. The query is embedded at request time, the documents were embedded weeks earlier, and the system has to guarantee both went through the same model and the same version of it.
- Changing the embedding model means rebuilding the whole index. Every stored vector has to be recomputed, which for a large corpus is a project in its own right, usually run as a parallel index that's evaluated against the current one before anything is cut over.