An embedding is a list of numbers that represents a piece of text, and it is built so that text with similar meaning produces a similar list. That is the entire idea. Everything else — vector databases, semantic search, retrieval-augmented generation — is machinery built on top of it.
The useful consequence is that comparing meaning becomes arithmetic. If "my laptop won't charge" and "notebook stops powering on when the cable is out" land close together, a computer can tell they are about the same problem without them sharing a single word.
This guide covers what an embedding actually is, why it was a breakthrough rather than a curiosity, what the numbers mean, and — the part most explanations skip — the three specific places embeddings fail badly enough to matter.
What an Embedding Actually Is
Take any text: a word, a sentence, a support ticket, a whole page. Run it through a trained model. What comes back is a fixed-length list of numbers, the same length every time regardless of how long the input was.

That list is called a vector. People use "embedding" and "vector" almost interchangeably, which is fine in practice.
The critical point is in the band above, and it is where most intuitions go wrong: no individual number means anything. You cannot inspect position 400 and discover that it encodes "formality" or "is about animals". The meaning is distributed across all of them, and it exists only relative to where every other piece of text sits. A single embedding in isolation is useless. A space full of them is powerful.
Why "Close Together" Means "Similar"
Picture a map. Cities near each other on the map are near each other in reality — the map's geometry encodes a real relationship. An embedding space does the same thing for meaning, except with hundreds or thousands of dimensions instead of two.
The relationships that emerge are more structured than simple clustering. The foundational demonstration came from Tomáš Mikolov, Wen-tau Yih and Geoffrey Zweig in 2013, who found that word vectors capture regularities as consistent directions through the space.

Their paper reported that the vectors were "surprisingly good at capturing syntactic and semantic regularities in language", answering nearly 40% of syntactic analogy questions correctly, and that "King – Man + Woman" produces a vector very close to "Queen".
Nobody wrote a rule about gender. The model learned it as a direction because of how those words are used in text. That is the same statistical-pattern mechanism underlying how large language models are trained — embeddings are the earlier, simpler expression of it.
Why This Was a Breakthrough, Not a Curiosity
Word-level vectors were interesting. The shift that made embeddings infrastructure was learning to embed whole sentences efficiently.
Here is the problem it solved, with real numbers. Early transformer models could judge whether two sentences were similar, but only by processing both together, as a pair. Nils Reimers and Iryna Gurevych spelled out the cost in their 2019 Sentence-BERT paper: finding the most similar pair in a collection of just 10,000 sentences required "about 50 million inference computations (~65 hours)" with BERT.
Sixty-five hours. For ten thousand sentences — a small corpus by any standard.
Sentence-BERT changed the shape of the problem by embedding each sentence once, independently, so that comparison became simple geometry on pre-computed vectors. Their reported result: the same task drops "from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT."
That is the moment semantic search became practical. You embed your documents once, store the vectors, and every subsequent query becomes a fast geometric lookup rather than a fresh comparison against every document.
How Many Numbers Are We Talking About?
Vector length — the dimension — is the main knob. More dimensions give the model more room to encode distinctions; they also cost more to store and compare.
Current production models sit in a familiar range. OpenAI's documentation gives text-embedding-3-small a default of 1,536 dimensions and text-embedding-3-large 3,072.
What is less widely known is that these can be shortened. The same documentation notes that developers can "shorten embeddings (i.e. remove some numbers from the end of the sequence) without the embedding losing its concept-representing properties" — and gives a striking example: a text-embedding-3-large embedding "can be shortened to a size of 256 while still outperforming an unshortened text-embedding-ada-002 embedding with a size of 1536."
A newer model at one sixth the vector length beating an older model at full length. If you are sizing a vector database, that is worth knowing before you provision storage.
What Embeddings Are Actually Used For
Where Embeddings Show Up
Five jobs, one mechanism — compare a vector against other vectors.
| Job | What gets compared | Why a vector helps |
|---|---|---|
| JobSemantic search | What gets comparedQuery against documents | Why a vector helpsFinds matches with no shared words |
| JobRAG | What gets comparedQuestion against your knowledge base | Why a vector helpsRetrieves context before the model answers |
| JobRecommendations | What gets comparedOne item against all items | Why a vector helpsSimilarity without hand-written rules |
| JobDeduplication | What gets comparedEverything against everything | Why a vector helpsCatches rewordings exact matching misses |
| JobClustering | What gets comparedItems against each other | Why a vector helpsGroups by topic with no labels needed |
The second row is the one most people meet first. Retrieval-augmented generation works by embedding your documents, embedding the user's question, retrieving the nearest chunks, and handing them to a language model as context. The retrieval quality of a RAG system is mostly embedding quality — which is why choosing between graph and vector retrieval is a real architectural decision rather than a detail.
It is also why RAG exists at all. A model can only consider what fits in its context window, so something has to decide which few thousand words out of your millions are worth including. Embeddings are how that decision gets made.
Where Embeddings Fail
This is the section that separates a working system from a demo. Embeddings have three well-documented weaknesses, and none of them is obvious from a successful first test.
They Are Close to Blind to Numbers
An embedding compresses a passage into a fixed-length vector. Precise values are among the first things to blur.
A 2026 study evaluating 13 widely used text embedding models measured this directly. The authors report that on their numerical-retrieval task, "the retrieval accuracy, averaged over the 13 embedding models under evaluation, is merely 0.54, slightly above random guessing (0.5)."
That number needs its context to land: the task was a binary choice between two candidate passages differing only in their numbers, so 0.5 is what blind guessing scores. The models managed 0.54.
If your documents are specifications, prices, dosages or dates, pure embedding search will not reliably tell "18 months" from "8 months".
There Is a Mathematical Ceiling
This one is structural rather than a training shortfall. A 2026 paper on the theoretical limitations of embedding-based retrieval shows that "the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding."
In plain terms: for a given vector length, there are combinations of documents the system cannot return together, no matter how good the model is. The authors built a deliberately simple stress-test dataset called LIMIT and observed that "even state-of-the-art models fail on this dataset despite the simple nature of the task."
Better training does not remove this. More dimensions raise the ceiling; they do not abolish it.
No Single Model Wins
It is tempting to take the top of a leaderboard and move on. The benchmark designed to test exactly that says otherwise.
MTEB, the Massive Text Embedding Benchmark, "spans 8 embedding tasks covering a total of 58 datasets and 112 languages". Its headline finding is blunt: "no particular text embedding method dominates across all tasks."
A model that leads on classification may trail on retrieval. One strong in English may be weak in the language you actually serve. This is the same trap as reading model leaderboards generally, which we cover in how AI benchmarks mislead — a single ranked number flattens away the thing you needed to know.
How to Choose a Model Without Guessing
A workable sequence, in order of how much time it costs:
- Start with a strong general model at default dimensions. Do not optimise before you have a baseline.
- Build a tiny evaluation set — 30 to 50 real queries from your own domain, each paired with the document you know should come back. This is the step people skip and the one that decides everything.
- Measure how often the right document appears in the top five results. That single percentage is your real benchmark.
- Only then consider alternatives — a different model, shorter dimensions for cost, or adding keyword search alongside.
- Re-check when your content changes. A model tuned on general web text may do poorly on legal clauses or clinical notes.
Step two is non-negotiable. Your 30 real queries tell you more than any public leaderboard, because they test the distribution you actually serve rather than an average across 112 languages.
The Bottom Line
An embedding converts text into a fixed-length list of numbers positioned so that similar meanings land near each other. That one property makes meaning comparable by arithmetic, which is what made semantic search, recommendations and RAG practical — Sentence-BERT's 65 hours to 5 seconds is the clearest measure of the shift.
But treat them as a strong default rather than a complete answer. They are close to blind to numbers, they carry a dimension-bound ceiling on what they can retrieve, and no single model leads everywhere. In practice that means embeddings usually belong alongside older techniques rather than replacing them — which is exactly the argument in our comparison of semantic versus keyword search.
For the wider picture of how these systems fit together, see our LLM hub.
Frequently Asked Questions
What is an embedding in simple terms?
It is a list of numbers that stands in for a piece of text, built so that texts with similar meanings produce similar lists. Because the lists are numbers, a computer can measure how close two pieces of text are without understanding either one. The numbers have no individual meaning — only the position relative to all other texts carries information.
What is the difference between an embedding and a vector?
In everyday use, none. "Vector" describes the mathematical object — an ordered list of numbers. "Embedding" describes what it is for — representing something, usually text, inside a meaning space. People say "we store the embeddings in a vector database", and both words point at the same list of numbers.
How many dimensions should an embedding have?
Common production models default to 1,536 or 3,072. More dimensions allow finer distinctions but cost more to store and compare, and OpenAI's documentation notes that a large model shortened to 256 dimensions can still outperform an older model at 1,536. Start with the default, measure on your own queries, and shorten only if storage or latency is an actual constraint.
Are embeddings good at handling numbers?
No, and this is a common source of production failures. A study across 13 widely used embedding models found retrieval accuracy of just 0.54 on a numerical task where random guessing scores 0.5. If your content turns on precise values — prices, dosages, specifications, dates — pair embeddings with exact matching rather than relying on them alone.
Which embedding model is best?
There is no single answer, and that is a measured finding rather than a hedge. MTEB, which spans 8 tasks across 58 datasets and 112 languages, reports that no embedding method dominates across all tasks. Build a small evaluation set of 30 to 50 real queries from your own domain and test candidates against it.
Do embeddings replace keyword search?
They should not. Embeddings find matches that share meaning but no words; keyword search nails exact identifiers, error codes and rare names that embeddings blur. Most production systems run both and merge the rankings, because the two approaches fail in opposite directions.
Sources
- Reimers & Gurevych — Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (arXiv:1908.10084)
- Mikolov, Yih & Zweig — Linguistic Regularities in Continuous Space Word Representations (NAACL 2013)
- Muennighoff et al. — MTEB: Massive Text Embedding Benchmark (arXiv:2210.07316)
- On the Theoretical Limitations of Embedding-Based Retrieval (arXiv:2508.21038)
- Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models (arXiv:2509.05691)
- OpenAI — Embeddings, official API documentation



