Keyword search matches the words you typed. Semantic search matches what you meant. The obvious conclusion is that the second one replaced the first — and the benchmark evidence says that conclusion is wrong.
In the most widely used evaluation of retrieval systems across unfamiliar domains, the decades-old keyword algorithm BM25 is still described as "a robust baseline", while the newer dense retrieval models "often underperform" it. That result has held up well enough that essentially every serious production search system now runs both methods and merges the results.
This comparison covers how each one actually works, the specific failures of each, what the research says about when the newer approach loses, and how to decide for your own system.
The Short Answer
| Keyword search | Semantic search | |
|---|---|---|
| Matches on | The words present | The meaning conveyed |
| Core method | Term frequency, e.g. BM25 | Vector distance between embeddings |
| Wins at | Exact IDs, codes, rare names, numbers | Paraphrase, synonyms, natural questions |
| Fails at | Any wording you did not anticipate | Precise values and exact strings |
| Setup cost | Low — index the text | Higher — embed, store, maintain vectors |
| Behaves predictably? | Yes, you can explain every hit | Less so, relevance is a distance |
If you only take one thing: they fail in opposite directions, which is precisely why combining them works so well.
How Each One Actually Works
Keyword search scores documents on the words they share with the query, weighting rare words more heavily than common ones and adjusting for document length. BM25, the standard implementation, dates from the 1990s and remains the default in Elasticsearch, OpenSearch and most search infrastructure.
Semantic search converts both the query and the documents into embeddings — fixed-length lists of numbers positioned so that similar meanings sit close together — and returns whatever is nearest. No words need to be shared at all.

The example in the figure is the whole argument in miniature. A user searching "laptop keeps shutting down when unplugged" wants a page that may well be titled "notebook powers off on battery" — no shared words, same problem. Keyword search misses it. Semantic search finds it.
Reverse the case. A user searches for error code 0x8007000E. Semantic search has no useful notion of what that string means and will happily return pages about other error codes that look similar. Keyword search returns the exact page, every time.
Where Semantic Search Loses
This is the part most comparisons underplay, so it is worth being specific. There are three documented failures, and all three are structural rather than fixable with a better model.
Out of Its Training Domain
BEIR is the standard benchmark for retrieval across domains a model was not tuned on. It uses "a careful selection of 18 publicly available datasets from diverse text retrieval tasks and domains", and its conclusion is direct: "BM25 is a robust baseline", while "dense and sparse-retrieval models are computationally more efficient but often underperform other approaches."
The practical reading: an embedding model tested on general web text can do noticeably worse on your legal clauses, clinical notes or internal jargon than a keyword index you could set up in an afternoon.
Numbers
Embeddings compress text into a fixed-length vector, and precise values blur early. A 2026 study across 13 widely used embedding models found that on a numerical retrieval task, "the retrieval accuracy, averaged over the 13 embedding models under evaluation, is merely 0.54, slightly above random guessing (0.5)."
The task was a binary choice between passages differing only in their numbers, so 0.5 is the score for guessing. If your catalogue turns on prices, capacities, dosages or model years, that is close to a coin flip.
A Hard Mathematical Ceiling
The third limit is not about training at all. A 2026 paper on the theoretical limitations of embedding-based retrieval shows that "the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding."
For any fixed vector length there are combinations of documents the system simply cannot return together. The authors built a deliberately easy stress test called LIMIT and found that "even state-of-the-art models fail on this dataset despite the simple nature of the task."
Where Keyword Search Loses
In fairness, its failures are just as real — they are only more familiar, because we have spent thirty years learning to work around them.
How Each One Fails
Both failure modes are predictable, which is what makes them manageable.
Keyword search misses
- Any synonym you did not anticipate
- Questions phrased conversationally
- Related concepts with different words
- Other languages entirely
- Spelling the user got wrong
Fails when wording varies
Semantic search misses
- Exact codes, SKUs and identifiers
- Precise numbers and specifications
- Rare proper nouns it never learned
- Negation — with and without
- Anything outside its training domain
Fails when precision matters
The two lists barely overlap
That is the case for running both. A query that defeats one is usually handled comfortably by the other.
The classic keyword failure is the synonym gap. A support site documenting "intermittent connectivity loss" will not surface for a user searching "wifi keeps dropping", however well the page answers the question. Teams historically patched this with manually maintained synonym lists — which works, and scales badly.
Hybrid Search: What Production Systems Actually Do
Because the two approaches fail in opposite directions, running both and merging the results outperforms either alone. The standard merge technique is older than most people assume.

The key insight is in the band. A BM25 score and a cosine similarity are different kinds of number — one is unbounded and corpus-dependent, the other sits between −1 and 1. Averaging them is meaningless. Ranks, however, are directly comparable: being third in a list means the same thing in both systems.
Reciprocal Rank Fusion, introduced by Cormack, Clarke and Buettcher in 2009, exploits exactly that. Each document scores the sum of 1 divided by its rank in each list, so a document placed respectably by both systems beats one placed first by a single system and nowhere by the other. The paper's title states its own finding: reciprocal rank fusion outperforms Condorcet and individual rank-learning methods.
The Third Stage Most Comparisons Skip
Framing this as a two-way choice leaves out the step that often matters most for quality: re-ranking.
Both methods above are built for speed. They have to be, because they compare the query against the entire collection. That constraint forces a compromise — the query and each document are processed separately, and only their finished representations are compared.
A more accurate approach exists. Models called cross-encoders read the query and a document together and judge the match directly, which captures interactions the separate approach cannot. The catch is the one Sentence-BERT quantified: doing that across a whole collection takes about 50 million inference computations, roughly 65 hours, for just 10,000 items. Unusable as a search method.
The resolution is to use each where it belongs. Fast retrieval — keyword, semantic, or both fused — narrows millions of documents to perhaps the top 50. The slow, accurate cross-encoder then re-scores only those 50 and reorders them. You pay the expensive computation 50 times instead of a million times.
If your search returns roughly the right documents but in a frustrating order, re-ranking is usually a bigger win than swapping the embedding model — and it is a far smaller change.
Which Should You Choose?
Pick by What Your Users Type
The deciding factor is query shape, not which technology is newer.
| If your queries are mostly | Start with | Why |
|---|---|---|
| If your queries are mostlyProduct codes, IDs, error strings | Start withKeyword | WhyExact matching is the requirement |
| If your queries are mostlyNatural questions in varied wording | Start withSemantic | WhyThe synonym gap is your main loss |
| If your queries are mostlySpecs, prices, dates, capacities | Start withKeyword | WhyEmbeddings score near chance on numbers |
| If your queries are mostlyA specialist domain with its own jargon | Start withKeyword first | WhyDense models often underperform out of domain |
| If your queries are mostlyA mix of all of the above | Start withHybrid | WhyEach covers the other gap |
If you are building RAG, this is a retrieval decision
The answer quality of a RAG system is capped by what retrieval hands the model. Fix retrieval before blaming the model.
A practical order of work:
- Start with keyword search. It is cheap, predictable, and it gives you a baseline you can explain.
- Collect 30 to 50 real failing queries. Not invented ones — actual searches that returned nothing useful.
- Check what kind of failure they are. If they are synonym and phrasing failures, semantic search will help. If they are exact-match failures, it will not.
- Add semantic retrieval alongside, not instead. Keep both indexes.
- Fuse by rank, not by score, for the reason above.
That third step is the one that saves money. Teams routinely add a vector database for a problem that was actually a tokenisation or synonym issue, and end up with two systems where a configuration change would have done.
The Bottom Line
Semantic search is a genuine advance, and for conversational queries it solves a problem keyword search never could. But it did not make keyword search obsolete, and the benchmark evidence is unusually clear on this: BM25 remains a robust baseline across unfamiliar domains, embeddings score barely above chance on numerical detail, and there is a mathematical ceiling on what any single vector can retrieve.
Treat the choice as a question about your queries rather than about which technology is newer. If users type identifiers and specifications, keyword search is not a legacy system, it is the right tool. If they type questions, semantic retrieval earns its place. If they type both — which is most real products — run both and fuse the rankings.
For the mechanism underneath semantic retrieval, see our guide to embeddings; for how retrieval feeds a language model, see what RAG is and GraphRAG versus vector RAG. More on how these systems fit together is in our LLM hub.
Frequently Asked Questions
What is the difference between semantic and keyword search?
Keyword search ranks documents by the words they share with the query, weighting rare words more heavily. Semantic search converts the query and documents into embeddings — lists of numbers representing meaning — and returns whichever documents sit closest. The practical difference is that semantic search can match a page with no words in common with the query, while keyword search cannot.
Is semantic search better than keyword search?
Not outright, which surprises people. On BEIR, the standard benchmark covering 18 datasets across unfamiliar domains, the keyword algorithm BM25 is reported as a robust baseline while dense retrieval models often underperform. Semantic search is better for varied phrasing; keyword search is better for exact strings and numbers.
What is hybrid search?
Running both a keyword index and a vector index over the same content, then merging the two result lists into one ranking. It is the standard production approach because the two methods fail in opposite directions, so a query that defeats one is usually handled well by the other.
What is Reciprocal Rank Fusion?
A method for merging ranked lists from different search systems, published by Cormack, Clarke and Buettcher in 2009. Each document is scored by summing 1 divided by its rank in each list. It uses positions rather than scores because a keyword relevance score and a cosine similarity are not on the same scale and cannot meaningfully be averaged.
Why does semantic search struggle with numbers?
Because an embedding compresses a passage into a fixed-length vector, and precise values are among the first details lost. A study across 13 widely used embedding models measured retrieval accuracy of 0.54 on a numerical task where random guessing scores 0.5 — close to a coin flip. For prices, specifications or dates, pair semantic retrieval with exact matching.
Do I need a vector database to do semantic search?
For a small collection, no — you can compute distances directly over stored vectors. A dedicated vector database earns its place when you have enough documents that exhaustive comparison becomes slow, since it uses approximate nearest-neighbour indexes to avoid comparing against everything. Start simple and add one when scale demands it.
Sources
- Thakur et al. — BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (arXiv:2104.08663)
- Cormack, Clarke & Buettcher — Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods (SIGIR 2009)
- On the Theoretical Limitations of Embedding-Based Retrieval (arXiv:2508.21038)
- Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models (arXiv:2509.05691)
- Reimers & Gurevych — Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (arXiv:1908.10084)



