Skip to main content
AI & Technology

Semantic vs Keyword Search: Which One Should You Use?

Semantic search understands meaning; keyword search matches words. The research shows neither wins outright — and the reason explains why production systems run both.

12 min read
Share:
Dark navy split illustration with stacked word tiles labelled exact words on the left and a linked constellation of nodes labelled nearby meaning on the right
Credit: PrimusSource (original illustration)

Keyword search matches the words you typed. Semantic search matches what you meant. The obvious conclusion is that the second one replaced the first — and the benchmark evidence says that conclusion is wrong.

In the most widely used evaluation of retrieval systems across unfamiliar domains, the decades-old keyword algorithm BM25 is still described as "a robust baseline", while the newer dense retrieval models "often underperform" it. That result has held up well enough that essentially every serious production search system now runs both methods and merges the results.

This comparison covers how each one actually works, the specific failures of each, what the research says about when the newer approach loses, and how to decide for your own system.

The Short Answer

Keyword searchSemantic search
Matches onThe words presentThe meaning conveyed
Core methodTerm frequency, e.g. BM25Vector distance between embeddings
Wins atExact IDs, codes, rare names, numbersParaphrase, synonyms, natural questions
Fails atAny wording you did not anticipatePrecise values and exact strings
Setup costLow — index the textHigher — embed, store, maintain vectors
Behaves predictably?Yes, you can explain every hitLess so, relevance is a distance

If you only take one thing: they fail in opposite directions, which is precisely why combining them works so well.

How Each One Actually Works

Keyword search scores documents on the words they share with the query, weighting rare words more heavily than common ones and adjusting for document length. BM25, the standard implementation, dates from the 1990s and remains the default in Elasticsearch, OpenSearch and most search infrastructure.

Semantic search converts both the query and the documents into embeddings — fixed-length lists of numbers positioned so that similar meanings sit close together — and returns whatever is nearest. No words need to be shared at all.

Two stacked cards comparing how each method handles the query laptop keeps shutting down when unplugged. The amber card, keyword search matches words, finds pages containing laptop, shutting and unplugged, misses notebook powers off on battery entirely, but nails an exact model number or error code. The blue card, semantic search matches meaning, finds the battery page even with no shared words, ranks by closeness in the embedding space, but can drift to pages that merely sound related. A green band notes that neither one is the better machine because they fail in opposite directions, which is why production systems run both
One query, two mechanisms. Note that each one's weakness is the other's strength.

The example in the figure is the whole argument in miniature. A user searching "laptop keeps shutting down when unplugged" wants a page that may well be titled "notebook powers off on battery" — no shared words, same problem. Keyword search misses it. Semantic search finds it.

Reverse the case. A user searches for error code 0x8007000E. Semantic search has no useful notion of what that string means and will happily return pages about other error codes that look similar. Keyword search returns the exact page, every time.

Where Semantic Search Loses

This is the part most comparisons underplay, so it is worth being specific. There are three documented failures, and all three are structural rather than fixable with a better model.

Out of Its Training Domain

BEIR is the standard benchmark for retrieval across domains a model was not tuned on. It uses "a careful selection of 18 publicly available datasets from diverse text retrieval tasks and domains", and its conclusion is direct: "BM25 is a robust baseline", while "dense and sparse-retrieval models are computationally more efficient but often underperform other approaches."

The practical reading: an embedding model tested on general web text can do noticeably worse on your legal clauses, clinical notes or internal jargon than a keyword index you could set up in an afternoon.

Numbers

Embeddings compress text into a fixed-length vector, and precise values blur early. A 2026 study across 13 widely used embedding models found that on a numerical retrieval task, "the retrieval accuracy, averaged over the 13 embedding models under evaluation, is merely 0.54, slightly above random guessing (0.5)."

The task was a binary choice between passages differing only in their numbers, so 0.5 is the score for guessing. If your catalogue turns on prices, capacities, dosages or model years, that is close to a coin flip.

A Hard Mathematical Ceiling

The third limit is not about training at all. A 2026 paper on the theoretical limitations of embedding-based retrieval shows that "the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding."

For any fixed vector length there are combinations of documents the system simply cannot return together. The authors built a deliberately easy stress test called LIMIT and found that "even state-of-the-art models fail on this dataset despite the simple nature of the task."

Where Keyword Search Loses

In fairness, its failures are just as real — they are only more familiar, because we have spent thirty years learning to work around them.

How Each One Fails

Both failure modes are predictable, which is what makes them manageable.

Keyword search misses

  1. Any synonym you did not anticipate
  2. Questions phrased conversationally
  3. Related concepts with different words
  4. Other languages entirely
  5. Spelling the user got wrong

Fails when wording varies

Semantic search misses

  1. Exact codes, SKUs and identifiers
  2. Precise numbers and specifications
  3. Rare proper nouns it never learned
  4. Negation — with and without
  5. Anything outside its training domain

Fails when precision matters

The two lists barely overlap

That is the case for running both. A query that defeats one is usually handled comfortably by the other.

The classic keyword failure is the synonym gap. A support site documenting "intermittent connectivity loss" will not surface for a user searching "wifi keeps dropping", however well the page answers the question. Teams historically patched this with manually maintained synonym lists — which works, and scales badly.

Hybrid Search: What Production Systems Actually Do

Because the two approaches fail in opposite directions, running both and merging the results outperforms either alone. The standard merge technique is older than most people assume.

Four numbered steps showing how hybrid search combines the two methods: send the query to both a keyword index and a vector index; each returns its own ranking, two ordered lists that often disagree; fuse by position rather than by score, where Reciprocal Rank Fusion adds up one divided by rank; and one merged list comes out, where what both methods liked rises to the top. A green band explains that a keyword score and a cosine similarity are not on the same scale and never will be, whereas positions are comparable
Reciprocal Rank Fusion was published in 2009. It now underpins hybrid retrieval in most search engines and RAG stacks.

The key insight is in the band. A BM25 score and a cosine similarity are different kinds of number — one is unbounded and corpus-dependent, the other sits between −1 and 1. Averaging them is meaningless. Ranks, however, are directly comparable: being third in a list means the same thing in both systems.

Reciprocal Rank Fusion, introduced by Cormack, Clarke and Buettcher in 2009, exploits exactly that. Each document scores the sum of 1 divided by its rank in each list, so a document placed respectably by both systems beats one placed first by a single system and nowhere by the other. The paper's title states its own finding: reciprocal rank fusion outperforms Condorcet and individual rank-learning methods.

The Third Stage Most Comparisons Skip

Framing this as a two-way choice leaves out the step that often matters most for quality: re-ranking.

Both methods above are built for speed. They have to be, because they compare the query against the entire collection. That constraint forces a compromise — the query and each document are processed separately, and only their finished representations are compared.

A more accurate approach exists. Models called cross-encoders read the query and a document together and judge the match directly, which captures interactions the separate approach cannot. The catch is the one Sentence-BERT quantified: doing that across a whole collection takes about 50 million inference computations, roughly 65 hours, for just 10,000 items. Unusable as a search method.

The resolution is to use each where it belongs. Fast retrieval — keyword, semantic, or both fused — narrows millions of documents to perhaps the top 50. The slow, accurate cross-encoder then re-scores only those 50 and reorders them. You pay the expensive computation 50 times instead of a million times.

If your search returns roughly the right documents but in a frustrating order, re-ranking is usually a bigger win than swapping the embedding model — and it is a far smaller change.

Which Should You Choose?

Pick by What Your Users Type

The deciding factor is query shape, not which technology is newer.

If your queries are mostlyProduct codes, IDs, error stringsStart withKeywordWhyExact matching is the requirement
If your queries are mostlyNatural questions in varied wordingStart withSemanticWhyThe synonym gap is your main loss
If your queries are mostlySpecs, prices, dates, capacitiesStart withKeywordWhyEmbeddings score near chance on numbers
If your queries are mostlyA specialist domain with its own jargonStart withKeyword firstWhyDense models often underperform out of domain
If your queries are mostlyA mix of all of the aboveStart withHybridWhyEach covers the other gap

If you are building RAG, this is a retrieval decision

The answer quality of a RAG system is capped by what retrieval hands the model. Fix retrieval before blaming the model.

Most real systems land on the last row — but start simple and add the second method when you can show it is needed.

A practical order of work:

  1. Start with keyword search. It is cheap, predictable, and it gives you a baseline you can explain.
  2. Collect 30 to 50 real failing queries. Not invented ones — actual searches that returned nothing useful.
  3. Check what kind of failure they are. If they are synonym and phrasing failures, semantic search will help. If they are exact-match failures, it will not.
  4. Add semantic retrieval alongside, not instead. Keep both indexes.
  5. Fuse by rank, not by score, for the reason above.

That third step is the one that saves money. Teams routinely add a vector database for a problem that was actually a tokenisation or synonym issue, and end up with two systems where a configuration change would have done.

The Bottom Line

Semantic search is a genuine advance, and for conversational queries it solves a problem keyword search never could. But it did not make keyword search obsolete, and the benchmark evidence is unusually clear on this: BM25 remains a robust baseline across unfamiliar domains, embeddings score barely above chance on numerical detail, and there is a mathematical ceiling on what any single vector can retrieve.

Treat the choice as a question about your queries rather than about which technology is newer. If users type identifiers and specifications, keyword search is not a legacy system, it is the right tool. If they type questions, semantic retrieval earns its place. If they type both — which is most real products — run both and fuse the rankings.

For the mechanism underneath semantic retrieval, see our guide to embeddings; for how retrieval feeds a language model, see what RAG is and GraphRAG versus vector RAG. More on how these systems fit together is in our LLM hub.

Frequently Asked Questions

What is the difference between semantic and keyword search?

Keyword search ranks documents by the words they share with the query, weighting rare words more heavily. Semantic search converts the query and documents into embeddings — lists of numbers representing meaning — and returns whichever documents sit closest. The practical difference is that semantic search can match a page with no words in common with the query, while keyword search cannot.

Is semantic search better than keyword search?

Not outright, which surprises people. On BEIR, the standard benchmark covering 18 datasets across unfamiliar domains, the keyword algorithm BM25 is reported as a robust baseline while dense retrieval models often underperform. Semantic search is better for varied phrasing; keyword search is better for exact strings and numbers.

What is hybrid search?

Running both a keyword index and a vector index over the same content, then merging the two result lists into one ranking. It is the standard production approach because the two methods fail in opposite directions, so a query that defeats one is usually handled well by the other.

What is Reciprocal Rank Fusion?

A method for merging ranked lists from different search systems, published by Cormack, Clarke and Buettcher in 2009. Each document is scored by summing 1 divided by its rank in each list. It uses positions rather than scores because a keyword relevance score and a cosine similarity are not on the same scale and cannot meaningfully be averaged.

Why does semantic search struggle with numbers?

Because an embedding compresses a passage into a fixed-length vector, and precise values are among the first details lost. A study across 13 widely used embedding models measured retrieval accuracy of 0.54 on a numerical task where random guessing scores 0.5 — close to a coin flip. For prices, specifications or dates, pair semantic retrieval with exact matching.

Do I need a vector database to do semantic search?

For a small collection, no — you can compute distances directly over stored vectors. A dedicated vector database earns its place when you have enough documents that exhaustive comparison becomes slow, since it uses approximate nearest-neighbour indexes to avoid comparing against everything. Start simple and add one when scale demands it.

Sources

Artificial Intelligence & LLMs#artificial intelligence#vector database#rag#large language models#machine learning
Share:
A workshop wall neatly racked with hand tools — chisels, files, hand planes, saws and pliers arranged in rows on wooden holders

AI & TechnologyComparison

Fine-Tuning vs RAG vs Prompting: Which Do You Actually Need?

They are not three ways to do the same thing — they fix three different failures. A diagnostic guide to choosing between them, what the head-to-head study really showed, and why 2026 narrowed the case for fine-tuning.

Sep 24, 202612 min