Writing · GenAI · 8 min read
Retrieval is the hard part
When a RAG system gives a bad answer, the model usually takes the blame. More often it was handed the wrong context. Embeddings, chunking, approximate search, hybrid ranking and reranking, seen from the retrieval side.
Retrieval-augmented generation is a simple idea. Before asking a language model a question, find the passages most likely to contain the answer and put them in front of it. The model writes the answer; retrieval decides what it gets to read.
I've shipped production GenAI applications that do semantic search and summarisation over domain-specific data, built on LLMs and vector databases like FAISS and Qdrant. The lesson I'd pass on first is that when an answer is bad, the model is rarely the place to start looking. It can only work with the context it was given, and a bad answer usually traces back to bad retrieval: the right passage wasn't found, or it was found alongside so much noise that the model followed the wrong thread.
The first retrieval system I built predates the name. As an intern at InterviewBit I built a tool that measured the semantic similarity between a student's question and solutions that already existed, and suggested the closest. For 72% of students, it cut the time to get a query resolved from two days to 13.3 minutes on average. Most of what follows I've relearned at larger scale since.
1What "similar" means to a machine
Semantic search starts with embeddings. An embedding model turns a piece of text into a long list of numbers, a point in a space with hundreds or thousands of dimensions, arranged so that texts with similar meanings land near each other. The question is embedded the same way, and "find the relevant passages" becomes "find the nearest points", usually by the angle between vectors: cosine similarity.
returned –best match –
Flattened to two dimensions it's a cartoon, but the behaviour is real. Retrieval returns the nearest neighbours whether or not any of them is relevant. Put the query between topics and the results become a mixture. Ask about something the corpus doesn't contain and you still get k confident-looking passages, because nothing in the vector maths says "I don't know". That has to be designed in, with a similarity threshold or a reranker that is allowed to return nothing.
2Chunking decides what can be found
Documents are too long to embed whole, so they're split into chunks, and the chunk size quietly sets a ceiling on retrieval quality. An embedding is a kind of average of what its text is about. Make chunks large and each one covers several topics, so its embedding blurs and a precise question matches it only weakly. Make them small and each chunk is sharp but short of context: the sentence that answers the question often doesn't say what it's about, because the sentence before it did.
retrieved chunk has the answer –relevant text in it –
There's no universally right size. It depends on how the documents are written and what people ask of them. Splitting along the document's own structure (sections, paragraphs, table rows) usually beats cutting every N tokens, and a little overlap between chunks stops an answer from being severed from the sentence that gives it meaning. The only reliable way to choose is to measure on real questions.
3Finding neighbours without checking everyone
Comparing the query with every stored vector is exact, and fine for a few thousand chunks. At millions it's too slow, so vector databases use approximate nearest-neighbour indexes. One of the most widely used, and the one Qdrant is built around, is HNSW: a hierarchical navigable small-world graph. FAISS offers it too, alongside other index types.
HNSW links every vector to a handful of its neighbours, then builds sparser layers on top, like express lanes. A search starts in the top layer and hops greedily towards the query until no neighbour is closer, then drops a layer and repeats. By the bottom layer it's in the right neighbourhood, having measured the distance to only a small fraction of the collection.
distances computed –checking everything –found the true nearest –
The approximation is a dial, not a flaw. The number of links per node and how widely the search explores trade speed and memory for recall. In Qdrant they're m, the links per node, 16 by default; ef_construct, how widely the graph is searched while it's being built, 100 by default; and hnsw_ef, how widely each query searches. It's worth measuring recall against exact search on your own data before trusting any of them, and Qdrant will run any query exactly if asked, so the measurement is one comparison:
approx = client.query_points("passages", query=vec, limit=10,
search_params=models.SearchParams(hnsw_ef=128)).points
exact = client.query_points("passages", query=vec, limit=10,
search_params=models.SearchParams(exact=True)).points
recall = len({p.id for p in approx} & {p.id for p in exact}) / len(exact)
4Words still matter
Embeddings are good at meaning and poor at exact strings. Domain data is full of identifiers, such as gene symbols, compound IDs and trial numbers, where the exact string is the whole point, and the embedding of "ERBB2" may sit uncomfortably close to "ERBB3". Keyword search, usually BM25, has the opposite strengths: it matches strings exactly and knows nothing about meaning. An embedding model may well have learned that HER2 is the protein the ERBB2 gene encodes; a keyword index won't know it unless someone tells it.
So run both and merge the lists. The merging is the subtle part, because the two scores can't be compared: a BM25 score is unbounded and depends on the corpus, while cosine similarity lives between −1 and 1. Reciprocal rank fusion sidesteps the problem by ignoring scores altogether and using only positions. Each list gives a passage 1/(k + rank), with k customarily 60, and the sums decide the order. A passage both lists rank well beats one that only one list loves, which is usually what you want.
right passages in the top 3: keywords –vectors –fused –
def rrf(*rankings, k=60):
scores = {}
for ranking in rankings: # each a list of ids, best first
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
fused = rrf(bm25_ids, vector_ids)
5Filter first
If the answer has to come from one organism, one study or one customer's documents, that isn't a similarity question, and it shouldn't be left to similarity. Fetch the nearest ten, throw away the ones from the wrong study, and you may be left with one passage, or none. Qdrant applies the filter inside the vector search, so the ten that come back are the nearest ten that match:
hits = client.query_points(
"passages", query=vec, limit=10,
query_filter=models.Filter(must=[
models.FieldCondition(key="study", match=models.MatchValue(value="study-42")),
]),
).points
6A second, slower opinion
Embedding search is fast because it never reads a question and a passage together: each was turned into a vector on its own, long before they met. A reranker does read them together. A cross-encoder takes the question and one passage as a single input and scores how well the passage answers it, which is far more precise, and far too slow to run over a whole collection. So the two are chained: fetch a generous set of candidates cheaply, then let the reranker reorder just those.
The catch is in the word "just". A reranker can only reorder what it's given. If the passage that answers the question was 80th by embedding similarity and only the top 20 were fetched, no reranker will ever see it. How many candidates to fetch is a trade between the first stage's recall and the second stage's cost, and it deserves a deliberate choice rather than whatever number the example code used.
right passages the model reads –never fetched –reranker calls –
7Measure it
Every choice above, the chunk size, the embedding model, the index settings, how many to fetch, whether to fuse or rerank, moves retrieval quality in a direction you can't see by trying a few questions by hand. What turns them into decisions is a small evaluation set: real questions, each with the passages a person agreed contain its answer. A few dozen is enough to start. Then every change gets measured the same way, and the number that matters most for RAG is recall at k: of the passages that answer the question, how many made it into the k the model will read.
def recall_at_k(retrieved, relevant, k):
return len(set(retrieved[:k]) & relevant) / len(relevant)
def mrr(retrieved, relevant): # 1 / the position of the first right answer
return next((1 / i for i, doc in enumerate(retrieved, 1) if doc in relevant), 0.0)
def evaluate(questions, search, k=5):
runs = [(search(q.text), q.relevant_ids) for q in questions]
return {"recall@k": sum(recall_at_k(r, rel, k) for r, rel in runs) / len(runs),
"mrr": sum(mrr(r, rel) for r, rel in runs) / len(runs)}
8The model is the easy part
It's tempting to treat the language model as the system and retrieval as plumbing. In production it's the other way round. The model is a component you can swap. The retrieval pipeline is where the domain knowledge lives: how documents are split, what gets indexed, which filters apply, and what "relevant" means to the people asking. Get that right and a modest model gives good answers. Get it wrong and the best model in the world will confidently summarise the wrong page.