Rohit Swami
India Resume ↗

Writing · GenAI · 8 min read

Retrieval is the hard part

When a RAG system gives a bad answer, the model usually takes the blame. More often it was handed the wrong context. Embeddings, chunking, approximate search, hybrid ranking and reranking, seen from the retrieval side.

Retrieval-augmented generation is a simple idea. Before asking a language model a question, find the passages most likely to contain the answer and put them in front of it. The model writes the answer; retrieval decides what it gets to read.

I've shipped production GenAI applications that do semantic search and summarisation over domain-specific data, built on LLMs and vector databases like FAISS and Qdrant. The lesson I'd pass on first is that when an answer is bad, the model is rarely the place to start looking. It can only work with the context it was given, and a bad answer usually traces back to bad retrieval: the right passage wasn't found, or it was found alongside so much noise that the model followed the wrong thread.

The first retrieval system I built predates the name. As an intern at InterviewBit I built a tool that measured the semantic similarity between a student's question and solutions that already existed, and suggested the closest. For 72% of students, it cut the time to get a query resolved from two days to 13.3 minutes on average. Most of what follows I've relearned at larger scale since.

1What "similar" means to a machine

Semantic search starts with embeddings. An embedding model turns a piece of text into a long list of numbers, a point in a space with hundreds or thousands of dimensions, arranged so that texts with similar meanings land near each other. The question is embedded the same way, and "find the relevant passages" becomes "find the nearest points", usually by the angle between vectors: cosine similarity.

top k

returned –best match –

Fig. 1 Drag the query. Real embeddings have hundreds of dimensions; here nearness on the page stands in for cosine similarity, which makes it a cartoon, but an honest one. Drag it into the empty space and see what comes back.

Flattened to two dimensions it's a cartoon, but the behaviour is real. Retrieval returns the nearest neighbours whether or not any of them is relevant. Put the query between topics and the results become a mixture. Ask about something the corpus doesn't contain and you still get k confident-looking passages, because nothing in the vector maths says "I don't know". That has to be designed in, with a similarity threshold or a reranker that is allowed to return nothing.

2Chunking decides what can be found

Documents are too long to embed whole, so they're split into chunks, and the chunk size quietly sets a ceiling on retrieval quality. An embedding is a kind of average of what its text is about. Make chunks large and each one covers several topics, so its embedding blurs and a precise question matches it only weakly. Make them small and each chunk is sharp but short of context: the sentence that answers the question often doesn't say what it's about, because the sentence before it did.

chunk size

retrieved chunk has the answer –relevant text in it –

Fig. 2 A toy model. Colours are topics, and the question is about the blue one. The answer is the sentence marked A, which doesn't name its topic; the sentence before it does. Each chunk scores by how much of it is about the question, shown as the bars above it, and the best one is retrieved.

There's no universally right size. It depends on how the documents are written and what people ask of them. Splitting along the document's own structure (sections, paragraphs, table rows) usually beats cutting every N tokens, and a little overlap between chunks stops an answer from being severed from the sentence that gives it meaning. The only reliable way to choose is to measure on real questions.

3Finding neighbours without checking everyone

Comparing the query with every stored vector is exact, and fine for a few thousand chunks. At millions it's too slow, so vector databases use approximate nearest-neighbour indexes. One of the most widely used, and the one Qdrant is built around, is HNSW: a hierarchical navigable small-world graph. FAISS offers it too, alongside other index types.

HNSW links every vector to a handful of its neighbours, then builds sparser layers on top, like express lanes. A search starts in the top layer and hops greedily towards the query until no neighbour is closer, then drops a layer and repeats. By the bottom layer it's in the right neighbourhood, having measured the distance to only a small fraction of the collection.

distances computed –checking everything –found the true nearest –

Fig. 3 Three hundred vectors in three layers. Each plane is the same space; higher layers hold fewer points with longer links. The red path is the search: greedy hops in each layer, then a drop to the one below.

The approximation is a dial, not a flaw. The number of links per node and how widely the search explores trade speed and memory for recall. In Qdrant they're m, the links per node, 16 by default; ef_construct, how widely the graph is searched while it's being built, 100 by default; and hnsw_ef, how widely each query searches. It's worth measuring recall against exact search on your own data before trusting any of them, and Qdrant will run any query exactly if asked, so the measurement is one comparison:

approx = client.query_points("passages", query=vec, limit=10,
                             search_params=models.SearchParams(hnsw_ef=128)).points
exact = client.query_points("passages", query=vec, limit=10,
                            search_params=models.SearchParams(exact=True)).points
recall = len({p.id for p in approx} & {p.id for p in exact}) / len(exact)

4Words still matter

Embeddings are good at meaning and poor at exact strings. Domain data is full of identifiers, such as gene symbols, compound IDs and trial numbers, where the exact string is the whole point, and the embedding of "ERBB2" may sit uncomfortably close to "ERBB3". Keyword search, usually BM25, has the opposite strengths: it matches strings exactly and knows nothing about meaning. An embedding model may well have learned that HER2 is the protein the ERBB2 gene encodes; a keyword index won't know it unless someone tells it.

So run both and merge the lists. The merging is the subtle part, because the two scores can't be compared: a BM25 score is unbounded and depends on the corpus, while cosine similarity lives between −1 and 1. Reciprocal rank fusion sidesteps the problem by ignoring scores altogether and using only positions. Each list gives a passage 1/(k + rank), with k customarily 60, and the sums decide the order. A passage both lists rank well beats one that only one list loves, which is usually what you want.

right passages in the top 3: keywords –vectors –fused –

Fig. 4 The rankings are an illustration; the fusion is the real formula, computed as you watch. Blue passages are the ones that answer the question. Hover a passage to see its two ranks and its fused score.
def rrf(*rankings, k=60):
    scores = {}
    for ranking in rankings:                       # each a list of ids, best first
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

fused = rrf(bm25_ids, vector_ids)

5Filter first

If the answer has to come from one organism, one study or one customer's documents, that isn't a similarity question, and it shouldn't be left to similarity. Fetch the nearest ten, throw away the ones from the wrong study, and you may be left with one passage, or none. Qdrant applies the filter inside the vector search, so the ten that come back are the nearest ten that match:

hits = client.query_points(
    "passages", query=vec, limit=10,
    query_filter=models.Filter(must=[
        models.FieldCondition(key="study", match=models.MatchValue(value="study-42")),
    ]),
).points

6A second, slower opinion

Embedding search is fast because it never reads a question and a passage together: each was turned into a vector on its own, long before they met. A reranker does read them together. A cross-encoder takes the question and one passage as a single input and scores how well the passage answers it, which is far more precise, and far too slow to run over a whole collection. So the two are chained: fetch a generous set of candidates cheaply, then let the reranker reorder just those.

The catch is in the word "just". A reranker can only reorder what it's given. If the passage that answers the question was 80th by embedding similarity and only the top 20 were fetched, no reranker will ever see it. How many candidates to fetch is a trade between the first stage's recall and the second stage's cost, and it deserves a deliberate choice rather than whatever number the example code used.

right passages the model reads –never fetched –reranker calls –

Fig. 5 An illustration. Six of two hundred passages answer the question; on the left they sit where embedding similarity ranked them, on a log scale. The reranker reads every fetched candidate with the question, and the five it scores highest go to the model. It's good, but not perfect: one passage that only looks relevant fools it too.

7Measure it

Every choice above, the chunk size, the embedding model, the index settings, how many to fetch, whether to fuse or rerank, moves retrieval quality in a direction you can't see by trying a few questions by hand. What turns them into decisions is a small evaluation set: real questions, each with the passages a person agreed contain its answer. A few dozen is enough to start. Then every change gets measured the same way, and the number that matters most for RAG is recall at k: of the passages that answer the question, how many made it into the k the model will read.

def recall_at_k(retrieved, relevant, k):
    return len(set(retrieved[:k]) & relevant) / len(relevant)

def mrr(retrieved, relevant):          # 1 / the position of the first right answer
    return next((1 / i for i, doc in enumerate(retrieved, 1) if doc in relevant), 0.0)

def evaluate(questions, search, k=5):
    runs = [(search(q.text), q.relevant_ids) for q in questions]
    return {"recall@k": sum(recall_at_k(r, rel, k) for r, rel in runs) / len(runs),
            "mrr": sum(mrr(r, rel) for r, rel in runs) / len(runs)}

8The model is the easy part

It's tempting to treat the language model as the system and retrieval as plumbing. In production it's the other way round. The model is a component you can swap. The retrieval pipeline is where the domain knowledge lives: how documents are split, what gets indexed, which filters apply, and what "relevant" means to the people asking. Get that right and a modest model gives good answers. Get it wrong and the best model in the world will confidently summarise the wrong page.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.