Rohit Swami
India Resume ↗

Writing · GenAI · 6 min read

Retrieval is the hard part

When a RAG system gives a bad answer, the model usually takes the blame. More often it was handed the wrong context. Embeddings, chunking and approximate search, seen from the retrieval side.

Retrieval-augmented generation is a simple idea. Before asking a language model a question, find the passages most likely to contain the answer and put them in front of it. The model writes the answer; retrieval decides what it gets to read.

I've shipped production GenAI applications that do semantic search and summarisation over domain-specific data, built on LLMs and vector databases like FAISS and Qdrant. The lesson I'd pass on first is that when an answer is bad, the model is rarely the place to start looking. It can only work with the context it was given, and a bad answer usually traces back to bad retrieval: the right passage wasn't found, or it was found alongside so much noise that the model followed the wrong thread.

1What "similar" means to a machine

Semantic search starts with embeddings. An embedding model turns a piece of text into a long list of numbers, a point in a space with hundreds or thousands of dimensions, arranged so that texts with similar meanings land near each other. The question is embedded the same way, and "find the relevant passages" becomes "find the nearest points", usually by the angle between vectors: cosine similarity.

top k

returned –best match –

Fig. 1 Drag the query. Real embeddings have hundreds of dimensions; here nearness on the page stands in for cosine similarity, which makes it a cartoon, but an honest one. Drag it into the empty space and see what comes back.

Flattened to two dimensions it's a cartoon, but the behaviour is real. Retrieval returns the nearest neighbours whether or not any of them is relevant. Put the query between topics and the results become a mixture. Ask about something the corpus doesn't contain and you still get k confident-looking passages, because nothing in the vector maths says "I don't know". That has to be designed in, with a similarity threshold or a reranker that is allowed to return nothing.

2Chunking decides what can be found

Documents are too long to embed whole, so they're split into chunks, and the chunk size quietly sets a ceiling on retrieval quality. An embedding is a kind of average of what its text is about. Make chunks large and each one covers several topics, so its embedding blurs and a precise question matches it only weakly. Make them small and each chunk is sharp but short of context: the sentence that answers the question often doesn't say what it's about, because the sentence before it did.

chunk size

retrieved chunk has the answer –relevant text in it –

Fig. 2 A toy model. Colours are topics, and the question is about the blue one. The answer is the sentence marked A, which doesn't name its topic; the sentence before it does. Each chunk scores by how much of it is about the question, shown as the bars above it, and the best one is retrieved.

There's no universally right size. It depends on how the documents are written and what people ask of them. Splitting along the document's own structure (sections, paragraphs, table rows) usually beats cutting every N tokens, and a little overlap between chunks stops an answer from being severed from the sentence that gives it meaning. The only reliable way to choose is to measure on real questions.

3Finding neighbours without checking everyone

Comparing the query with every stored vector is exact, and fine for a few thousand chunks. At millions it's too slow, so vector databases use approximate nearest-neighbour indexes. One of the most widely used, and the one Qdrant is built around, is HNSW: a hierarchical navigable small-world graph. FAISS offers it too, alongside other index types.

HNSW links every vector to a handful of its neighbours, then builds sparser layers on top, like express lanes. A search starts in the top layer and hops greedily towards the query until no neighbour is closer, then drops a layer and repeats. By the bottom layer it's in the right neighbourhood, having measured the distance to only a small fraction of the collection.

distances computed –checking everything –found the true nearest –

Fig. 3 Three hundred vectors in three layers. Each plane is the same space; higher layers hold fewer points with longer links. The red path is the search: greedy hops in each layer, then a drop to the one below.

The approximation is a dial, not a flaw. The number of links per node and how widely the search explores trade speed and memory for recall, and it's worth measuring recall against exact search on your own data before trusting the defaults.

4What actually helps

In practice, the improvements that matter most tend to be the unglamorous ones:

5The model is the easy part

It's tempting to treat the language model as the system and retrieval as plumbing. In production it's the other way round. The model is a component you can swap. The retrieval pipeline is where the domain knowledge lives: how documents are split, what gets indexed, which filters apply, and what "relevant" means to the people asking. Get that right and a modest model gives good answers. Get it wrong and the best model in the world will confidently summarise the wrong page.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.