Writing · GenAI · 6 min read
Retrieval is the hard part
When a RAG system gives a bad answer, the model usually takes the blame. More often it was handed the wrong context. Embeddings, chunking and approximate search, seen from the retrieval side.
Retrieval-augmented generation is a simple idea. Before asking a language model a question, find the passages most likely to contain the answer and put them in front of it. The model writes the answer; retrieval decides what it gets to read.
I've shipped production GenAI applications that do semantic search and summarisation over domain-specific data, built on LLMs and vector databases like FAISS and Qdrant. The lesson I'd pass on first is that when an answer is bad, the model is rarely the place to start looking. It can only work with the context it was given, and a bad answer usually traces back to bad retrieval: the right passage wasn't found, or it was found alongside so much noise that the model followed the wrong thread.
1What "similar" means to a machine
Semantic search starts with embeddings. An embedding model turns a piece of text into a long list of numbers, a point in a space with hundreds or thousands of dimensions, arranged so that texts with similar meanings land near each other. The question is embedded the same way, and "find the relevant passages" becomes "find the nearest points", usually by the angle between vectors: cosine similarity.
returned –best match –
Flattened to two dimensions it's a cartoon, but the behaviour is real. Retrieval returns the nearest neighbours whether or not any of them is relevant. Put the query between topics and the results become a mixture. Ask about something the corpus doesn't contain and you still get k confident-looking passages, because nothing in the vector maths says "I don't know". That has to be designed in, with a similarity threshold or a reranker that is allowed to return nothing.
2Chunking decides what can be found
Documents are too long to embed whole, so they're split into chunks, and the chunk size quietly sets a ceiling on retrieval quality. An embedding is a kind of average of what its text is about. Make chunks large and each one covers several topics, so its embedding blurs and a precise question matches it only weakly. Make them small and each chunk is sharp but short of context: the sentence that answers the question often doesn't say what it's about, because the sentence before it did.
retrieved chunk has the answer –relevant text in it –
There's no universally right size. It depends on how the documents are written and what people ask of them. Splitting along the document's own structure (sections, paragraphs, table rows) usually beats cutting every N tokens, and a little overlap between chunks stops an answer from being severed from the sentence that gives it meaning. The only reliable way to choose is to measure on real questions.
3Finding neighbours without checking everyone
Comparing the query with every stored vector is exact, and fine for a few thousand chunks. At millions it's too slow, so vector databases use approximate nearest-neighbour indexes. One of the most widely used, and the one Qdrant is built around, is HNSW: a hierarchical navigable small-world graph. FAISS offers it too, alongside other index types.
HNSW links every vector to a handful of its neighbours, then builds sparser layers on top, like express lanes. A search starts in the top layer and hops greedily towards the query until no neighbour is closer, then drops a layer and repeats. By the bottom layer it's in the right neighbourhood, having measured the distance to only a small fraction of the collection.
distances computed –checking everything –found the true nearest –
The approximation is a dial, not a flaw. The number of links per node and how widely the search explores trade speed and memory for recall, and it's worth measuring recall against exact search on your own data before trusting the defaults.
4What actually helps
In practice, the improvements that matter most tend to be the unglamorous ones:
- Hybrid search. Embeddings are good at meaning and bad at exact strings. Domain data is full of identifiers (gene symbols, compound IDs, trial numbers) where the exact string is the whole point. Combining vector search with keyword search such as BM25 catches what each one misses.
- Filters before similarity. If the answer must come from one organism, one study or one customer's documents, filter on that metadata first. Qdrant can filter inside the vector search itself, which beats retrieving broadly and throwing most of it away.
- Reranking. Fetch a generous set of candidates cheaply, then reorder them with a slower, more precise model that reads the question and each passage together.
- Evaluation. A small set of real questions with known good passages, and recall measured against it whenever chunking, embeddings or index settings change. Without it, every change is a guess.
5The model is the easy part
It's tempting to treat the language model as the system and retrieval as plumbing. In production it's the other way round. The model is a component you can swap. The retrieval pipeline is where the domain knowledge lives: how documents are split, what gets indexed, which filters apply, and what "relevant" means to the people asking. Get that right and a modest model gives good answers. Get it wrong and the best model in the world will confidently summarise the wrong page.