WritingRetrieval5 min read2026

Building production RAG with Qdrant and LangGraph

Why naive vector search breaks down in production, how to blend dense embeddings with lexical search, and how to structure an agent that refuses to answer without a citation.

Most retrieval-augmented generation systems fail in the same quiet way. The demo works. Someone asks a question the corpus can answer, the top five chunks contain the answer, the model paraphrases it well, and everyone in the room nods. Then the system meets real users, real documents and real questions, and it starts returning confident answers that are almost right.

Almost right is the most expensive failure mode in AI. A wrong answer that looks wrong gets caught. An answer that is ninety percent correct, cites a plausible source and reads fluently gets trusted. This essay is about the engineering that closes that gap: how retrieval actually breaks, why hybrid search is the default rather than an optimisation, and how to structure an agent so that it would rather refuse than guess.

Why naive vector search breaks

The standard tutorial pipeline is: split documents into fixed-size chunks, embed each chunk, store the vectors, embed the question, take the nearest neighbours, stuff them into a prompt. Every step in that pipeline hides an assumption that production will violate.

Fixed-size chunking assumes meaning is evenly distributed through a document. It is not. A legal clause, a function definition, a table row: these are units of meaning with hard boundaries, and a 512-token window cuts straight through them. The retriever then returns half a clause, and the model fills in the other half from its priors. That is where a lot of so-called hallucination actually comes from. The model was not making things up from nothing; it was completing a fragment you handed it.

Dense embeddings assume that semantic similarity is what the user wants. Often it is not. When someone searches for a specific statute number, a product code, an error string or a person's name, they want an exact match, and embedding models are notoriously bad at exact matches. Two clause numbers that differ by one digit sit almost on top of each other in vector space. Lexical search, the boring BM25 kind, handles this trivially.

Top-k assumes the answer lives in a fixed number of chunks. Some questions need one passage. Comparison questions need passages from several documents. Questions about absence (does this contract contain a termination clause?) need you to know that you searched properly and found nothing, which top-k cannot tell you at all.

Chunk on structure, not on token counts

The first fix is upstream of the vector database. Parse documents into their real structure before you chunk: headings, sections, clauses, lists, tables. Keep the hierarchy as metadata, so every chunk knows which document, section and parent heading it belongs to. When a chunk is too long, split it along its own internal structure. When it is too short to stand alone, carry its parent heading with it.

This sounds like plumbing, and it is, but it changes what the rest of the system can do. Once chunks are meaningful units, a citation can point at something a human recognises, such as clause 14.2, instead of characters 18,230 to 18,742. Once structure is metadata, you can filter by it before you rank anything.

Hybrid retrieval in Qdrant

My default is hybrid retrieval: every chunk is stored with a dense vector for semantic similarity and a sparse vector for lexical matching, and a query runs against both. Qdrant supports this natively with named vectors, so the two representations live on the same point alongside the payload you filter on.

The results from the two searches are fused, usually with reciprocal rank fusion, which only cares about each result's rank in each list and so avoids the problem of comparing scores that are on different scales. The fused list then goes through a cross-encoder re-ranker, which is slower but reads the query and passage together and is far better at judging actual relevance than either first-stage retriever.

Payload filters do a lot of quiet work here. If the user is asking about one matter, one client or one jurisdiction, filter to it before ranking. Retrieval quality improves more from narrowing the haystack than from any clever ranking trick, and it is also a security boundary: a user should never be able to retrieve a passage they are not entitled to read, and filtering at query time is the simplest place to enforce that.

The agent decides whether it has enough

A single retrieval call is a guess about what the model will need. For anything beyond simple lookups I use a LangGraph state machine instead: the agent retrieves, inspects what came back, and decides whether the evidence is sufficient. If it is not, it reformulates the query, searches a different scope, or pulls the neighbouring sections of a promising passage.

The important word is state machine. I want the agent's possible moves to be explicit nodes and edges: retrieve, assess, expand, compare, answer, refuse. Each transition has a condition. There is a hard limit on how many retrieval rounds it can run. This is less flexible than letting a model improvise with a bag of tools, and that is the point. When something goes wrong, I can read the trace and see exactly which node made which decision with which evidence in front of it.

Citations as a contract, not a feature

The design decision that matters most is treating citations as a hard contract. Every claim in the final answer must point to the span of source text that supports it. The answer is assembled claim by claim, and a verification step checks each claim against its cited span. Claims that cannot be supported are removed. If what remains does not answer the question, the system says so.

This changes the behaviour of the whole system. The model can no longer smooth over gaps with fluent filler, because filler has no source span. It also changes how people use the output. A reviewer does not have to trust the answer; they can click through to the evidence and check it in seconds. In domains like law, that is the difference between a tool professionals can use and a toy they cannot.

Refusal is a feature

A system that answers everything is a system that is wrong some of the time with no way to tell when. I would much rather ship one that refuses clearly when retrieval confidence is low, when the question falls outside the corpus, or when the evidence conflicts. Users adapt quickly to a system that says it does not know. They do not forgive one that confidently misleads them.

Refusal has to be designed, not just allowed. It needs its own path through the graph, its own wording and its own tests. A refusal should explain what was searched and why it was not enough, so the user can rephrase or go to the source themselves.

Evaluate it like software

None of this holds without evaluation. I keep a golden dataset of question, answer and citation triples, including questions the system should refuse, and run it on every change to chunking, retrieval, prompts or models. It scores groundedness, citation precision and refusal behaviour separately, because a change can improve one while quietly breaking another.

Production RAG is not a prompt with a vector database attached. It is a retrieval system, a reasoning loop and a verification layer, each of which can be measured on its own. Build it that way and it stops being almost right.

Next step

Have a system that needs
to hold up in production?