Building production RAG with Qdrant and LangGraph
Why naive vector search breaks down in production, how to blend dense embeddings with lexical search, and how to structure an agent that refuses to answer without a citation.
Writing
Long-form thinking about what happens after the demo works: grounding, evaluation, cost and failure.
Why naive vector search breaks down in production, how to blend dense embeddings with lexical search, and how to structure an agent that refuses to answer without a citation.
Agents are rarely tested the way other software is. A case for evaluation suites that treat an agent's state machine with the rigour of a compiler test suite.
Confidence is not correctness. Fallback thresholds, uncertainty signals and refusal paths that keep a system dependable when the stakes are high.
Frontier model calls are expensive and slow. How routing, caching and smaller models turn an unpredictable bill into an engineered cost curve.
Next step