AI Engineer Path

Week 31

RAG (1/2): building retrieval

Build naive RAG, watch it fail, and fix retrieval stage by stage.

  • Never cut

Why this week matters

RAG is the single most important node in the track. Let each stage fail before fixing it, so you know why each technique exists.

Done when

A working RAG pipeline with hybrid search, and a written list of the failure cases you've observed.

Concepts

6 lessons · tick each one once you could explain it

Indexing (offline): load documents, split into chunks, embed each chunk, store vectors with text and metadata. Querying (online): embed the question, retrieve the top-k most similar chunks, put them in the prompt with instructions to answer from them, generate.

Typical failures: the answer is split across chunks; the relevant chunk ranks 7th when k = 5; keyword-specific questions (error codes, product names) miss with pure embeddings; the question's phrasing doesn't match the document's; the model ignores the context or answers from its own memory; it confidently answers when nothing relevant was retrieved.

Log the retrieved chunks for every query. Most RAG failures are retrieval failures, and you can only see them if you look. Step through the stages in the simulation.

python
def answer(question: str, k: int = 5) -> str:
    q = embed(question)
    hits = store.search(q, k=k)
    context = "\n\n".join(f"<doc id='{h.id}'>{h.text}</doc>" for h in hits)
    prompt = (f"Answer using only the documents below. If they don't contain the answer, say so.\n"
              f"{context}\n\nQuestion: {question}")
    return llm(prompt)

Going deeper

The RAG survey's taxonomy (naive → advanced → modular RAG) maps exactly onto weeks 31–32: each stage adds pre-retrieval (query transformation), retrieval (hybrid, filtering) and post-retrieval (reranking, compression) improvements.

Where this comes back

  • Week 30Embedding choice and recall@k decide what retrieval can find.

Fixed-size chunks (e.g. 500 tokens with 50–100 overlap) are simple but cut sentences and tables mid-way. Recursive splitting tries separators in order (sections, paragraphs, sentences) to respect structure. Semantic chunking splits where the embedding similarity between consecutive sentences drops, at the cost of more computation.

Small chunks give precise retrieval but lose context; large chunks keep context but dilute the embedding and waste window space. Respect document structure: headings, lists, tables, code blocks. Prepend the document title and section heading to each chunk.

There is no universal best size. Measure recall@k for 2–3 configurations on your query set.

Going deeper

Late chunking embeds the whole document with a long-context embedding model first, then pools token embeddings per chunk, so each chunk's vector carries document context.

For structured sources (Markdown, HTML, code), split on structure (headings, functions) and keep the heading path as metadata, which is often the single biggest chunking win.

Common pitfalls

  • Splitting tables row by row so headers are lost.
  • Choosing a chunk size without measuring.

Decouple what you search from what you give the model. Parent-document retrieval: index small child chunks, but when one matches, return its larger parent section. Sentence-window retrieval: index single sentences, return the matched sentence plus a few neighbours on each side.

You get the precision of small embeddings and the context of big passages.

Deduplicate when several children map to the same parent, and watch the token budget.

Going deeper

RAPTOR builds a tree of recursive summaries over chunks, so retrieval can return a high-level summary for broad questions and leaf chunks for specific ones.

Best resources for this lesson

Store structured metadata with every chunk: source, document type, date, product, language, and access-control fields. Filter on it: 'only 2024 policies', 'only docs for product X', and above all 'only documents this user may see'.

Filters can be explicit (UI facets) or extracted from the question by an LLM (self-querying), which you should validate.

Permission filtering is non-negotiable in multi-user systems: retrieval that ignores ACLs leaks data straight into answers.

Going deeper

Multi-tenant RAG should isolate tenants structurally (separate collections, partitions, or row-level security), not only with a filter clause someone can forget. Vector and embedding weaknesses, including cross-tenant leakage, are on the OWASP LLM Top 10.

Common pitfalls

  • Post-filtering after top-k, which can leave zero results.
  • Forgetting tenant or permission filters.
Backend engineer tip: Row-level security for retrieval. Enforce it in the data layer, not in the prompt.

Chunks lose context: 'Revenue grew 3% over the previous quarter' doesn't say which company or quarter. Contextual retrieval asks an LLM, given the whole document and one chunk, to write a sentence or two situating the chunk, and prepends that context before embedding and keyword indexing.

Anthropic reported large reductions in retrieval failures from contextual embeddings plus contextual BM25, and further gains with reranking. Prompt caching makes generating context for every chunk affordable, since the document is a cached prefix.

It's an indexing-time cost (one LLM call per chunk) for a query-time gain. Measure recall@k on your own query set to decide.

Best resources for this lesson

Where this comes back

  • Week 28Prompt caching is what makes per-chunk context generation cheap.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Watch naive RAG fail

    Build naive RAG over your own documents, write 30 questions, and log retrieved chunks. Categorise every failure.

  2. Warm-up

    Chunking bake-off

    Compare fixed, recursive and structure-aware chunking on recall@5 with the same embedding model.

  3. Stretch

    Hybrid + contextual

    Add BM25 with RRF fusion, then contextual chunk headers. Report recall@5 after each step.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 211Monday19 Apr2 h planned

    Build naive RAG end-to-end - and watch it fail

  2. Day 212Tuesday20 Apr2 h planned

    Chunking: fixed, recursive, semantic

  3. Day 213Wednesday21 Apr2 h planned

    Chunking: parent-document and sentence-window

  4. Day 214Thursday22 Apr2 h planned

    Hybrid search: BM25 + dense

  5. Day 215Friday23 Apr2 h planned

    Metadata filtering

  6. Day 216Saturday24 Apr3 h planned

    Read the RAG paper + work an LLM Zoomcamp module

  7. Day 217Sunday25 AprReview

    Review the week, finish anything unfinished, rest

Watch

Generative AI using LangChainPrimary

CampusX · playlist

RAG From ScratchPrimary

LangChain · playlist

Paper of the week

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

The paper that named RAG. Compare its setup with the production pipeline you build in weeks 31–32.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Your RAG system gives wrong answers. How do you debug it?
  • Separate retrieval vs generation: check retrieved chunks first
  • Measure recall@k on labelled queries
  • Fix chunking, hybrid search, filters, reranking; then prompt/grounding
Why use hybrid search?
  • Dense: semantics and paraphrase
  • BM25: exact terms, IDs, rare words
  • Fuse with RRF; consistently more robust

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Users search for error code 'E-4021' and RAG misses it. Best fix?

  2. 2.Reciprocal rank fusion combines lists using…

  3. 3.Parent-document retrieval searches over…

  4. 4.Most RAG failures are…

  5. 5.Why must retrieval apply permission filters?