AI Engineer Path

Week 32

RAG (2/2) + Flagship 1

Advanced retrieval, grounded answers, and a production RAG system with an eval suite.

  • Never cut
  • Flagship project

Why this week matters

Know cold: RAG for knowledge that changes, needs citations or is too large; fine-tuning for format, tone and behaviour. Usually both.

Done when

Flagship 1: production RAG with reranking, streaming, citations, a Ragas eval suite, and recorded baseline numbers.

Milestone: Flagship 1: production RAG + evals

Concepts

7 lessons · tick each one once you could explain it

User queries are short, ambiguous and phrased differently from documents. Query rewriting uses an LLM to turn a conversational follow-up ('what about for contractors?') into a standalone query, essential in chat.

Multi-query generates several paraphrases, retrieves for each and merges results (RAG-fusion merges them with RRF). HyDE asks the LLM to write a hypothetical answer and embeds that instead of the question, because answers look more like documents than questions do.

Each adds an LLM call (latency and cost) and can drift from the user's intent. Measure the recall gain before keeping it.

Going deeper

Query decomposition breaks multi-part questions into sub-questions, retrieves for each and synthesises, the building block of agentic and multi-hop RAG.

Bi-encoders embed query and document separately: fast, but they can't model fine interactions. A cross-encoder reads the query and a candidate together and outputs a relevance score: far more accurate, too slow to run over the whole corpus.

Two-stage pattern: retrieve the top 30–100 with hybrid search, rerank them with a cross-encoder (open models or hosted rerank APIs), keep the top 3–8 for the prompt.

Reranking is often the single largest quality jump in a RAG pipeline, and it lets you send fewer, better chunks, which cuts cost.

python
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("BAAI/bge-reranker-base")
candidates = hybrid_search(question, k=50)
scores = reranker.predict([(question, c.text) for c in candidates])
top = [c for _, c in sorted(zip(scores, candidates), key=lambda x: -x[0])][:5]

Going deeper

LLM-based rerankers (listwise or pointwise prompting) can outperform cross-encoders on complex relevance at much higher cost. Cross-encoders remain the standard latency/quality trade-off.

Give each chunk an ID in the prompt, and require the answer to cite IDs for each claim. Render citations as links to the source passage so users can verify. Optionally check after generation that cited chunks actually support the sentence.

Instruct the model explicitly: if the documents don't contain the answer, say so. Add a retrieval-score threshold so when nothing relevant is found you abstain before calling the model at all.

Measure both sides: faithfulness (answers supported by context) and abstention quality (does it decline on unanswerable questions without declining answerable ones?).

Going deeper

Some APIs offer native citation features that return exact source spans for each claim, which is more reliable than asking the model to emit chunk IDs in free text.

Common pitfalls

  • Citations that look right but point to chunks that don't support the claim.

Retrieval metrics: context precision (are retrieved chunks relevant?) and context recall (was everything needed retrieved?). Generation metrics: faithfulness (are claims supported by the context?) and answer relevancy (does it address the question?). Answer correctness needs reference answers.

Ragas computes these, mostly with an LLM as judge. Spot-check judge scores against your own judgement on 20–30 examples before trusting them.

Build an eval set of 50–200 questions (real ones if possible, including unanswerable ones), record a baseline, and re-run on every change: chunking, embeddings, reranker, prompt. Report the numbers in the README.

Going deeper

Ragas metrics are LLM-judged, so validate them on a sample against your own labels, and keep deterministic checks (citation present, abstained when expected) alongside them.

Where this comes back

  • Week 36This suite becomes a CI deploy gate.

Choose RAG when knowledge changes often, answers need citations, the corpus is large or private, or access control matters. Updating is just re-indexing.

Choose fine-tuning for consistent format, tone or style, domain-specific behaviour, distilling a large model's skill into a smaller cheaper one, or shortening long prompts. It is poor at injecting facts reliably and makes them hard to update or cite.

Often both: fine-tune a small model to follow your format and use retrieved context well, with RAG supplying the facts. Start with prompting + RAG, measure, and fine-tune only for a specific measured gap.

Going deeper

Long-context models add a third option for small corpora: put the whole document set in the prompt (with caching). It's simple and strong up to a point, then cost and context rot make RAG better again.

Best resources for this lesson

Where this comes back

  • Week 33You'll run one QLoRA fine-tune and judge whether it was worth it.

Linear RAG is query → retrieve → answer. Agentic RAG puts retrieval behind tools in an agent loop: the model plans, decides whether to search, rewrites queries, retrieves again after reading results, combines sources (vector store, SQL, web) and verifies claims before answering.

Named patterns: Self-RAG (the model critiques its own retrieval and generation), Corrective RAG (a retrieval evaluator triggers re-search or web fallback when results look weak), Adaptive RAG (a classifier picks no retrieval, single-shot or multi-step by question complexity) and multi-hop decomposition (sub-questions answered in sequence).

Costs: more LLM calls, higher latency, harder evaluation. Use it for complex, multi-part or multi-source questions; keep simple RAG for straightforward lookups, ideally with a router choosing between them.

Common pitfalls

  • Unbounded retrieval loops: cap iterations and cost.
  • Evaluating only final answers: also log and score each retrieval step.

Where this comes back

  • Week 34Agentic RAG is an agent whose main tools are retrievers.

Vector search finds passages similar to the question, but struggles with global questions ('what are the main themes across these 2,000 reports?') and multi-hop ones ('which suppliers of our top customer had incidents?').

GraphRAG uses an LLM to extract entities and relationships into a knowledge graph, clusters it into communities, and pre-writes summaries per community. Queries can then traverse relationships (local search) or combine community summaries (global search).

Trade-offs: expensive indexing (many LLM calls), more moving parts, and extraction errors that propagate. Worth it for relationship-heavy domains and corpus-level questions; overkill for FAQ-style lookups.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Add a reranker

    Retrieve 50 with hybrid search, rerank with a cross-encoder, keep 5. Measure the change in answer quality and token cost.

  2. Core

    Flagship 1 evals

    Build a 100-question eval set (including unanswerable questions), run Ragas plus deterministic checks, and record baseline numbers in the README.

  3. Stretch

    Agentic upgrade

    Wrap retrieval as a tool in an agent loop with a 3-search cap; compare quality, latency and cost against the linear pipeline on multi-part questions.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 218Monday26 Apr2 h planned

    Query rewriting, HyDE, multi-query, RAG-fusion

  2. Day 219Tuesday27 Apr2 h planned

    Re-ranking with cross-encoders

  3. Day 220Wednesday28 Apr2 h planned

    Citations, grounding, and answering 'I don't know'

  4. Day 221Thursday29 Apr2 h planned

    FLAGSHIP 1: build production RAG (retrieval + rerank)

  5. Day 222Friday30 Apr2 h planned

    FLAGSHIP 1: streaming, citations, cost and latency logging

  6. Day 223Saturday1 May3 h planned

    FLAGSHIP 1: Ragas eval suite; record the baseline numbers

  7. Day 224Sunday2 MayReview

    Review the week, finish anything unfinished, rest

Watch

RAG From ScratchPrimary

LangChain · playlist

Building Production-Ready RAG Applications

Jerry Liu · AI Engineer

Flagship project

Flagship 1: production RAG with evals

A RAG system that cites its sources, says 'I don't know' when it should, and comes with numbers that prove it works.

  • Chunking strategy chosen by measurement
  • Hybrid search (BM25 + dense) and cross-encoder reranking
  • Streaming answers with citations
  • Ragas eval suite with recorded baseline numbers
  • Cost and latency logging; hardened against indirect prompt injection (week 37)
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Design a production RAG system for internal documents.
  • Ingestion: parsing, structure-aware chunking, metadata/ACLs
  • Retrieval: hybrid + rerank, permission filters
  • Generation: citations, abstain, streaming
  • Evals, tracing, cost/latency, injection hardening
RAG or fine-tuning for a support bot over changing docs?
  • RAG: changing knowledge, citations, ACLs
  • Fine-tuning: tone/format, smaller cheaper model
  • Often both; start with RAG
When is agentic RAG worth its cost?
  • Multi-part, multi-hop, multi-source questions
  • Needs self-correction for high accuracy
  • Route simple questions to plain RAG

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why use a cross-encoder only on top candidates?

  2. 2.HyDE embeds…

  3. 3.Faithfulness measures…

  4. 4.Company policies change monthly and answers need sources. Approach?

  5. 5.Your eval set should include…