Indexing (offline): load documents, split into chunks, embed each chunk, store vectors with text and metadata. Querying (online): embed the question, retrieve the top-k most similar chunks, put them in the prompt with instructions to answer from them, generate.
Typical failures: the answer is split across chunks; the relevant chunk ranks 7th when k = 5; keyword-specific questions (error codes, product names) miss with pure embeddings; the question's phrasing doesn't match the document's; the model ignores the context or answers from its own memory; it confidently answers when nothing relevant was retrieved.
Log the retrieved chunks for every query. Most RAG failures are retrieval failures, and you can only see them if you look. Step through the stages in the simulation.
def answer(question: str, k: int = 5) -> str:
q = embed(question)
hits = store.search(q, k=k)
context = "\n\n".join(f"<doc id='{h.id}'>{h.text}</doc>" for h in hits)
prompt = (f"Answer using only the documents below. If they don't contain the answer, say so.\n"
f"{context}\n\nQuestion: {question}")
return llm(prompt)Going deeper
The RAG survey's taxonomy (naive → advanced → modular RAG) maps exactly onto weeks 31–32: each stage adds pre-retrieval (query transformation), retrieval (hybrid, filtering) and post-retrieval (reranking, compression) improvements.
Best resources for this lesson
Where this comes back
- Week 30Embedding choice and recall@k decide what retrieval can find.