AI Engineer Path

Concept Lab · Week 31 · AI Engineer Track

A RAG pipeline

Chunk → embed → retrieve → rerank → answer, with each stage visible.

The idea: Naive RAG end to end, and how it fails

Indexing (offline): load documents, split into chunks, embed each chunk, store vectors with text and metadata. Querying (online): embed the question, retrieve the top-k most similar chunks, put them in the prompt with instructions to answer from them, generate.

Typical failures: the answer is split across chunks; the relevant chunk ranks 7th when k = 5; keyword-specific questions (error codes, product names) miss with pure embeddings; the question's phrasing doesn't match the document's; the model ignores the context or answers from its own memory; it confidently answers when nothing relevant was retrieved.

Log the retrieved chunks for every query. Most RAG failures are retrieval failures, and you can only see them if you look. Step through the stages in the simulation.

Open the full lesson in week 31

Next simulation: LoRA: low-rank updates