AI Engineer Path

Week 33

Memory, fine-tuning (light) and AI automation

Give systems memory, run one honest fine-tune, and build an automated workflow.

Why this week matters

'Light' is correct: you'll consume fine-tunes far more often than you produce them. One QLoRA run with an honest before/after eval, then move on.

Done when

Project 5: a QLoRA fine-tune with an honest before/after evaluation, plus one working automated workflow.

Milestone: Project 5: QLoRA fine-tune with honest eval

Concepts

7 lessons · tick each one once you could explain it

LLMs are stateless: memory is whatever you put back into the context. Short-term memory is the running conversation, managed with buffers (last N turns), token windows and summaries (week 28 compaction).

Long-term memory persists across sessions: user preferences, facts learned, past decisions. It is stored externally and retrieved into context when relevant.

Kinds worth distinguishing: semantic (facts about the user or world), episodic (what happened in past interactions), procedural (learned instructions or how-tos).

Going deeper

Memory raises product questions as much as technical ones: what the user can see and delete, how long memories live, and whether memories cross contexts (work vs personal).

Where this comes back

  • Week 28Compaction manages short-term memory.

A memory layer extracts candidate facts from conversations with an LLM ('prefers Python', 'is preparing for interviews in June'), deduplicates and updates them when they change, and stores them with embeddings. Each new turn retrieves relevant memories into the prompt.

Tools like Mem0 package this; LangGraph offers stores for the same purpose.

Hard parts: deciding what is worth remembering, resolving contradictions, expiring stale facts, and privacy: users must be able to see and delete what's remembered.

Going deeper

Memory writes are an injection surface: if an attacker can get text into a user's memories, it influences every future session. Treat memory updates as untrusted writes.

Common pitfalls

  • Remembering sensitive data the user never meant to share.
  • Stale memories overriding newer facts.

Best resources for this lesson

Full fine-tuning updates every weight: billions of parameters with gradients and optimiser state for each. LoRA freezes the pretrained weight WW and learns an update ΔW=BA\Delta W = BA, where BB is d×rd \times r and AA is r×kr \times k with a small rank rr (8–64). Trainable parameters drop by 100–1000×.

This works because the adjustment needed for a new task appears to be low-rank: the week 4 linear algebra, applied. Adapters are small files that can be swapped per task over one base model, or merged into the weights for zero inference overhead.

QLoRA loads the frozen base in 4-bit precision and trains LoRA adapters in 16-bit, so a 7–8B model fine-tunes on a single consumer or free Colab GPU. Count the parameters saved in the simulation.

h=Wx+αrBAx,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)h = Wx + \frac{\alpha}{r}BAx,\qquad B\in\mathbb{R}^{d\times r},\ A\in\mathbb{R}^{r\times k},\ r \ll \min(d,k)

Going deeper

Target modules matter: applying LoRA to all linear layers (attention and MLP) usually beats attention-only. Rank 8–64 with alpha ≈ rank or 2×rank are common starting points.

Variants: DoRA (decomposes weights into magnitude and direction), LoRA+ (different learning rates for A and B), and multi-adapter serving (one base model, many LoRA adapters hot-swapped per request).

Where this comes back

  • Week 4Rank and low-rank decomposition.
  • Week 17Transfer learning: freeze most, train a little.

Pick a narrow task where prompting alone falls short (a strict output format, a domain style, or classifying with your labels). Prepare a few hundred to a few thousand high-quality examples in the model's chat template. Data quality matters far more than quantity.

Unsloth's notebooks set up 4-bit loading, LoRA config (rank, alpha, target modules), and TRL's SFT trainer. Train 1–3 epochs with a held-out validation split and save the adapter.

Watch for overfitting (validation loss rising) and catastrophic forgetting (general ability degrading). Test a few off-task prompts too.

python
from unsloth import FastLanguageModel
model, tok = FastLanguageModel.from_pretrained("unsloth/Llama-3.2-3B-Instruct", load_in_4bit=True)
model = FastLanguageModel.get_peft_model(model, r=16, lora_alpha=16,
          target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"])
# then trl.SFTTrainer(model=model, train_dataset=train, eval_dataset=val, ...).train()

Going deeper

Use the model's chat template exactly (tokenizer.apply_chat_template); mismatched templates are the most common cause of a fine-tune that 'learned nothing'. Train only on assistant tokens (mask the prompt) for instruction tuning.

Compare on one held-out test set the model never saw: (1) the base model with your best prompt, including few-shot examples; (2) the fine-tuned model; (3) a strong frontier model with prompting; (4) your classic baseline where relevant.

Report quality, latency, cost per 1,000 requests and operational burden (hosting, updating). A fine-tune that beats a weak prompt but not a good one hasn't earned its place.

Write the verdict plainly: worth it or not, and why. That paragraph is the interview story.

Going deeper

Check for regressions outside the target task with a small general benchmark or a set of off-task prompts. Narrow fine-tunes can quietly degrade instruction following or safety behaviour.

Common pitfalls

  • Comparing against a lazy zero-shot prompt.
  • Test examples leaking into training data.

Where this comes back

  • Week 13The baseline principle, again.

Many 'agents' should really be workflows: a fixed sequence (trigger → fetch → LLM classify → branch → act → notify) where only some steps need an LLM. Workflows are cheaper, more predictable and easier to test.

n8n builds them visually with hundreds of integrations: good for business automation and fast iteration. LangGraph builds them in code with state, branching, retries and human approval steps: good when it's part of a product.

Build one end to end: for example, new support email → classify and extract with structured output → create a ticket → draft a reply for human approval.

Going deeper

Anthropic's workflow patterns (prompt chaining, routing, parallelisation, orchestrator–workers, evaluator–optimiser) are a useful checklist before reaching for a fully autonomous agent.

Where this comes back

  • Week 34When a workflow isn't enough, you need an agent.

Knowledge distillation trains a small 'student' to imitate a large 'teacher', classically by matching its output probabilities, and today often by fine-tuning on the teacher's generated answers. For a narrow, high-volume task, a distilled 7B model can approach a frontier model's quality at a fraction of the cost and latency.

Synthetic data generation (instructions, answers, preference pairs, retrieval queries) is now standard. Quality control is everything: filter with validators and judges, deduplicate, keep diversity, and mix in real data to avoid collapse onto the teacher's quirks.

Check the teacher model's terms of service, since some providers restrict using outputs to train competing models.

Where this comes back

  • Week 40Distillation is a lever in 'cut the LLM bill by 60%'.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Memory that forgets correctly

    Add long-term memory to a chatbot; test that it remembers a preference across sessions, updates it when the user changes their mind, and deletes it on request.

  2. Stretch

    Project 5: QLoRA

    Fine-tune a small model with Unsloth on a narrow task; compare against the base model with your best prompt, plus a frontier model, on a held-out set with cost per 1k requests.

  3. Warm-up

    Parameter arithmetic

    For a 7B model with d = 4096 and 32 layers, compute LoRA trainable parameters at r = 16 on all attention and MLP projections.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 225Monday3 May2 h planned

    Short-term vs long-term memory; buffers and summaries

  2. Day 226Tuesday4 May2 h planned

    Vector-backed long-term memory (Mem0)

  3. Day 227Wednesday5 May2 h planned

    Fine Tuning (Light): LoRA / QLoRA concepts + read the LoRA paper

  4. Day 228Thursday6 May2 h planned

    QLoRA fine-tune a small model with Unsloth

  5. Day 229Friday7 May2 h planned

    PROJECT 5: evaluate the fine-tune honestly (before / after)

  6. Day 230Saturday8 May3 h planned

    AI Automation: build an n8n or LangGraph workflow

  7. Day 231Sunday9 MayReview

    Review the week, finish anything unfinished, rest

Watch

What is LoRA, explained by the inventor

Edward Hu

LoRA explained (and a bit about precision and quantization)

DeepFindr

Low-rank Adaptation: Key Concepts Behind LoRA

Chris Alexiuk

Project 5

QLoRA fine-tune with an honest eval

One QLoRA run on a small open model, with a before/after evaluation you would defend in an interview.

  • Training data card: source, size, cleaning
  • QLoRA run with Unsloth (free Colab is enough)
  • Before / after scores on a held-out set
  • Written verdict: was fine-tuning worth it versus prompting or RAG?
Track it on the Projects page

Paper of the week

LoRA: Low-Rank Adaptation of Large Language Models

The rank idea from week 4 linear algebra, applied to fine-tuning at a fraction of the cost.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Explain LoRA and QLoRA.
  • Freeze W; learn low-rank ΔW = BA
  • Trainable params r(d+k) instead of d·k
  • QLoRA: 4-bit frozen base, 16-bit adapters → single-GPU fine-tuning
How do you know a fine-tune was worth it?
  • Held-out test vs best-prompted base and frontier models
  • Quality, latency, cost per 1k requests
  • Check regressions on off-task behaviour

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.LoRA reduces trainable parameters by…

  2. 2.QLoRA's key addition over LoRA is…

  3. 3.A fair baseline for your fine-tune is…

  4. 4.Long-term memory differs from short-term memory because it…

  5. 5.'Classify incoming emails and open a ticket' is best built as…