AI Engineer Path

Week 22

Unsupervised DL, deployment and review

Representation learning, a deployed demo, and a written review of everything so far.

  • First to cut if behind
  • Light week + buffer

Why this week matters

A deliberately light week with a buffer day. If you've slipped, absorb it here. Writing the review exposes what you don't actually know.

Done when

A written summary of everything learned so far, and a model deployed on Hugging Face Spaces.

Concepts

6 lessons · tick each one once you could explain it

In plain words

An autoencoder squeezes its input into a small code and then tries to rebuild the original from that code. Because the code is small, the network must keep only what matters most. If something rebuilds badly, it's unusual, which makes autoencoders useful for spotting anomalies.

Summarising a book in one paragraph, then trying to rewrite the book from your summary: what survives is what mattered.

In detail

An autoencoder has an encoder that maps input to a low-dimensional latent code and a decoder that reconstructs the input from it. Training minimises reconstruction error, so the bottleneck forces the code to capture the most important structure: a non-linear generalisation of PCA.

Variants: denoising autoencoders reconstruct clean input from corrupted input; sparse autoencoders penalise active units.

Uses: anomaly detection (high reconstruction error = unusual), compression, pretraining. Sparse autoencoders are now a key tool for interpreting what LLM features represent.

z=fθ(x),x^=gϕ(z),L=∥x−x^∥2z = f_\theta(x),\quad \hat x = g_\phi(z),\quad \mathcal{L} = \|x - \hat x\|^2
xx
the input
fθf_\theta
the encoder, with weights θ\theta
zz
the latent code: the compressed representation
gϕg_\phi
the decoder, with weights ϕ\phi
x^\hat x
the reconstruction
∥x−x^∥2\|x - \hat x\|^2
the reconstruction error

Worked example

Compressing digits and catching anomalies

  1. MNIST digits: 784 pixels → encoder → 32 numbers → decoder → 784 pixels, a 24.5× compression.
  2. Train to minimise reconstruction error between input and output.
  3. On normal digits, the average error is around 0.01.
  4. Feed it a letter 'A' it never saw: the error jumps to about 0.08, because the code has no way to represent it.
  5. Set the anomaly threshold at, say, the 99th percentile of errors on normal validation data.

A narrow bottleneck forces the code to capture structure; high reconstruction error flags the unfamiliar.

Common mistakes

  • A bottleneck as wide as the input: the network just learns to copy and captures nothing useful.
  • Trusting the anomaly threshold without checking it on labelled examples of real anomalies.

Check yourself

How is an autoencoder related to PCA?Show answer

A linear autoencoder with squared error learns the same subspace as PCA. Non-linear layers let it capture curved structure PCA can't.

What's a denoising autoencoder?Show answer

One trained to reconstruct the clean input from a corrupted version, which forces it to learn robust structure rather than copy pixels.

Going deeper

Sparse autoencoders trained on an LLM's internal activations decompose them into many interpretable 'features' (concepts the model represents), now a central tool in mechanistic interpretability.

Best resources for this lesson

In plain words

A variational autoencoder encodes each input as a small cloud of possible codes rather than one exact point, and keeps all those clouds close to a standard bell curve. The result is a smooth latent space: pick any point and the decoder produces something sensible, so you can generate new examples.

Instead of pinning each photo to one exact spot on a map, a VAE gives it a small fuzzy area, and the areas overlap so the whole map is covered with no gaps.

In detail

A variational autoencoder encodes each input to a distribution (a mean and variance) rather than a point, samples from it using the reparameterisation trick, and adds a KL-divergence term that pulls the latent distribution toward a standard normal.

Result: a smooth, continuous latent space where nearby points decode to similar outputs and you can generate new samples. Latent diffusion models (like Stable Diffusion) run diffusion in a VAE's latent space.

The broader idea, representation learning, is that good internal representations transfer across tasks. It is the thread running from Word2Vec through BERT to the embeddings you'll put in vector databases.

L=Eq(z∣x)[log⁡p(x∣z)]−DKL(q(z∣x) ∥ N(0,I))\mathcal{L} = \mathbb{E}_{q(z|x)}[\log p(x|z)] - D_{KL}\big(q(z|x)\,\|\,\mathcal{N}(0, I)\big)
q(z∣x)q(z|x)
the encoder's distribution over codes for input xx
p(x∣z)p(x|z)
the decoder's probability of reconstructing xx from code zz
Eq(z∣x)[log⁡p(x∣z)]\mathbb{E}_{q(z|x)}[\log p(x|z)]
the reconstruction term: rebuild the input well
DKLD_{KL}
KL divergence: how far the code distribution is from the target
N(0,I)\mathcal{N}(0, I)
the standard normal: the shape every code cloud is pulled towards

Worked example

One sample and its KL penalty

  1. The encoder outputs mean μ = 0.5 and standard deviation σ = 0.2 for one input (1D for simplicity).
  2. Reparameterisation: draw ε = 1.5 from a standard normal, so z=μ+σϵ=0.5+0.2×1.5=0.8z = \mu + \sigma\epsilon = 0.5 + 0.2 \times 1.5 = 0.8. Gradients flow through μ and σ.
  3. The decoder reconstructs from z = 0.8.
  4. KL to the standard normal: 12(σ2+μ2−1−ln⁡σ2)=12(0.04+0.25−1+3.22)≈1.25\tfrac12(\sigma^2 + \mu^2 - 1 - \ln\sigma^2) = \tfrac12(0.04 + 0.25 - 1 + 3.22) \approx 1.25, which pulls σ up and μ towards 0.
  5. Walking from one image's code to another's decodes into a smooth blend between them.

Encode to a distribution, sample with the reparameterisation trick, and the KL term keeps the latent space smooth enough to sample from.

Common mistakes

  • Too strong a KL weight: the decoder learns to ignore the code ('posterior collapse').
  • Expecting sharp images from a plain VAE: its outputs tend to be blurry.

Check yourself

Why is the reparameterisation trick needed?Show answer

Sampling is random, so gradients can't flow through it directly. Writing z=μ+σϵz = \mu + \sigma\epsilon moves the randomness into ε, and gradients flow through μ and σ.

Where do VAEs appear in modern image generation?Show answer

Latent diffusion models such as Stable Diffusion run the diffusion process inside a VAE's compressed latent space, then decode the result to pixels.

Going deeper

The reparameterisation trick writes a sample as z=μ+σ⊙ϵz = \mu + \sigma\odot\epsilon with ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I), so gradients flow through μ and σ even though sampling is random.

β-VAE increases the KL weight to encourage disentangled latents. Too much KL causes 'posterior collapse', where the decoder ignores the latent code.

In plain words

Object detection finds where things are in an image as well as what they are: it outputs boxes, each with a label and a confidence score. IoU measures how much two boxes overlap, and is used both to remove duplicate boxes and to decide whether a predicted box counts as correct.

Not just 'there's a cat in this photo', but drawing a rectangle round each cat and labelling it.

In detail

Detection outputs bounding boxes with classes and confidence scores. IoU (intersection over union) measures box overlap; non-max suppression removes duplicate boxes; mAP summarises precision–recall across classes and IoU thresholds.

Families: two-stage (Faster R-CNN) and single-stage (YOLO), plus transformer-based DETR.

Light touch only: know the vocabulary, since vision-language models now handle many detection-like tasks via prompting.

IoU(A,B)=∣A∩B∣∣A∪B∣\text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|}
A, BA,\ B
two boxes
∣A∩B∣|A \cap B|
the area where they overlap
∣A∪B∣|A \cup B|
the area covered by either box

Worked example

IoU and non-max suppression

  1. Box A covers (0,0)–(4,4), area 16. Box B covers (2,2)–(6,6), area 16.
  2. Their overlap is (2,2)–(4,4), area 4. The union is 16 + 16 − 4 = 28.
  3. IoU = 4 / 28 ≈ 0.14: little overlap.
  4. Two predictions on the same cat with IoU 0.8 and scores 0.90 and 0.75: non-max suppression keeps 0.90 and drops 0.75.
  5. A prediction usually counts as correct if its IoU with the true box is at least 0.5; mAP averages precision over classes and thresholds.

IoU measures overlap, NMS removes duplicates, and mAP summarises detection quality.

Common mistakes

  • Comparing mAP numbers computed at different IoU thresholds.
  • Training a custom detector when a vision-language model prompted with the classes would be good enough.

Check yourself

Two identical boxes: what is their IoU?Show answer

1, because the overlap equals the union. Boxes that don't touch at all have IoU 0.

What problem does non-max suppression solve?Show answer

Detectors often fire several overlapping boxes for one object. NMS keeps the most confident and removes the others that overlap it heavily.

Going deeper

Open-vocabulary detectors (OWL-ViT, Grounding DINO) detect objects described in text, combining CLIP-style alignment with detection, which is useful when classes aren't fixed in advance.

Best resources for this lesson

In plain words

Gradio builds a simple web interface around any Python function: an upload box, a text field, a results panel. Hugging Face Spaces hosts it for free and gives you a public link. A working demo makes a project far more convincing than a screenshot.

Turning your notebook into a shop window people can walk up to and try.

In detail

Gradio builds a web UI around a Python function in a few lines: inputs, outputs and examples. Hugging Face Spaces hosts it for free from a Git repo with an app.py and requirements.txt.

A live demo link turns a project from 'trust me' into 'try it'. Every portfolio project from here on should have one.

Mind the limits of free hardware: load the model once, keep it small (or quantised), cache example outputs.

Worked example

From model to public demo

  1. Write classify(img) that returns a dict of label → probability.
  2. gr.Interface(classify, gr.Image(type="pil"), gr.Label(num_top_classes=3), examples=["cat.jpg", "dog.jpg"]).launch() gives a local web app.
  3. Create a Space and push app.py and requirements.txt. It builds and serves at a public URL.
  4. Load the model once at module level, not inside classify, or every click reloads it.
  5. Enable caching for the examples: most visitors click an example, and cached results load instantly.

Wrap the function, push to a Space, load the model once and cache the examples.

python
import gradio as gr
def classify(img):
    probs = predict(img)                 # your model
    return {label: float(p) for label, p in probs}
gr.Interface(classify, gr.Image(type="pil"), gr.Label(num_top_classes=3),
             examples=["cat.jpg", "dog.jpg"]).launch()

Common mistakes

  • Loading the model inside the prediction function, so every request takes seconds.
  • A large unquantised model on free CPU hardware: the demo times out for visitors.

Check yourself

What two files does a basic Gradio Space need?Show answer

app.py with the interface and requirements.txt listing the dependencies.

When would you use gr.Blocks instead of gr.Interface?Show answer

When you need a custom layout: several inputs and outputs, tabs, or steps that react to each other.

Going deeper

ZeroGPU Spaces share GPUs on demand, so free demos of heavier models become possible. gr.Blocks builds multi-component apps beyond a single Interface.

Add example inputs and cache their outputs: most visitors click an example, and cached results load instantly.

Best resources for this lesson

In plain words

Before moving on to building GPT, write a summary of everything so far as if teaching a colleague, without looking at your notes. Wherever you hesitate or wave your hands, you've found a gap. Fix the weakest two before starting the next phase.

Like checking a map before a long drive: the parts you can't trace from memory are where you'll get lost.

In detail

Write a summary of weeks 1–21 as if teaching a colleague: one paragraph per major idea, no notes, then check against the lessons. Wherever you hesitate or hand-wave, that is a gap to revisit before the Build-a-GPT phase.

Suggested spine: data and evaluation discipline → classical models → representations (TF-IDF to embeddings) → neural network training → why attention.

Use the buffer day to fix the two weakest gaps, not to start week 23 early.

Worked example

A one-evening review

  1. Write one paragraph each on data and evaluation discipline, classical models, representations (TF-IDF to embeddings), neural network training, and why attention.
  2. Now check each paragraph against the lessons and mark the wrong or vague sentences.
  3. Typical gaps: 'why does layer norm help?' or 'what exactly does the causal mask block?'
  4. Pick the two weakest and redo the lesson, the simulation or the exercise for each.
  5. Rewrite those paragraphs. If they now come out clear and simple, the gap is closed.

Writing from memory finds the gaps; fix the two weakest before you build on them.

Common mistakes

  • Re-reading notes instead of writing from memory, which feels productive but hides the gaps.
  • Using the buffer day to start the next week early instead of fixing weak spots.

Check yourself

What is the Feynman technique?Show answer

Explain the concept in simple words as if teaching, find where the explanation breaks, return to the source, and simplify again.

Why write the review without notes?Show answer

Recall shows what you actually know; reading notes only shows what you recognise, which feels like understanding but isn't.

Going deeper

The Feynman technique: explain a concept in simple language as if teaching, find where your explanation breaks, go back to the source, and simplify again.

Best resources for this lesson

In plain words

A diffusion model learns to remove noise. During training, real images are blurred with random noise of varying strength, and the network learns to predict exactly what noise was added. To generate, start from pure noise and remove a little at a time until an image appears, with a text prompt steering each step.

A sculptor who learned, by watching statues erode, how to reverse erosion, and can now carve a statue out of a random block of stone.

In detail

A diffusion model is trained by corrupting data with increasing amounts of Gaussian noise and teaching a network to predict the noise that was added. Generation runs the process in reverse: start from random noise and repeatedly denoise.

Latent diffusion (Stable Diffusion) runs this in a VAE's compressed latent space for efficiency, and conditions each denoising step on a text embedding via cross-attention, so the text prompt steers the image.

Flow matching and diffusion transformers (DiT) are the current frontier for image and video generation. For an AI engineer the practical knowledge is the API surface (prompts, guidance scale, steps, seeds) and the cost/latency of each step.

L=Ex0,ϵ,t∥ϵ−ϵθ(αˉt x0+1−αˉt ϵ, t)∥2\mathcal{L} = \mathbb{E}_{x_0,\epsilon,t}\big\|\epsilon - \epsilon_\theta(\sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\epsilon,\ t)\big\|^2
x0x_0
the clean training image
ϵ\epsilon
the random noise that was added
tt
the noise step: higher means noisier
αˉt\bar\alpha_t
how much of the original signal survives at step tt
ϵθ(⋅,t)\epsilon_\theta(\cdot, t)
the network's guess of the noise

Worked example

One training example, one pixel

  1. Original pixel value x0=0.8x_0 = 0.8. Noise level αˉt=0.5\bar\alpha_t = 0.5. Random noise ε = −1.0.
  2. Noisy value: 0.5×0.8+0.5×(−1.0)≈0.566−0.707=−0.141\sqrt{0.5} \times 0.8 + \sqrt{0.5} \times (-1.0) \approx 0.566 - 0.707 = -0.141.
  3. The network sees −0.141 and the step t, and predicts the noise: −0.9.
  4. Loss = (−1.0−(−0.9))2=0.01(-1.0 - (-0.9))^2 = 0.01. Training pushes the prediction towards the real noise.
  5. Generation: start from pure noise and run, say, 30 denoising steps. More steps is usually better quality at higher cost and latency.

Train a network to predict added noise; generate by starting from noise and denoising step by step.

Common mistakes

  • Treating generation cost as fixed: steps, resolution and guidance settings change cost and latency a lot.
  • Assuming a fixed seed gives identical output across library versions or hardware.

Check yourself

Why run diffusion in a VAE's latent space?Show answer

The latent is far smaller than the full image, so every denoising step is much cheaper; the VAE decodes the final latent to pixels.

How does the text prompt steer the image?Show answer

A text encoder embeds the prompt, and each denoising step attends to that embedding through cross-attention; guidance scale controls how strongly.

Where this comes back

  • Week 25Multimodal models connect vision and language in the other direction.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Autoencoder on MNIST

    Train a small autoencoder, visualise reconstructions and the 2D latent space coloured by digit.

  2. Core

    Ship a demo

    Deploy your week 20 classifier to Hugging Face Spaces with Gradio, example inputs and a short description.

  3. Stretch

    Write the review

    Write a 2–3 page summary of weeks 1–21 from memory, then mark every place you had to look something up.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 148Monday15 Feb2 h planned

    Unsupervised DL: autoencoders

  2. Day 149Tuesday16 Feb2 h planned

    VAEs and representation learning

  3. Day 150Wednesday17 Feb2 h planned

    CV using DL: object detection concepts (light)

  4. Day 151Thursday18 Feb2 h planned

    Deploy the image classifier to HF Spaces with Gradio

  5. Day 152Friday19 Feb2 h planned

    Buffer / catch-up day

  6. Day 153Saturday20 Feb3 h planned

    PHASE REVIEW: write a summary of everything learned so far

  7. Day 154Sunday21 FebReview

    Review the week, finish anything unfinished, rest

Watch

Practical Deep Learning for Coders

fast.ai · Jeremy Howard · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

What's the difference between an autoencoder and a VAE?
  • AE: deterministic code, reconstruction loss
  • VAE: distribution over codes + KL to a prior
  • VAE's smooth latent space allows sampling new data
How do diffusion models generate images?
  • Trained to predict added noise at many noise levels
  • Generate by iteratively denoising from pure noise
  • Text conditioning via cross-attention; latent space for efficiency

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.An autoencoder's bottleneck forces it to…

  2. 2.What does the KL term in a VAE do?

  3. 3.IoU measures…

  4. 4.High reconstruction error from an autoencoder trained on normal data suggests…

  5. 5.Main purpose of the week 22 written review?