AI Engineer Path
Phase 4 · Deep Learning11 Jan – 17 Jan

Week 17

CNNs and transfer learning

Understand convolutional networks and, more importantly, transfer learning.

  • First to cut if behind

Why this week matters

Transfer learning here is the conceptual rehearsal for LLM fine-tuning in week 33. Same idea, different scale.

First on the cut list if you're behind: keep the transfer-learning experiment, skim the architecture history.

Done when

A scratch model and a pretrained model both trained on the same task, with the gap measured.

Concepts

4 lessons · tick each one once you could explain it

A conv layer learns many kernels; each produces a feature map. Because the same kernel slides everywhere (weight sharing), the layer has few parameters and detects a pattern wherever it appears (translation equivariance).

Stride skips positions, shrinking output. Padding adds a border so kernels can centre on edge pixels. Pooling (max or average) downsamples, giving some translation invariance and a larger receptive field for later layers.

Output size: ⌊(n+2p−k)/s⌋+1\lfloor (n + 2p - k)/s \rfloor + 1 for input size n, padding p, kernel k, stride s.

nout=⌊n+2p−ks⌋+1n_{\text{out}} = \Big\lfloor \frac{n + 2p - k}{s} \Big\rfloor + 1

Going deeper

1×1 convolutions mix channels without spatial context, a cheap way to change channel count. Depthwise-separable convolutions (MobileNet) split spatial and channel mixing for big efficiency wins.

Receptive field grows with depth, stride and dilation; deep layers 'see' large parts of the image even with 3×3 kernels.

Best resources for this lesson

LeNet (1998) recognised digits; AlexNet (2012) won ImageNet with GPUs and ReLU; VGG stacked small 3×3 convolutions deeper. Beyond ~20 layers, plain networks got worse, even on training data: an optimisation problem, not overfitting.

ResNet (2015) added skip connections: each block computes F(x)F(x) and outputs x+F(x)x + F(x). Learning a small correction is easy, the identity is the default, and gradients flow straight through the additions. Networks of 100+ layers became trainable.

Residual connections are now everywhere, including around every attention and MLP sub-layer in a transformer.

y=x+F(x; W)y = x + F(x;\,W)

Going deeper

EfficientNet scales depth, width and resolution together with a compound coefficient; ConvNeXt modernised CNNs with transformer-era design choices and matches ViTs at similar compute.

For most practical vision tasks today, start from a pretrained backbone (ResNet, ConvNeXt, ViT) and fine-tune. Architecture search from scratch is rarely worth it.

Where this comes back

  • Week 19Transformer blocks wrap every sub-layer in a residual connection.

Random crops, flips, small rotations, colour jitter and blur teach invariances: a cat flipped horizontally is still a cat. It acts as a regulariser and is especially valuable with small datasets.

Choose augmentations that preserve the label. Flipping a '6' or a road sign can change its meaning; heavy crops can cut the object out.

Augment only the training set. Validation and test data stay untouched.

Going deeper

Mixup and CutMix blend images and labels, acting as strong regularisers. RandAugment and TrivialAugment pick random augmentation policies with few hyperparameters.

Test-time augmentation (average predictions over flipped or cropped inputs) buys a little accuracy at extra inference cost.

Best resources for this lesson

A network trained on ImageNet has already learned general visual features: edges, textures, shapes. For a new task, replace its final classification layer and train on your data.

Feature extraction: freeze the backbone, train only the new head. Fast, works with little data. Fine-tuning: also unfreeze some or all backbone layers with a small learning rate, typically later layers first, since they are the most task-specific.

With a few thousand images, a fine-tuned pretrained model usually beats a scratch model by a wide margin. Measuring that gap is this week's exercise, and the same reasoning (pretrain broadly, adapt cheaply) underlies LLM fine-tuning and LoRA.

python
import torch.nn as nn, torchvision
m = torchvision.models.resnet18(weights="IMAGENET1K_V1")
for p in m.parameters():
    p.requires_grad = False            # freeze backbone
m.fc = nn.Linear(m.fc.in_features, num_classes)  # new trainable head

Going deeper

Discriminative learning rates (smaller for early layers, larger for the head) and gradual unfreezing (unfreeze top layers first) are fast.ai's practical recipe and work well with small data.

Linear probing (frozen backbone, train only a linear head) is also the standard way to measure how good a representation is, and appears in embedding-model evaluation.

Where this comes back

  • Week 21Fine-tuning BERT is the same move for text.
  • Week 33LoRA fine-tunes LLMs by training small adapters on a frozen base.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Output-size drill

    For 5 conv configurations (kernel, stride, padding), predict output shapes by formula, then verify in PyTorch.

  2. Core

    Scratch vs pretrained

    Train a small CNN from scratch and fine-tune a pretrained ResNet on the same 2,000-image dataset. Report accuracy, training time and the gap.

  3. Stretch

    Augmentation ablation

    Measure the effect of 4 augmentation settings on validation accuracy with everything else fixed.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 113Monday11 Jan2 h planned

    Convolution, stride, padding, pooling

  2. Day 114Tuesday12 Jan2 h planned

    CNN architectures: LeNet -> VGG -> ResNet

  3. Day 115Wednesday13 Jan2 h planned

    Data augmentation

  4. Day 116Thursday14 Jan2 h planned

    Transfer learning: freezing layers vs fine-tuning

  5. Day 117Friday15 Jan2 h planned

    Build an image classifier from scratch

  6. Day 118Saturday16 Jan3 h planned

    Rebuild it with a pretrained ResNet; compare honestly

  7. Day 119Sunday17 JanReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Deep LearningPrimary

CampusX · playlist

But what is a convolution?

3Blue1Brown

Neural Networks / Deep Learning

StatQuest · playlist

Convolutional Neural Networks (DLS Course 4)

Andrew Ng · DeepLearning.AI · playlist

Image Classification with CNNs

StatQuest

NYU Deep Learning (LeCun & Canziani)

Alfredo Canziani · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

When would you freeze a pretrained backbone vs fine-tune it fully?
  • Small data or similar domain → freeze, train head
  • More data or different domain → unfreeze progressively with small LR
  • Watch for overfitting and forgetting
Why did ResNets enable much deeper networks?
  • Skip connections: learn residuals, identity by default
  • Gradients flow directly through additions
  • Fixed the degradation problem of very deep plain nets

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why do convolutional layers have few parameters?

  2. 2.ResNet's key idea is…

  3. 3.You have 800 labelled images. Best starting point?

  4. 4.Which augmentation could change the label for digit recognition?

  5. 5.Input 32×32, kernel 3, padding 1, stride 1. Output size?