AI Engineer Path

Week 20

PyTorch

Write idiomatic PyTorch training code.

Why this week matters

The entire LLM ecosystem is PyTorch. This isn't a preference.

Done when

The week 17 classifier rewritten in idiomatic PyTorch.

Concepts

6 lessons · tick each one once you could explain it

In plain words

A PyTorch tensor is like a NumPy array that can live on a GPU and remembers the operations used to build it. Mark a tensor as needing gradients, compute a loss from it, call .backward(), and PyTorch fills in the gradient for you, which is exactly the backprop you did by hand.

A tensor with requires_grad is a receipt that lists every step of the calculation, so the bill can be traced back to each item.

In detail

A torch.Tensor is an n-dimensional array with a dtype and a device (cpu, cuda, mps). The API mirrors NumPy, broadcasting included. Move data and model to the same device or you get the most common PyTorch error.

Tensors with requires_grad=True record the operations applied to them in a dynamic graph. Calling .backward() on a scalar runs backpropagation and fills .grad on every leaf tensor. This is the engine you'll rebuild from scratch in week 23.

Gradients accumulate across .backward() calls, so zero them every step. Wrap inference in torch.no_grad() (or torch.inference_mode()) to skip graph-building.

Worked example

One gradient, computed by PyTorch

  1. w = (2, −1) with requires_grad=True, x = (3, 4).
  2. Forward: w⋅x=6−4=2w \cdot x = 6 - 4 = 2, and loss =(2−1)2=1= (2 - 1)^2 = 1.
  3. loss.backward() fills w.grad with 2(w⋅x−1) x=2×1×(3,4)=(6,8)2(w\cdot x - 1)\,x = 2 \times 1 \times (3, 4) = (6, 8).
  4. Run the forward and backward again without zeroing and w.grad becomes (12, 16): gradients accumulate.
  5. Model on the GPU, data on the CPU: you get 'Expected all tensors to be on the same device'. Move both with .to(device).

Tensors record their history, so .backward() gives exact gradients. Zero them each step and keep everything on one device.

python
import torch
w = torch.tensor([2.0, -1.0], requires_grad=True)
x = torch.tensor([3.0, 4.0])
loss = ((w * x).sum() - 1) ** 2
loss.backward()
print(w.grad)   # d(loss)/dw = 2*(w·x - 1) * x = tensor([6., 8.])

Common mistakes

  • Forgetting to zero gradients between steps, so they add up across batches.
  • Running evaluation without torch.no_grad(), which wastes memory building graphs you never use.

Check yourself

What does .detach() do, and when do you need it?Show answer

It returns a tensor cut out of the computation graph, so no gradient flows through it. Use it for targets, logged values and frozen features.

Does model.eval() stop gradients from being computed?Show answer

No. It only switches layers like dropout and batch norm to inference behaviour. Use torch.no_grad() or torch.inference_mode() to skip gradients.

Going deeper

.detach() cuts a tensor out of the graph (used for targets, logging and frozen features); .requires_grad_(False) freezes parameters. In-place operations on tensors needed for backward raise errors, which is PyTorch protecting you from wrong gradients.

torch.compile traces your model into an optimised graph, often giving large speedups with one line. It's worth trying once your code is correct.

Where this comes back

  • Week 23micrograd is a 100-line version of exactly this.

In plain words

In PyTorch, a model is a class that inherits from nn.Module. You create its layers in __init__ and describe how data flows in forward. PyTorch then finds every weight automatically, so one call moves the whole model to the GPU and one call hands all its weights to the optimiser.

An nn.Module is a set of nesting boxes: open the outer box and every smaller box, and every weight inside, is registered and reachable.

In detail

Subclass nn.Module, create layers in __init__, and define computation in forward. Sub-modules register their parameters automatically, so model.parameters() finds them all and model.to(device) moves them all.

Losses live in torch.nn (CrossEntropyLoss takes raw logits and integer labels). Optimisers live in torch.optim (AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)).

model.train() and model.eval() switch dropout and batch-norm behaviour. They don't stop gradients; no_grad does that.

Worked example

Counting a small MLP's parameters

  1. MLP(784, 256, 10): Linear(784 → 256), ReLU, Dropout, Linear(256 → 10).
  2. First layer: 784 × 256 weights + 256 biases = 200,960.
  3. Second layer: 256 × 10 + 10 = 2,570. Total: 203,530.
  4. sum(p.numel() for p in model.parameters()) prints 203,530, so all weights were found.
  5. Store layers in a plain Python list instead and parameters() finds none of them; the optimiser silently never updates them. Use nn.ModuleList.

Build layers in __init__, compute in forward, and let PyTorch register parameters through Module containers.

python
import torch.nn as nn
class MLP(nn.Module):
    def __init__(self, d_in, d_hidden, n_classes):
        super().__init__()
        self.net = nn.Sequential(nn.Linear(d_in, d_hidden), nn.ReLU(),
                                 nn.Dropout(0.1), nn.Linear(d_hidden, n_classes))
    def forward(self, x):
        return self.net(x)          # logits

Common mistakes

  • Keeping layers in plain lists or dicts: their parameters aren't registered or trained.
  • Applying softmax in forward and then using CrossEntropyLoss, which expects raw logits.

Check yourself

What does register_buffer do?Show answer

It stores a tensor that isn't a trainable parameter (such as a mask or running statistics) but should still move with .to(device) and be saved in the state dict.

Why put biases and LayerNorm weights in a separate optimiser group?Show answer

So you can exclude them from weight decay, which is meant for the large weight matrices; decaying biases and norm scales usually hurts.

Going deeper

Register non-parameter tensors (masks, running stats) with register_buffer so they move with .to(device) and save with the state dict. Use nn.ModuleList/ModuleDict, never plain lists, or parameters won't be found.

Parameter groups in the optimiser let you set different learning rates or weight decay per group, such as no decay on biases and LayerNorm weights.

Best resources for this lesson

In plain words

A Dataset knows how to fetch one example by its index. A DataLoader wraps it to serve shuffled batches, loading them in parallel worker processes so the GPU isn't kept waiting. For text, a collate function pads sequences in a batch to the same length.

The Dataset is the warehouse that can fetch any single item; the DataLoader is the delivery service that packs items into boxes and keeps the trucks coming.

In detail

Implement __len__ and __getitem__ on a Dataset to return one (input, label) pair, applying per-example transforms like augmentation there. A DataLoader wraps it with batching, shuffling (training only), parallel num_workers and pin_memory for faster GPU transfer.

A custom collate_fn handles variable-length data, such as padding token sequences to the longest in the batch and returning an attention mask.

Data loading is often the real bottleneck: if GPU utilisation is low, profile the loader first.

Worked example

Batching 10,000 examples of uneven length

  1. 10,000 examples with batch size 32 gives 313 batches per epoch (the last batch has 16).
  2. shuffle=True for training so each epoch sees a new order; no shuffle for validation.
  3. Three sequences of lengths 5, 3 and 7 in one batch: the collate function pads all to 7.
  4. Attention masks mark the real tokens: (1,1,1,1,1,0,0), (1,1,1,0,0,0,0) and (1,1,1,1,1,1,1).
  5. GPU utilisation at 30%? Raise num_workers and set pin_memory=True; often the loader, not the model, is the bottleneck.

The Dataset returns one item, the DataLoader batches, shuffles and parallelises, and collate handles padding.

python
from torch.utils.data import Dataset, DataLoader
class TextDS(Dataset):
    def __init__(self, enc, labels): self.enc, self.labels = enc, labels
    def __len__(self): return len(self.labels)
    def __getitem__(self, i): return self.enc[i], self.labels[i]
loader = DataLoader(TextDS(enc, y), batch_size=32, shuffle=True, num_workers=4)

Common mistakes

  • Shuffling the validation or test loader, which makes per-example debugging confusing (the metrics are unaffected).
  • Doing heavy preprocessing in __getitem__ on every epoch when it could be done once and cached.

Check yourself

Why return an attention mask alongside padded sequences?Show answer

So the model ignores padding positions: attention gives them zero weight and the loss skips them.

Your GPU sits at 20% utilisation during training. What do you check first?Show answer

The data pipeline: profile the DataLoader, raise num_workers, enable pin_memory, and move preprocessing out of the training loop.

Going deeper

num_workers > 0 loads batches in parallel processes; persistent_workers=True avoids re-spawning them each epoch; pin_memory=True speeds host-to-GPU copies. Profile before tuning.

For huge datasets, streaming (IterableDataset, Hugging Face streaming, WebDataset shards) avoids loading everything into memory or disk.

Best resources for this lesson

In plain words

Every PyTorch training loop repeats the same five steps for each batch: clear old gradients, run the model, compute the loss, backpropagate, and update the weights. After each epoch you evaluate on validation data and save a checkpoint if it's the best so far, so you can resume or roll back.

A training loop is a rehearsal routine: perform, get notes, adjust, repeat, and record the best run.

In detail

The canonical loop per batch: move the batch to the device, optimizer.zero_grad(), compute logits, compute the loss, loss.backward(), optionally clip gradients, optimizer.step(), then scheduler.step().

After each epoch, evaluate under model.eval() and torch.no_grad(), track the validation metric, and save a checkpoint (model and optimiser state dicts, epoch, best metric) when it improves, so you can resume or roll back.

Seed everything (torch.manual_seed, NumPy, Python's random) for reproducibility, and log metrics to a tracker (Weights & Biases or MLflow).

Worked example

Three epochs with checkpointing

  1. Per batch: move to device → optimizer.zero_grad() → loss = loss_fn(model(xb), yb) → loss.backward() → clip gradients → optimizer.step() → scheduler.step().
  2. After epoch 1, under model.eval() and torch.no_grad(), validation F1 = 0.81. Save a checkpoint.
  3. Epoch 2: F1 = 0.85, a new best, so overwrite the checkpoint.
  4. Epoch 3: F1 = 0.84, worse, so keep the epoch-2 checkpoint.
  5. The checkpoint holds model and optimiser state dicts, the epoch and the best metric, so you can resume exactly where you left off.

zero_grad → forward → loss → backward → step, then evaluate in eval mode and checkpoint the best.

python
for epoch in range(epochs):
    model.train()
    for xb, yb in train_loader:
        xb, yb = xb.to(device), yb.to(device)
        optimizer.zero_grad(set_to_none=True)
        loss = loss_fn(model(xb), yb)
        loss.backward()
        torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
        optimizer.step(); scheduler.step()
    val = evaluate(model, val_loader)
    if val > best:
        best = val
        torch.save({"model": model.state_dict(), "opt": optimizer.state_dict(),
                    "epoch": epoch, "val": val}, "best.pt")

Common mistakes

  • Forgetting zero_grad: gradients accumulate across batches.
  • Evaluating without model.eval().

Check yourself

Why save the optimiser's state as well as the model's?Show answer

AdamW keeps running averages per weight. Without them, resuming restarts those statistics and training behaves as if freshly started.

What three things make a run reproducible?Show answer

Seeds for PyTorch, NumPy and Python's random module (plus DataLoader workers), deterministic settings where possible, and recorded library versions.

Going deeper

Reproducibility needs more than a seed: set torch.backends.cudnn.deterministic, seed DataLoader workers, and record library versions. Even then, GPU non-determinism can cause small differences.

High-level libraries (PyTorch Lightning, Hugging Face Accelerate) remove loop boilerplate and add multi-GPU and mixed precision. Write the loop by hand once first so you know what they do.

In plain words

Computers can store numbers in 32 bits or, with less precision, in 16. Doing most of the training maths in 16-bit roughly doubles speed and halves memory, while a few sensitive operations stay in 32-bit. bfloat16 is the usual choice for LLMs because it handles the same range of sizes as 32-bit.

Doing rough sums on a calculator with fewer decimal places: much faster, and fine as long as you keep the final ledger in full precision.

In detail

Mixed precision runs most operations in float16 or bfloat16 while keeping master weights and sensitive operations (reductions, softmax) in float32. Use torch.autocast; with float16 add a GradScaler to prevent small gradients underflowing. bfloat16 has float32's range and usually needs no scaler: it is the default for LLM training.

Debugging checklist: overfit one batch; check for NaN/inf in the loss; print shapes; verify labels and inputs line up; check the learning rate actually used.

torch.profiler shows where time goes. Common culprits: the data loader, CPU↔GPU copies inside the loop, and .item() calls that force synchronisation every step.

Worked example

Memory and underflow in numbers

  1. A 1-billion-parameter model: weights take 4 GB in float32 and 2 GB in bfloat16.
  2. float16's smallest positive value is about 6 × 10⁻⁸. A gradient of 10⁻⁸ rounds to zero (underflow), and learning silently stops for that weight.
  3. A GradScaler multiplies the loss by, say, 65,536, so that gradient becomes about 6.6 × 10⁻⁴, which is representable. It's divided back before the update.
  4. bfloat16 covers the same range as float32 (up to about 3 × 10³⁸) with fewer significant digits, so it usually needs no scaler.
  5. Profile: if .item() is called every step, each call forces the GPU to finish; log every 100 steps instead.

Use autocast with bfloat16 for speed and memory; scale the loss if you must use float16; profile before optimising.

Common mistakes

  • Using float16 without a GradScaler: small gradients underflow to zero.
  • Optimising code before profiling: the slow part is often the data loader or CPU–GPU syncs.

Check yourself

Why is bfloat16 preferred over float16 for LLM training?Show answer

It has float32's exponent range, so gradients rarely underflow or overflow; it just has fewer digits of precision, which training tolerates.

Which operations are usually kept in float32 under autocast?Show answer

Numerically sensitive ones such as reductions (sums, means), softmax, layer norm and the master copy of the weights.

Going deeper

FP8 training (on recent GPUs) pushes precision further for large-model training; inference increasingly runs in 8-bit or 4-bit formats (week 26).

Tensor cores need dimensions that are multiples of 8 (or 16, or 64) to run at full speed: one reason model widths are round numbers.

Where this comes back

  • Week 26Precision and quantisation determine the VRAM an LLM needs.

In plain words

When a batch is too big for GPU memory, gradient accumulation runs several small batches and adds up their gradients before taking one step. With several GPUs, data parallelism gives each GPU a copy of the model and part of the batch. When the model itself doesn't fit, sharding (FSDP or ZeRO) splits its weights and optimiser state across GPUs.

Accumulation is paying for a big purchase in instalments; data parallelism is several cashiers serving one queue; sharding is splitting a huge book across several shelves and fetching pages as needed.

In detail

Gradient accumulation simulates a large batch on small memory: run several micro-batches, sum their gradients, then step the optimiser once. The maths matches a big batch (scale the loss accordingly).

Distributed Data Parallel (DDP) puts a full model copy on each GPU, splits each batch across them and averages gradients with an all-reduce. It's simple and efficient when the model fits on one GPU.

FSDP / ZeRO shards parameters, gradients and optimiser state across GPUs, gathering each layer only when needed. That's how models far larger than one GPU's memory are trained. Tensor and pipeline parallelism split individual layers and layer stacks for the very largest models.

Worked example

Reaching an effective batch of 256

  1. Only 32 examples fit on one GPU. Accumulate over 8 micro-batches (dividing each loss by 8), then step once: an effective batch of 256.
  2. With 4 GPUs under DDP, each takes 32 per micro-batch and accumulates 2: 4 × 32 × 2 = 256, finishing about 4× sooner.
  3. A 7-billion-parameter model with AdamW needs roughly 16 bytes per parameter for weights, gradients and optimiser state: about 112 GB, more than any single GPU.
  4. FSDP shards that across 8 GPUs, about 14 GB each plus activations, gathering each layer's full weights only while it's being computed.
  5. Clip gradients and step the scheduler once per optimiser step, not once per micro-batch.

Accumulate when the batch doesn't fit, replicate (DDP) when the model fits, and shard (FSDP/ZeRO) when it doesn't.

python
accum = 8
for i, (xb, yb) in enumerate(loader):
    with torch.autocast("cuda", dtype=torch.bfloat16):
        loss = loss_fn(model(xb.cuda()), yb.cuda()) / accum
    loss.backward()
    if (i + 1) % accum == 0:
        torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
        optimizer.step(); optimizer.zero_grad(set_to_none=True)

Common mistakes

  • Forgetting to divide the loss by the number of accumulation steps, which inflates the effective learning rate.
  • Stepping the scheduler every micro-batch, so the learning rate decays far too fast.

Check yourself

Is gradient accumulation mathematically the same as a bigger batch?Show answer

Essentially yes, for the gradient. The exception is batch-norm statistics, which still see only the micro-batch.

When do you need FSDP rather than DDP?Show answer

When one GPU can't hold the full model plus its gradients and optimiser state. DDP needs a full copy per GPU; FSDP splits it.

Where this comes back

  • Week 25nanoGPT trains with DDP and gradient accumulation.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Autograd sanity check

    Recompute your week 15 network's gradients with autograd and confirm they match your manual NumPy gradients.

  2. Core

    Idiomatic rewrite

    Rewrite the week 17 classifier with Dataset/DataLoader, nn.Module, AdamW, cosine schedule, AMP, checkpointing and a validation loop.

  3. Stretch

    Profile and speed up

    Profile one epoch, find the bottleneck (often the loader), fix it, and try torch.compile. Report before/after throughput.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 134Monday1 Feb2 h planned

    Tensors and autograd

  2. Day 135Tuesday2 Feb2 h planned

    nn.Module, layers, losses, optimisers

  3. Day 136Wednesday3 Feb2 h planned

    Dataset, DataLoader, transforms

  4. Day 137Thursday4 Feb2 h planned

    Training loop, checkpointing, device handling

  5. Day 138Friday5 Feb2 h planned

    Mixed precision, debugging, profiling

  6. Day 139Saturday6 Feb3 h planned

    Reimplement the week 17 classifier in idiomatic PyTorch

  7. Day 140Sunday7 FebReview

    Review the week, finish anything unfinished, rest

Watch

PyTorch for Deep Learning & Machine Learning: Full CoursePrimary

freeCodeCamp

Practical Deep Learning for Coders

fast.ai · Jeremy Howard · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Write out a PyTorch training step and explain each line.
  • zero_grad → forward → loss → backward → (clip) → step
  • model.train()/eval() and no_grad for validation
  • Device placement, AMP, checkpointing
Your model doesn't fit in GPU memory. What are your options?
  • Smaller batch + gradient accumulation; mixed precision
  • Gradient checkpointing
  • FSDP/ZeRO sharding; quantised or LoRA fine-tuning

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why call optimizer.zero_grad() each step?

  2. 2.nn.CrossEntropyLoss expects…

  3. 3.What does torch.no_grad() do?

  4. 4.Why is bfloat16 popular for LLM training?

  5. 5.GPU utilisation sits at 30%. First suspect?