In plain words
A PyTorch tensor is like a NumPy array that can live on a GPU and remembers the operations used to build it. Mark a tensor as needing gradients, compute a loss from it, call .backward(), and PyTorch fills in the gradient for you, which is exactly the backprop you did by hand.
A tensor with requires_grad is a receipt that lists every step of the calculation, so the bill can be traced back to each item.
In detail
A torch.Tensor is an n-dimensional array with a dtype and a device (cpu, cuda, mps). The API mirrors NumPy, broadcasting included. Move data and model to the same device or you get the most common PyTorch error.
Tensors with requires_grad=True record the operations applied to them in a dynamic graph. Calling .backward() on a scalar runs backpropagation and fills .grad on every leaf tensor. This is the engine you'll rebuild from scratch in week 23.
Gradients accumulate across .backward() calls, so zero them every step. Wrap inference in torch.no_grad() (or torch.inference_mode()) to skip graph-building.
Worked example
One gradient, computed by PyTorch
w = (2, −1)withrequires_grad=True,x = (3, 4).- Forward: , and loss .
loss.backward()fillsw.gradwith .- Run the forward and backward again without zeroing and
w.gradbecomes (12, 16): gradients accumulate. - Model on the GPU, data on the CPU: you get 'Expected all tensors to be on the same device'. Move both with
.to(device).
Tensors record their history, so .backward() gives exact gradients. Zero them each step and keep everything on one device.
import torch
w = torch.tensor([2.0, -1.0], requires_grad=True)
x = torch.tensor([3.0, 4.0])
loss = ((w * x).sum() - 1) ** 2
loss.backward()
print(w.grad) # d(loss)/dw = 2*(w·x - 1) * x = tensor([6., 8.])Common mistakes
- Forgetting to zero gradients between steps, so they add up across batches.
- Running evaluation without
torch.no_grad(), which wastes memory building graphs you never use.
Check yourself
What does .detach() do, and when do you need it?Show answer
It returns a tensor cut out of the computation graph, so no gradient flows through it. Use it for targets, logged values and frozen features.
Does model.eval() stop gradients from being computed?Show answer
No. It only switches layers like dropout and batch norm to inference behaviour. Use torch.no_grad() or torch.inference_mode() to skip gradients.
Going deeper
.detach() cuts a tensor out of the graph (used for targets, logging and frozen features); .requires_grad_(False) freezes parameters. In-place operations on tensors needed for backward raise errors, which is PyTorch protecting you from wrong gradients.
torch.compile traces your model into an optimised graph, often giving large speedups with one line. It's worth trying once your code is correct.
Best resources for this lesson
Where this comes back
- Week 23micrograd is a 100-line version of exactly this.