A conv layer learns many kernels; each produces a feature map. Because the same kernel slides everywhere (weight sharing), the layer has few parameters and detects a pattern wherever it appears (translation equivariance).
Stride skips positions, shrinking output. Padding adds a border so kernels can centre on edge pixels. Pooling (max or average) downsamples, giving some translation invariance and a larger receptive field for later layers.
Output size: for input size n, padding p, kernel k, stride s.
Going deeper
1×1 convolutions mix channels without spatial context, a cheap way to change channel count. Depthwise-separable convolutions (MobileNet) split spatial and channel mixing for big efficiency wins.
Receptive field grows with depth, stride and dilation; deep layers 'see' large parts of the image even with 3×3 kernels.
Best resources for this lesson
- CourseCS231n: convolutional neural networks
- InteractiveCNN Explainer