A colour image is a 3D array: height × width × 3 channels (red, green, blue), each value 0–255 or rescaled to 0–1. A batch of images adds a fourth dimension. PyTorch orders it (batch, channels, height, width); OpenCV and PIL use (height, width, channels), and OpenCV uses BGR order. Mixing them up is a classic bug.
Preprocessing for models: resize to the expected size, convert to float, normalise with the dataset mean and standard deviation the model was trained with.
Spend two days here, no more. The point is to make the tensor shapes in later weeks unsurprising.
import cv2
img = cv2.imread("invoice.png") # (H, W, 3), uint8, BGR
rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
x = rgb.astype("float32") / 255.0
x = x.transpose(2, 0, 1) # (3, H, W) for PyTorchGoing deeper
Normalisation constants matter: ImageNet models expect inputs normalised with specific per-channel means and standard deviations. Using the wrong ones silently degrades accuracy.
Image resolution drives compute quadratically for CNNs and ViTs alike (more patches), and for multimodal LLMs it drives token count and cost.