AI Engineer Path

Week 14

Computer vision (trimmed) + time series (awareness)

Just enough classical CV, then the multimodal models that matter now. Plus three time-series rules.

Why this week matters

SIFT and Hough transforms connect to nothing downstream. Visual understanding via multimodal models connects to a lot, and invoice/form extraction is one of the highest-demand real use cases right now.

Done when

Project 3 turns a document image into structured JSON, and you can state the time-series split rule in one sentence.

Milestone: Project 3: document → JSON via a vision model

Concepts

5 lessons · tick each one once you could explain it

A colour image is a 3D array: height × width × 3 channels (red, green, blue), each value 0–255 or rescaled to 0–1. A batch of images adds a fourth dimension. PyTorch orders it (batch, channels, height, width); OpenCV and PIL use (height, width, channels), and OpenCV uses BGR order. Mixing them up is a classic bug.

Preprocessing for models: resize to the expected size, convert to float, normalise with the dataset mean and standard deviation the model was trained with.

Spend two days here, no more. The point is to make the tensor shapes in later weeks unsurprising.

python
import cv2
img = cv2.imread("invoice.png")           # (H, W, 3), uint8, BGR
rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
x = rgb.astype("float32") / 255.0
x = x.transpose(2, 0, 1)                  # (3, H, W) for PyTorch

Going deeper

Normalisation constants matter: ImageNet models expect inputs normalised with specific per-channel means and standard deviations. Using the wrong ones silently degrades accuracy.

Image resolution drives compute quadratically for CNNs and ViTs alike (more patches), and for multimodal LLMs it drives token count and cost.

Best resources for this lesson

A convolution slides a small grid of weights (a kernel, e.g. 3×3) over the image. At each position it multiplies the kernel with the patch underneath and sums the result into one output pixel.

Hand-designed kernels do classic image processing: a blur averages neighbours; a Sobel kernel responds to horizontal or vertical edges; sharpening amplifies the centre relative to neighbours. Try them in the simulation.

CNNs (week 17) learn their kernels from data instead: early layers learn edge detectors that look much like Sobel; deeper layers learn textures, parts and objects.

(I∗K)(i,j)=∑m∑nI(i+m, j+n) K(m,n)(I * K)(i, j) = \sum_{m}\sum_{n} I(i+m,\ j+n)\,K(m, n)

Going deeper

Technically, deep-learning 'convolution' is cross-correlation (the kernel isn't flipped); since kernels are learned, the distinction doesn't matter.

Convolution is equivalent to multiplying by a large, sparse, weight-shared matrix, which links CNNs back to the linear algebra of week 4.

Where this comes back

  • Week 17CNNs stack learned convolutions.

A Vision Transformer (ViT) cuts an image into 16×16 patches, flattens each into a vector, adds position information and feeds the sequence into a standard transformer, exactly as if the patches were words.

CLIP trains an image encoder and a text encoder together on hundreds of millions of image–caption pairs with a contrastive objective: matching pairs get high cosine similarity, mismatched pairs low. The result is a shared space where an image and a sentence describing it land close together.

This enables zero-shot classification ('a photo of a cat' vs 'a photo of a dog'), text-to-image search, and it is the vision front-end idea behind most multimodal LLMs.

L=−1N∑ilog⁡exp⁡(cos⁡(Ii,Ti)/τ)∑jexp⁡(cos⁡(Ii,Tj)/τ)\mathcal{L} = -\frac{1}{N}\sum_{i}\log\frac{\exp(\cos(I_i, T_i)/\tau)}{\sum_j \exp(\cos(I_i, T_j)/\tau)}

Going deeper

ViTs lack CNNs' built-in locality bias, so they need much more data (or strong augmentation and distillation) to match CNNs, but they scale better and share architecture with text models.

CLIP-style embeddings power image search, zero-shot classification and content moderation. SigLIP replaced the softmax contrastive loss with a sigmoid one that scales more easily.

Where this comes back

  • Week 25Large multimodal models connect a vision encoder like this to an LLM.

Traditional pipelines ran OCR, then layout analysis, then rules or NER to pull fields out. Modern vision-language models read the page image directly and can return structured output against a schema you provide.

Project 3 pattern: define a Pydantic schema (vendor, date, line items, total), send the image with instructions and the JSON Schema, validate the response, and retry with the validation error if it fails.

Evaluate honestly: hand-label 20–50 documents and measure field-level accuracy (exact match for IDs, tolerance for amounts), plus cost and latency per page. Check arithmetic (line items sum to the total) as a free consistency test.

Going deeper

OCR-free document models (Donut, and general VLMs) read pixels directly; OCR-plus-LLM pipelines are cheaper and more controllable for clean scans. Hybrid approaches pass both the OCR text and the image.

Bounding-box grounding (returning where each field was found) makes human review fast and builds trust in extraction systems.

Common pitfalls

  • Trusting extracted numbers without validation.
  • Evaluating on the same five documents you used to write the prompt.

Where this comes back

  • Week 28Project 4 generalises this into a robust extraction service.

Rule 1: order matters. Observations are not independent: today depends on yesterday. Shuffling destroys the signal and breaks most standard assumptions.

Rule 2: split chronologically, never randomly. Train on the past, validate on the future, using expanding or sliding windows (TimeSeriesSplit). A random split lets the model peek at the future and makes results look far better than they will be. This is the one most likely to come up in an interview.

Rule 3: in production, the winning answer is usually gradient boosting on lag features: yesterday's value, last week's value, rolling means, calendar features, all computed only from information available at prediction time.

python
df["lag_1"] = df["sales"].shift(1)
df["lag_7"] = df["sales"].shift(7)
df["roll_7"] = df["sales"].shift(1).rolling(7).mean()  # shift first: no peeking
from sklearn.model_selection import TimeSeriesSplit
cv = TimeSeriesSplit(n_splits=5)

Going deeper

Backtesting with a rolling origin (train up to month t, test on month t+1, then roll forward) shows how a forecaster performs over time, not just once.

Foundation models for time series (TimesFM, Chronos) offer zero-shot forecasts, but a tuned gradient-boosting model on good lag features remains a hard baseline to beat.

Common pitfalls

  • Rolling features computed without shifting first, which includes the value being predicted.

Best resources for this lesson

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Kernel playground

    Apply blur, sharpen and Sobel kernels to your own photo with OpenCV and explain each result.

  2. Core

    Project 3: invoice → JSON

    Extract structured fields from 20 document images with a vision model and Pydantic validation. Report field-level accuracy, cost per page and the three worst failures.

  3. Stretch

    Zero-shot classification with CLIP

    Classify 100 images into 5 categories using CLIP text prompts. Try prompt variations and report the accuracy change.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 92Monday21 Dec2 h planned

    Image basics: pixels, channels, tensors; OpenCV intro

  2. Day 93Tuesday22 Dec2 h planned

    Convolution, filters, edge detection, transformations (skim only)

  3. Day 94Wednesday23 Dec2 h planned

    VLM concepts: CLIP and Vision Transformers

  4. Day 95Thursday24 Dec2 h planned

    PROJECT 3: document image -> structured JSON with a vision model

  5. Day 96Friday25 Dec2 h planned

    PROJECT 3: finish and evaluate

  6. Day 97Saturday26 Dec3 h planned

    Time Series AWARENESS (half day): order matters, chronological split, lag features, GBM baseline. Then buffer.

  7. Day 98Sunday27 DecReview

    Review the week, finish anything unfinished, rest

Watch

But what is a convolution?

3Blue1Brown

OpenAI CLIP: Connecting Text and Images

Yannic Kilcher

An Image is Worth 16x16 Words (ViT)

Yannic Kilcher

Project 3

Document understanding via a vision model

Turn a document image (invoice, receipt or form) into validated structured JSON.

  • Pydantic schema for the target document type
  • Vision-language model call that returns schema-valid JSON
  • Field-level accuracy on a small hand-labelled set
  • Cost per document and latency recorded
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How would you split data for a demand-forecasting model?
  • Chronologically: train on past, validate on future
  • Rolling-origin backtests
  • Lag features computed only from past data
OCR pipeline vs vision-language model for document extraction?
  • OCR + rules/LLM: cheaper, controllable, good on clean scans
  • VLM: handles layout and messy docs, costs more
  • Validate outputs, measure field accuracy, keep humans in the loop

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.A PyTorch image batch tensor is usually shaped…

  2. 2.CLIP is trained to…

  3. 3.State the time-series split rule.

  4. 4.A ViT treats an image as…

  5. 5.A free consistency check for invoice extraction?