AI Engineer Path

Concept Lab · Week 33 · AI Engineer Track

LoRA: low-rank updates

Compare a full weight update with W + BA and count the parameters saved.

The idea: LoRA and QLoRA

Full fine-tuning updates every weight: billions of parameters with gradients and optimiser state for each. LoRA freezes the pretrained weight WW and learns an update ΔW=BA\Delta W = BA, where BB is d×rd \times r and AA is r×kr \times k with a small rank rr (8–64). Trainable parameters drop by 100–1000×.

This works because the adjustment needed for a new task appears to be low-rank: the week 4 linear algebra, applied. Adapters are small files that can be swapped per task over one base model, or merged into the weights for zero inference overhead.

QLoRA loads the frozen base in 4-bit precision and trains LoRA adapters in 16-bit, so a 7–8B model fine-tunes on a single consumer or free Colab GPU. Count the parameters saved in the simulation.

Open the full lesson in week 33

Next simulation: The ReAct agent loop