Concept Lab · Week 33 · AI Engineer Track
LoRA: low-rank updates
Compare a full weight update with W + BA and count the parameters saved.
The idea: LoRA and QLoRA
Full fine-tuning updates every weight: billions of parameters with gradients and optimiser state for each. LoRA freezes the pretrained weight and learns an update , where is and is with a small rank (8–64). Trainable parameters drop by 100–1000×.
This works because the adjustment needed for a new task appears to be low-rank: the week 4 linear algebra, applied. Adapters are small files that can be swapped per task over one base model, or merged into the weights for zero inference overhead.
QLoRA loads the frozen base in 4-bit precision and trains LoRA adapters in 16-bit, so a 7–8B model fine-tunes on a single consumer or free Colab GPU. Count the parameters saved in the simulation.
Next simulation: The ReAct agent loop