AI Engineer Path

Concept Lab · Week 19 · Deep Learning

Self-attention, Q·K·V

Pick a word and see which other words it attends to, and why.

The idea: Self-attention with queries, keys and values

In self-attention a sequence attends to itself. Each token's embedding is projected three ways: a query (what am I looking for?), a key (what do I contain?) and a value (what do I pass on if selected?).

Token i's output is a weighted average of all values, weighted by softmax of the dot products between its query and every key. The scores are divided by dk\sqrt{d_k} to keep them from growing with dimension and saturating the softmax. In matrix form, all positions are computed at once: the parallelism RNNs lacked.

For generation, a causal mask sets scores for future positions to −∞ so each token attends only to itself and earlier tokens. Pick a word in the simulation and see where it attends.

Open the full lesson in week 19

Next simulation: Byte-pair encoding