Concept Lab · Week 19 · Deep Learning
Self-attention, Q·K·V
Pick a word and see which other words it attends to, and why.
The idea: Self-attention with queries, keys and values
In self-attention a sequence attends to itself. Each token's embedding is projected three ways: a query (what am I looking for?), a key (what do I contain?) and a value (what do I pass on if selected?).
Token i's output is a weighted average of all values, weighted by softmax of the dot products between its query and every key. The scores are divided by to keep them from growing with dimension and saturating the softmax. In matrix form, all positions are computed at once: the parallelism RNNs lacked.
For generation, a causal mask sets scores for future positions to −∞ so each token attends only to itself and earlier tokens. Pick a word in the simulation and see where it attends.
Next simulation: Byte-pair encoding