AI Engineer Path
Phase 4 · Deep Learning25 Jan – 31 Jan

Week 19

Seq2seq and attention

Understand the transformer block completely.

  • Hinge week: slow down here
  • Never cut

Why this week matters

Everything in the AI engineer track descends from this week: LLMs, embeddings, RAG, agents. If you're behind, slow down here rather than skim.

Done when

You can draw a transformer block from memory and explain what each component is for.

Milestone: Draw the transformer block from memory

Concepts

7 lessons · tick each one once you could explain it

Sequence-to-sequence models map an input sequence to an output sequence of different length: translation, summarisation. The encoder reads the input; the decoder generates the output one token at a time, each step conditioned on the tokens generated so far.

During training, the decoder is fed the correct previous tokens (teacher forcing); at inference it consumes its own outputs, decoded greedily or with beam search.

Transformers keep this split. The original transformer was encoder–decoder (like T5 today). BERT is encoder-only (understanding), GPT is decoder-only (generation).

Going deeper

Beam search keeps the k best partial sequences at each step instead of only the single best (greedy). It suits translation and summarisation, where there is a 'right' answer; for open-ended generation it produces bland, repetitive text.

Cross-attention (decoder queries attending to encoder keys and values) is how the decoder reads the source; it's also how many multimodal models inject image features into a language model.

Where this comes back

  • Week 24You'll build a decoder-only transformer: GPT.

Bahdanau attention (2014) fixed the seq2seq bottleneck: at each decoder step, score every encoder state for relevance, turn the scores into weights with softmax, and take the weighted average as a context vector. The decoder now gets a fresh, focused summary at every step.

Luong attention simplified the scoring to a dot product. The weights are interpretable: in translation they trace word alignments between languages.

The general pattern, score → softmax → weighted sum, is all of attention. The transformer's leap was to use it everywhere and drop recurrence entirely.

αti=softmaxi(score(st,hi)),ct=∑iαtihi\alpha_{ti} = \text{softmax}_i\big(\text{score}(s_t, h_i)\big),\qquad c_t = \sum_i \alpha_{ti} h_i

Going deeper

Additive (Bahdanau) attention scores with a small MLP; multiplicative (Luong, and the transformer's scaled dot product) uses dot products, which are faster on hardware and became the standard.

Attention weights are tempting as explanations but are not faithful explanations of model reasoning on their own.

In self-attention a sequence attends to itself. Each token's embedding is projected three ways: a query (what am I looking for?), a key (what do I contain?) and a value (what do I pass on if selected?).

Token i's output is a weighted average of all values, weighted by softmax of the dot products between its query and every key. The scores are divided by dk\sqrt{d_k} to keep them from growing with dimension and saturating the softmax. In matrix form, all positions are computed at once: the parallelism RNNs lacked.

For generation, a causal mask sets scores for future positions to −∞ so each token attends only to itself and earlier tokens. Pick a word in the simulation and see where it attends.

A library search: your query is matched against every book's catalogue entry (keys), and you read a blend of the books (values) in proportion to how well they matched.

Attention(Q,K,V)=softmax ⁣(QK⊤dk+M)V\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V

Going deeper

Write it once with einsum to fix the shapes in your head: scores = einsum('bhqd,bhkd->bhqk', Q, K) / √d; weights = softmax(scores + mask); out = einsum('bhqk,bhkd->bhqd', weights, V).

Interpretability research found 'induction heads': attention patterns that copy what followed an earlier occurrence of the current token. They're a key mechanism behind in-context learning.

Where this comes back

  • Week 4QKᵀ is a matrix of dot products: similarity again.
  • Week 26Keys and values are what the KV cache stores during generation.

One attention pattern must average over everything a token cares about. Multi-head attention splits the model dimension into h heads (say 12 heads of 64 dimensions for a 768-wide model). Each head has its own Q, K and V projections and attends independently.

Heads specialise: some track the previous token, some syntax (verb → subject), some coreference (pronoun → noun). Outputs are concatenated and mixed by a final linear projection.

Cost is the same as one full-width head; you get several relationship types for free.

MultiHead(X)=Concat(head1,…,headh)WO,headi=Attention(XWiQ,XWiK,XWiV)\text{MultiHead}(X) = \text{Concat}(\text{head}_1,\dots,\text{head}_h)W^O,\quad \text{head}_i = \text{Attention}(XW_i^Q, XW_i^K, XW_i^V)

Going deeper

Grouped-query attention (GQA) shares each key/value head across several query heads, shrinking the KV cache several-fold with little quality loss. Multi-query attention (MQA) is the extreme of one shared KV head. Most modern LLMs use GQA.

Not all heads matter equally: many can be pruned after training with small loss, evidence of redundancy.

Self-attention treats its input as a set: shuffle the tokens and each token's output is unchanged apart from the reordering. Word order must be added.

The original transformer added fixed sinusoidal encodings of different frequencies to the embeddings. GPT-2 used learned position embeddings. Modern LLMs mostly use RoPE (rotary embeddings), which rotate queries and keys by an angle proportional to position, so their dot product depends on relative distance.

Positional scheme determines how well a model handles contexts longer than it was trained on, which is why context-extension techniques often adjust RoPE.

PE(pos,2i)=sin⁡ ⁣(pos/100002i/d),PE(pos,2i+1)=cos⁡ ⁣(pos/100002i/d)PE_{(pos,2i)} = \sin\!\big(pos/10000^{2i/d}\big),\quad PE_{(pos,2i+1)} = \cos\!\big(pos/10000^{2i/d}\big)

Going deeper

RoPE rotates pairs of query/key dimensions by angles proportional to position, so the attention dot product depends only on relative distance. Context-extension methods (position interpolation, YaRN) rescale those angles to reach longer contexts than training.

ALiBi skips position embeddings and adds a distance-proportional penalty to attention scores, with good length extrapolation.

A (pre-norm, decoder-only) block: x = x + MultiHeadAttention(LayerNorm(x)), then x = x + MLP(LayerNorm(x)). The MLP is two linear layers with a GELU between them, expanding to about 4× the model width and back.

Roles: attention moves information between positions (communication); the MLP processes each position independently (computation, and where much factual knowledge is stored); residual connections keep a clean gradient path and let each block make small edits to a shared 'residual stream'; layer norm keeps scales stable.

A full model is: token embeddings + positions → N blocks → final LayerNorm → a linear layer to vocabulary logits → softmax. Draw it from memory until you can do it without hesitating.

python
class Block(nn.Module):
    def __init__(self, d, n_heads):
        super().__init__()
        self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
        self.attn = CausalSelfAttention(d, n_heads)
        self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
    def forward(self, x):
        x = x + self.attn(self.ln1(x))   # communicate
        x = x + self.mlp(self.ln2(x))    # compute
        return x

Going deeper

The residual stream view: each block reads from and writes back to a shared vector per position. Attention heads move information between positions; MLPs transform it in place, and much factual association appears to live in MLP weights.

Parameter split in a standard block: attention ≈ 4d², MLP ≈ 8d² (with a 4× expansion), so MLPs hold about two-thirds of a transformer's parameters.

Best resources for this lesson

Where this comes back

  • Week 24You'll implement and train exactly this.
  • Week 26Attention cost grows with context length: the root of LLM cost mechanics.

The score matrix QK⊤QK^\top has n × n entries for a sequence of n tokens. Doubling the context quadruples attention compute and, naively, memory. At 100k tokens that matrix has ten billion entries per head per layer.

FlashAttention computes exact attention without ever materialising the full matrix: it tiles Q, K and V into blocks that fit in fast on-chip memory and uses an online softmax. Same result, far less memory traffic, often 2–4× faster. It's now the default in PyTorch (scaled_dot_product_attention).

Other approaches change the computation: sliding-window attention (each token sees a local window), sparse attention, linear attention and state-space models trade some modelling power for linear scaling. During generation the KV cache (week 26) makes each new token cost O(n) rather than recomputing everything.

FLOPsattn≈4 n2d per layer,FLOPsMLP≈16 nd2\text{FLOPs}_{\text{attn}} \approx 4\,n^2 d \ \text{per layer}, \qquad \text{FLOPs}_{\text{MLP}} \approx 16\,n d^2

Common pitfalls

  • Assuming a model handles long context as well as short just because the window is large (see context rot, week 28).

Best resources for this lesson

Where this comes back

  • Week 26Context length drives KV-cache memory and cost.
  • Week 28Long windows don't guarantee good use of them.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Draw it from memory

    Draw a decoder-only transformer block (LayerNorm, attention, residual, MLP) with tensor shapes on every arrow. Check against the Annotated Transformer.

  2. Core

    Attention in NumPy

    Implement scaled dot-product attention, a causal mask and multi-head splitting in NumPy, and verify against torch.nn.functional.scaled_dot_product_attention.

  3. Stretch

    Paper with the annotated code

    Read Attention Is All You Need alongside The Annotated Transformer and write a one-page summary of what each component is for.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 127Monday25 Jan2 h planned

    Encoder-decoder architecture

  2. Day 128Tuesday26 Jan2 h planned

    Attention mechanism (Bahdanau, Luong)

  3. Day 129Wednesday27 Jan2 h planned

    Read The Illustrated Transformer - twice

  4. Day 130Thursday28 Jan2 h planned

    Self-attention, multi-head attention, Q/K/V

  5. Day 131Friday29 Jan2 h planned

    Positional encodings, residuals, layer norm; read Attention Is All You Need

  6. Day 132Saturday30 Jan3 h planned

    The Annotated Transformer: line-by-line code walkthrough

  7. Day 133Sunday31 JanReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Deep LearningPrimary

CampusX · playlist

Stanford CS224N: NLP with Deep Learning

Stanford Online · playlist

Lectures 1–2 cover word vectors in depth.

Neural Networks

3Blue1Brown · playlist

Sequence Models (DLS Course 5)

Andrew Ng · DeepLearning.AI · playlist

Seq2seq Encoder-Decoder Networks

StatQuest

Attention for Neural Networks

StatQuest

Transformer Neural Networks, Clearly Explained

StatQuest

Decoder-Only Transformers

StatQuest

Transformers, the tech behind LLMs

3Blue1Brown

Attention in transformers, step-by-step

3Blue1Brown

How might LLMs store facts

3Blue1Brown

Attention is all you need: model explanation, inference and training

Umar Jamil

Stanford CS25: Transformers United

Stanford Online · playlist

Guest lectures from the people building these models.

Paper of the week

Attention Is All You Need

The architecture every model in this path is built on. Read it after The Illustrated Transformer, not before.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Explain self-attention with queries, keys and values.
  • Each token projects to Q, K, V
  • Scores = QKᵀ/√d, softmax → weights
  • Output = weighted sum of V; masked for causality in decoders
Why scale by √d_k?
  • Dot products grow with dimension
  • Large scores saturate softmax → tiny gradients
  • Scaling keeps variance roughly constant
What's the computational complexity of attention, and how is long context handled?
  • O(n²·d) compute, O(n²) naive memory
  • FlashAttention: exact, memory-efficient tiling
  • Sliding window, sparse, linear attention; KV cache for generation

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why divide attention scores by √d_k?

  2. 2.What does the causal mask do?

  3. 3.Without positional information, self-attention…

  4. 4.In a transformer block, which part moves information between positions?

  5. 5.GPT is…