AI Engineer Path
Phase 4 · Deep Learning18 Jan – 24 Jan

Week 18

Sequence models: why RNNs lost

Learn RNNs as a problem statement, not as tools.

Why this week matters

You are learning these to understand why they lost: a sequential bottleneck, vanishing gradients over long context, no parallelism. That failure is the entire motivation for attention.

Done when

You can explain, unprompted, why an RNN can't be parallelised across a sequence and a transformer can.

Concepts

6 lessons · tick each one once you could explain it

A recurrent network processes a sequence one element at a time, updating a hidden state: ht=tanh⁡(Whht−1+Wxxt+b)h_t = \tanh(W_h h_{t-1} + W_x x_t + b). The same weights are reused at every step, and the hidden state is the network's memory of everything so far.

Training unrolls the recurrence into a deep feed-forward graph, one layer per time step, and backpropagates through all of it: backpropagation through time (BPTT). Long sequences are often truncated to bound cost.

The structure is inherently sequential: you cannot compute h100h_{100} before h99h_{99}.

ht=tanh⁡(Whht−1+Wxxt+b),y^t=Wyhth_t = \tanh(W_h h_{t-1} + W_x x_t + b),\qquad \hat y_t = W_y h_t

Going deeper

Truncated BPTT backpropagates through only the last k steps while carrying the hidden state forward, bounding memory at the cost of long-range learning.

State-space models (S4, Mamba) revive recurrence with structured, parallelisable updates: linear in sequence length at inference, trainable in parallel. They're the most serious current alternative to attention.

The gradient flowing from step t back to step t−k passes through k copies of the recurrent Jacobian. If its largest eigenvalue is below 1, the product shrinks exponentially (vanishes); above 1, it grows exponentially (explodes).

Consequence: plain RNNs struggle to learn dependencies more than a few dozen steps apart. In 'The keys that the man near the old cabinets left on the table are lost', the verb agreement depends on a word many steps back.

Clipping fixes explosions. Vanishing needs an architectural fix: gates (LSTM/GRU), and ultimately attention, which connects any two positions directly.

∂ht∂ht−k=∏i=1k∂ht−i+1∂ht−i\frac{\partial h_t}{\partial h_{t-k}} = \prod_{i=1}^{k}\frac{\partial h_{t-i+1}}{\partial h_{t-i}}

Going deeper

Exploding gradients are easy to see (NaN losses) and fix (clipping). Vanishing gradients are silent: the model simply never learns long-range dependencies, which is far more dangerous.

Orthogonal initialisation of recurrent weights keeps eigenvalues near magnitude 1 at the start, delaying the problem but not removing it.

Best resources for this lesson

Where this comes back

  • Week 4Repeated multiplication is governed by eigenvalues.

An LSTM adds a cell state ctc_t that runs through time with only element-wise, gated updates, a 'conveyor belt' along which gradients flow more easily. Three sigmoid gates, each between 0 and 1, control it.

The forget gate decides what to erase from the cell; the input gate decides what new candidate information to write; the output gate decides what part of the cell to expose as the hidden state.

Because the cell update is additive (ct=ft⊙ct−1+it⊙c~tc_t = f_t \odot c_{t-1} + i_t \odot \tilde c_t), gradients don't have to pass through a squashing non-linearity at every step: the same insight as ResNet's skip connections.

ct=ft⊙ct−1+it⊙c~t,ht=ot⊙tanh⁡(ct)c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t,\qquad h_t = o_t \odot \tanh(c_t)

Going deeper

The forget-gate bias is often initialised to 1 so the cell remembers by default early in training, a small trick with large effect.

Peephole connections and coupled forget/input gates are variants; empirical studies found the standard LSTM hard to beat.

Best resources for this lesson

The GRU merges the cell and hidden state and uses an update gate (how much of the old state to keep vs replace) and a reset gate (how much of the past to use when proposing the new state).

Fewer parameters, faster training, and often comparable results to LSTM. Neither consistently dominates; the choice was usually empirical.

Both are now mostly historical for language, but still appear in small on-device and streaming models.

ht=(1−zt)⊙ht−1+zt⊙h~th_t = (1 - z_t)\odot h_{t-1} + z_t \odot \tilde{h}_t

Going deeper

With fewer parameters, GRUs train faster and can generalise better on small datasets; LSTMs sometimes win on tasks needing precise counting or long memory.

Best resources for this lesson

A bidirectional RNN runs one RNN left-to-right and another right-to-left and concatenates their states, so each position sees both past and future context. Great for tagging and classification; impossible for generation, where the future doesn't exist yet.

Stacked RNNs feed one layer's hidden states as the next layer's input sequence, building more abstract representations.

Encoder side of seq2seq (next week) was typically a bidirectional, stacked LSTM. BERT's 'bidirectional' name echoes this.

Going deeper

ELMo (2018) used deep bidirectional LSTM language models to produce contextual word embeddings, the direct predecessor of BERT's idea of pretraining and then fine-tuning.

Best resources for this lesson

No parallelism across time. Step t needs step t−1, so a 1,000-token sequence needs 1,000 sequential steps. GPUs, which excel at huge parallel matrix multiplies, sit mostly idle. Transformers process all positions at once during training.

A fixed-size bottleneck. Everything the model knows about the past must squeeze into one hidden vector. In seq2seq translation, the whole source sentence was compressed into a single vector before decoding began.

Long-range dependencies are hard. Even with gates, information from far back degrades. Attention lets any position look directly at any other in one step, with a path length of 1. Hold these three failures in mind: next week each one gets solved.

Going deeper

Hardware decided it as much as algorithms: attention turns sequence processing into large matrix multiplications, exactly what GPUs and TPUs do best, so transformers scale smoothly with compute.

The trade-off flipped at inference: RNNs need constant memory per step, while transformers' KV cache grows with context. That's why recurrent and hybrid architectures are being revisited for long contexts.

Best resources for this lesson

Where this comes back

  • Week 19Attention removes all three limitations at once.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Char-RNN

    Train a character-level LSTM on a small text corpus and sample from it at different temperatures.

  2. Stretch

    Watch gradients vanish

    Plot the gradient norm reaching each time step in a vanilla RNN vs an LSTM on a long sequence.

  3. Warm-up

    Explain it in writing

    Write one paragraph explaining why RNNs can't parallelise across time and transformers can. This is the week's done-when test.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 120Monday18 Jan2 h planned

    RNN architecture and backprop through time

  2. Day 121Tuesday19 Jan2 h planned

    Vanishing gradients in RNNs

  3. Day 122Wednesday20 Jan2 h planned

    LSTM: the gate mechanism

  4. Day 123Thursday21 Jan2 h planned

    GRU and how it compares to LSTM

  5. Day 124Friday22 Jan2 h planned

    Bidirectional and stacked RNNs

  6. Day 125Saturday23 Jan3 h planned

    Why RNNs lost: sequential bottleneck, no parallelism

  7. Day 126Sunday24 JanReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Deep LearningPrimary

CampusX · playlist

Neural Networks / Deep Learning

StatQuest · playlist

Sequence Models (DLS Course 5)

Andrew Ng · DeepLearning.AI · playlist

Recurrent Neural Networks, Clearly Explained

StatQuest

LSTM, Clearly Explained

StatQuest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why did transformers replace RNNs?
  • Parallel computation across positions during training
  • Direct paths between any two tokens (no fixed-size bottleneck)
  • Better long-range dependencies; scales with hardware
How do LSTMs mitigate vanishing gradients?
  • Additive cell-state updates controlled by gates
  • Forget gate lets information persist
  • Gradient flows without repeated squashing

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why can't an RNN be parallelised across a sequence during training?

  2. 2.The LSTM's cell-state update is mostly…

  3. 3.A bidirectional RNN is unsuitable for…

  4. 4.In classic seq2seq, the source sentence passed to the decoder as…

  5. 5.Gradients across many RNN steps vanish when…