A recurrent network processes a sequence one element at a time, updating a hidden state: . The same weights are reused at every step, and the hidden state is the network's memory of everything so far.
Training unrolls the recurrence into a deep feed-forward graph, one layer per time step, and backpropagates through all of it: backpropagation through time (BPTT). Long sequences are often truncated to bound cost.
The structure is inherently sequential: you cannot compute before .
Going deeper
Truncated BPTT backpropagates through only the last k steps while carrying the hidden state forward, bounding memory at the cost of long-range learning.
State-space models (S4, Mamba) revive recurrence with structured, parallelisable updates: linear in sequence length at inference, trainable in parallel. They're the most serious current alternative to attention.