Sequence-to-sequence models map an input sequence to an output sequence of different length: translation, summarisation. The encoder reads the input; the decoder generates the output one token at a time, each step conditioned on the tokens generated so far.
During training, the decoder is fed the correct previous tokens (teacher forcing); at inference it consumes its own outputs, decoded greedily or with beam search.
Transformers keep this split. The original transformer was encoder–decoder (like T5 today). BERT is encoder-only (understanding), GPT is decoder-only (generation).
Going deeper
Beam search keeps the k best partial sequences at each step instead of only the single best (greedy). It suits translation and summarisation, where there is a 'right' answer; for open-ended generation it produces bland, repetitive text.
Cross-attention (decoder queries attending to encoder keys and values) is how the decoder reads the source; it's also how many multimodal models inject image features into a language model.
Best resources for this lesson
Where this comes back
- Week 24You'll build a decoder-only transformer: GPT.