Concept Lab · Week 25 · Build a GPT
Byte-pair encoding
Merge the most frequent pair, again and again, and watch tokens form.
The idea: Byte-pair encoding
BPE training: start with the text as a sequence of bytes (256 base tokens, so any string is representable). Count every adjacent pair, merge the most frequent pair into a new token, replace it everywhere, and repeat until the vocabulary reaches the target size (GPT-2: about 50k; modern models: 100k–200k+).
Encoding new text applies the learned merges in order. Decoding concatenates the bytes of each token. Frequent words become single tokens; rare words split into pieces; nothing is ever out-of-vocabulary. Step through merges in the simulation.
Production tokenizers first split text with a regex (so merges don't cross word, number or punctuation boundaries) and add special tokens like <|endoftext|>.
Next simulation: Temperature and top-p