Training data is a long stream of tokens. Sample random chunks of length block_size; the targets are the same chunk shifted by one. A single chunk of 256 tokens therefore holds 256 training examples: predict token 2 from token 1, token 3 from tokens 1–2, and so on.
The causal mask makes this possible in one forward pass: each position may only attend to earlier ones, so no position can cheat by looking at its own target.
Loss is cross-entropy over the vocabulary at every position, averaged. This is the pretraining objective of every GPT-style LLM.
def get_batch(data, block_size, batch_size):
ix = torch.randint(len(data) - block_size, (batch_size,))
x = torch.stack([data[i:i + block_size] for i in ix])
y = torch.stack([data[i + 1:i + block_size + 1] for i in ix]) # shifted by one
return x, yGoing deeper
Teacher forcing during training (always feeding the true previous tokens) creates exposure bias: at generation time the model conditions on its own possibly wrong outputs, so errors compound.
Packing many short documents into fixed-length sequences, separated by end-of-text tokens, keeps GPUs fully utilised during pretraining.