AI Engineer Path
Phase 6 · Build a GPT8 Mar – 14 Mar

Week 25

Tokenizer, nanoGPT and multimodal models

Open the tokenizer black box, read production GPT code, and see how vision enters a language model.

  • Off-roadmap insertion
  • Third to cut if behind

Why this week matters

BPE by hand is where tokenisation stops being a black box and token costs start making sense. Many 'LLM weirdness' bugs (bad arithmetic, trouble spelling words, odd behaviour with whitespace) are tokenizer bugs.

Done when

All from-scratch code is on GitHub with a README.

Milestone: All from-scratch code on GitHub

Concepts

6 lessons · tick each one once you could explain it

BPE training: start with the text as a sequence of bytes (256 base tokens, so any string is representable). Count every adjacent pair, merge the most frequent pair into a new token, replace it everywhere, and repeat until the vocabulary reaches the target size (GPT-2: about 50k; modern models: 100k–200k+).

Encoding new text applies the learned merges in order. Decoding concatenates the bytes of each token. Frequent words become single tokens; rare words split into pieces; nothing is ever out-of-vocabulary. Step through merges in the simulation.

Production tokenizers first split text with a regex (so merges don't cross word, number or punctuation boundaries) and add special tokens like <|endoftext|>.

python
def get_stats(ids):
    counts = {}
    for pair in zip(ids, ids[1:]):
        counts[pair] = counts.get(pair, 0) + 1
    return counts

def merge(ids, pair, new_id):
    out, i = [], 0
    while i < len(ids):
        if i < len(ids) - 1 and (ids[i], ids[i + 1]) == pair:
            out.append(new_id); i += 2
        else:
            out.append(ids[i]); i += 1
    return out

ids = list("low lower lowest".encode("utf-8"))
for k in range(10):
    pair = max(get_stats(ids).items(), key=lambda kv: kv[1])[0]
    ids = merge(ids, pair, 256 + k)

Going deeper

Byte-level BPE (GPT-2 onward) starts from 256 bytes, so any UTF-8 text is encodable; a pre-tokenization regex stops merges from crossing word, number and punctuation boundaries.

SentencePiece treats the input as a raw stream (spaces become '▁'), making it language-agnostic. Its Unigram algorithm prunes a large vocabulary down instead of growing a small one.

Where this comes back

  • Week 26Token counts set cost and context limits.

Spelling and character counting are hard because 'strawberry' may be one or two tokens; the model never directly sees its letters. Arithmetic is awkward because numbers split inconsistently ('1234' vs '12' + '34').

Non-English languages and code often take 2–4× more tokens for the same content, so they cost more and fill context faster. Leading spaces matter: ' hello' and 'hello' are different tokens.

Trailing whitespace in prompts, unusual unicode and rare 'glitch tokens' can produce odd behaviour. When a model misbehaves on a specific string, inspect its tokens first.

Going deeper

Newer tokenizers split numbers into single digits or fixed-size chunks to make arithmetic more consistent, one reason maths ability jumped between model generations.

Common pitfalls

  • Estimating cost with words instead of tokens.
  • Assuming the same text costs the same across different model families' tokenizers.

Best resources for this lesson

nanoGPT is a compact, readable codebase that reproduces GPT-2. Compared with your from-scratch model: fused, efficient attention (scaled_dot_product_attention, i.e. FlashAttention), weight tying between the embedding and output layers, careful initialisation, AdamW with weight decay applied only to 2D weights, cosine LR schedule with warmup, gradient accumulation, mixed precision, torch.compile, and distributed data parallel training.

Reading it end to end is the bridge from 'I understand GPT' to 'I can read production training code'.

Optional: Karpathy's 'Let's reproduce GPT-2 (124M)' walks through the same engineering live.

Going deeper

Reproducing GPT-2 124M now takes hours on rented GPUs rather than weeks, thanks to FlashAttention, bf16, torch.compile and better hyperparameters. The 'speedrun' community has pushed it further still.

Best resources for this lesson

Pretraining: next-token prediction over a huge filtered web, code and book corpus. The result is a base model that continues documents but doesn't follow instructions.

Post-training: supervised fine-tuning (SFT) on curated conversations teaches the assistant format; preference optimisation (RLHF or DPO) on human or AI preference judgements shapes helpfulness and harmlessness; reinforcement learning on verifiable tasks (maths, code) builds reasoning.

Karpathy's 'Deep Dive into LLMs' covers this whole pipeline and is the bridge into week 26.

Going deeper

Data quality dominates: deduplication, quality filtering and careful mixing of web, code, books and maths shape capability as much as architecture. Post-training data is small but extremely curated.

Reinforcement learning with verifiable rewards (GRPO and relatives) on maths and code is how reasoning models learn long chains of thought; DeepSeek-R1's report is the most readable public account.

Where this comes back

  • Week 26Pretraining vs SFT vs RLHF/DPO is the first topic of LLM 101.

The common recipe: a pretrained vision encoder (a ViT, often CLIP-style) turns an image into patch embeddings; a small projector maps them into the LLM's embedding space; the LLM then processes these 'visual tokens' alongside text tokens with ordinary attention.

Training typically aligns the projector first (image–caption pairs), then fine-tunes on visual instruction data. Some models are natively multimodal from pretraining.

Practical consequences: images cost tokens (often hundreds per image, scaling with resolution), small text in images may be missed at low resolution, and document understanding (Project 3) relies on exactly this.

Going deeper

Two main designs: 'unified embedding' (image tokens concatenated with text tokens, as in LLaVA) and 'cross-attention' (the LLM attends to image features through extra layers). Raschka's survey compares them clearly.

Where this comes back

  • Week 14CLIP and ViT are the vision half of this recipe.

Kaplan et al. (2020) found that language-model loss decreases as a smooth power law with model size, dataset size and compute, over many orders of magnitude. This made it possible to predict a big model's performance from small runs.

The Chinchilla paper (2022) showed earlier models were undertrained: for a fixed compute budget, parameters and training tokens should grow roughly equally, around 20 tokens per parameter. Training compute is roughly 6ND6ND FLOPs for N parameters and D tokens.

In practice, models are now trained far past Chinchilla-optimal, because a smaller model trained longer is cheaper to serve. Inference cost, not just training cost, drives the choice.

C≈6 ND,L(N,D)≈E+ANα+BDβC \approx 6\,N D,\qquad L(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}

Where this comes back

  • Week 26Model size and training choices shape the models you'll pick from.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Tokenizer forensics

    In Tiktokenizer, compare token counts for the same paragraph in English, Hindi and Python. Estimate the cost difference per million words.

  2. Core

    Implement BPE

    Implement BPE training, encode and decode (minbpe-style) and verify round-tripping on random unicode text.

  3. Stretch

    Publish the from-scratch repo

    Push micrograd, makemore, GPT and tokenizer code to GitHub with a README that explains what each part taught you.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 169Monday8 Mar2 h planned

    Karpathy 8: build the GPT tokenizer (minbpe) part 1

  2. Day 170Tuesday9 Mar2 h planned

    Karpathy 8: tokenizer part 2 - BPE implementation

  3. Day 171Wednesday10 Mar2 h planned

    nanoGPT walkthrough

  4. Day 172Thursday11 Mar2 h planned

    Karpathy: Deep Dive into LLMs (overview lecture)

  5. Day 173Friday12 Mar2 h planned

    Large Multimodal Models: how vision gets into an LLM

  6. Day 174Saturday13 Mar3 h planned

    PHASE REVIEW: push all code to GitHub with a proper README

  7. Day 175Sunday14 MarReview

    Review the week, finish anything unfinished, rest

Watch

Neural Networks: Zero to HeroPrimary

Andrej Karpathy · playlist

Type every line. Watching alone is worthless; typing it is transformative.

Let's build the GPT TokenizerPrimary

Andrej Karpathy

Let's reproduce GPT-2 (124M)

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

Andrej Karpathy

Build a Large Language Model (From Scratch)

Sebastian Raschka · playlist

Stanford CS336: Language Modeling from Scratch (2025)

Stanford Online · playlist

Graduate-level depth. Optional.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why do LLMs struggle with spelling and arithmetic?
  • They operate on tokens, not characters or digits
  • Numbers tokenize inconsistently
  • Mitigations: tools (calculators, code), digit-level tokenization
What are scaling laws and why do they matter?
  • Loss falls as power law in N, D, compute
  • Chinchilla: ~20 tokens/parameter compute-optimal
  • Production models over-train small models to cut inference cost

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.BPE training repeatedly…

  2. 2.Why can byte-level BPE encode any string?

  3. 3.Why do LLMs struggle to count the r's in 'strawberry'?

  4. 4.A base model (pretrained only) typically…

  5. 5.How does a multimodal LLM 'see' an image?