Tokenisation splits text into units: words, punctuation, sub-words or characters. Splitting on spaces fails quickly: 'don't', 'U.S.', emoji, URLs, and languages without spaces. Libraries like spaCy handle these rules.
Classic cleaning steps: lowercasing, removing punctuation and URLs, normalising unicode, expanding contractions. Each throws information away. Lowercasing merges 'Apple' the company with 'apple' the fruit.
For classical models (bag-of-words, TF-IDF), heavier cleaning helps. For transformers, don't clean: they come with their own sub-word tokenizer and need the raw text, casing and punctuation included.
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Dr. Smith didn't visit the U.S. in 2023!")
print([t.text for t in doc])
# ['Dr.', 'Smith', 'did', "n't", 'visit', 'the', 'U.S.', 'in', '2023', '!']Going deeper
Unicode normalisation (NFC/NFKC) makes visually identical strings byte-identical: 'café' can be one code point or two. Skip it and deduplication, matching and token counts quietly go wrong.
Language identification, boilerplate removal (navigation, cookie banners) and deduplication are the unglamorous steps that decide data quality for search indexes and LLM training corpora alike.
Best resources for this lesson
- BookSpeech and Language Processing, ch. 2: words and tokens · Jurafsky & Martin
- DocsspaCy 101
Where this comes back
- Week 25LLMs use learned sub-word tokenizers (BPE) instead of rules.