AI Engineer Path

Week 12

NLP (1/2): representing text

Turn text into numbers, from counting words to learned embeddings.

Why this week matters

This branch is the direct ancestor of the whole AI engineer track. TF-IDF and Word2Vec are the conceptual roots of the vector databases you'll build ten weeks from now. Learn them properly and embeddings will feel inevitable rather than magical.

Done when

You can explain why TF-IDF down-weights 'the', and why Word2Vec puts 'king' near 'queen'.

Concepts

6 lessons · tick each one once you could explain it

Tokenisation splits text into units: words, punctuation, sub-words or characters. Splitting on spaces fails quickly: 'don't', 'U.S.', emoji, URLs, and languages without spaces. Libraries like spaCy handle these rules.

Classic cleaning steps: lowercasing, removing punctuation and URLs, normalising unicode, expanding contractions. Each throws information away. Lowercasing merges 'Apple' the company with 'apple' the fruit.

For classical models (bag-of-words, TF-IDF), heavier cleaning helps. For transformers, don't clean: they come with their own sub-word tokenizer and need the raw text, casing and punctuation included.

python
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Dr. Smith didn't visit the U.S. in 2023!")
print([t.text for t in doc])
# ['Dr.', 'Smith', 'did', "n't", 'visit', 'the', 'U.S.', 'in', '2023', '!']

Going deeper

Unicode normalisation (NFC/NFKC) makes visually identical strings byte-identical: 'café' can be one code point or two. Skip it and deduplication, matching and token counts quietly go wrong.

Language identification, boilerplate removal (navigation, cookie banners) and deduplication are the unglamorous steps that decide data quality for search indexes and LLM training corpora alike.

Best resources for this lesson

Where this comes back

  • Week 25LLMs use learned sub-word tokenizers (BPE) instead of rules.

Stemming chops suffixes by rule ('running' → 'run', 'studies' → 'studi'): fast and crude. Lemmatisation uses vocabulary and part of speech to return the dictionary form ('better' → 'good', 'was' → 'be'): slower but correct.

Stopwords ('the', 'is', 'and') are frequent and usually uninformative for topic-style tasks. But removing them can destroy meaning: 'not good' becomes 'good'. For sentiment, keep negations.

spaCy's pipeline gives you tokens, lemmas, part-of-speech tags and entities in one pass.

Going deeper

Lemmatisation depends on part of speech: 'meeting' is a noun in 'the meeting ran long' and a verb in 'we are meeting'. That's why spaCy tags POS before lemmatising.

For search, stemming increases recall (more matches) but hurts precision; modern search often skips it and lets embeddings handle morphology.

Common pitfalls

  • Removing 'not', 'no' and 'never' before sentiment analysis.

Best resources for this lesson

Bag-of-words builds a vocabulary and represents each document as a vector of word counts. Word order is thrown away: 'dog bites man' and 'man bites dog' look identical.

n-grams recover some order by also counting sequences of n words: bigrams like 'not good' or 'New York'. The vocabulary grows fast, so cap it with min_df, max_features.

The result is a sparse vector: tens of thousands of dimensions, almost all zero. That sparsity is the key contrast with the dense embeddings coming at the end of the week.

python
from sklearn.feature_extraction.text import CountVectorizer
vec = CountVectorizer(ngram_range=(1, 2), min_df=2)
X = vec.fit_transform(["the food was not good", "the food was good"])

Going deeper

Character n-grams (e.g. 3–5 characters) are robust to typos and morphology and work well for short, noisy text and language identification.

Sparse matrices (CSR format) store only non-zeros, so a 100,000-word vocabulary costs almost nothing per document. Keep them sparse all the way into the model.

Best resources for this lesson

Term frequency (TF) says a word matters to a document if it appears often there. Inverse document frequency (IDF) says a word is informative if it appears in few documents. Multiply them: 'the' scores near zero everywhere, while 'transformer' scores high in the few documents about it.

Rows are usually L2-normalised, so cosine similarity between TF-IDF vectors measures topical overlap. That is a complete, working search engine in a few lines.

BM25, the ranking function behind classic search engines, is a refined TF-IDF with saturation (the 10th occurrence adds less than the 2nd) and document-length normalisation. You will use it for hybrid search in RAG.

tf-idf(t,d)=tf(t,d)⋅log⁡N1+df(t)\text{tf-idf}(t, d) = \text{tf}(t, d)\cdot \log\frac{N}{1 + \text{df}(t)}

Going deeper

sublinear_tf=True uses 1 + log(tf), so the tenth occurrence of a word counts far less than the first, the same intuition as BM25's saturation.

BM25's parameters: k1 (how fast term frequency saturates, ~1.2–2) and b (how much to normalise for document length, ~0.75). It remains a strong first-stage retriever and the keyword half of hybrid search.

Best resources for this lesson

Where this comes back

  • Week 13TF-IDF + logistic regression is Project 2, your permanent baseline.
  • Week 31BM25 (a TF-IDF descendant) powers keyword search in hybrid RAG.

The distributional hypothesis: 'you shall know a word by the company it keeps'. Word2Vec turns that into a training task. Skip-gram predicts surrounding words from a centre word; CBOW predicts the centre word from its context. The trained weights become dense embeddings of ~100–300 dimensions.

Words used in similar contexts end up close together, and directions encode relationships: vector('king') − vector('man') + vector('woman') ≈ vector('queen'). Training uses negative sampling to stay cheap: tell real context words apart from random ones.

Limitation: one vector per word, regardless of context. 'Bank' gets a single vector whether it means river or money. Transformers fix this with contextual embeddings.

max⁡∑t∑−c≤j≤c, j≠0log⁡P(wt+j∣wt),P(o∣c)=exp⁡(uo⊤vc)∑wexp⁡(uw⊤vc)\max \sum_{t}\sum_{-c\le j\le c,\ j\neq 0} \log P(w_{t+j}\mid w_t),\quad P(o\mid c) = \frac{\exp(u_o^\top v_c)}{\sum_{w}\exp(u_w^\top v_c)}

Going deeper

Skip-gram with negative sampling is implicitly factorising a shifted PMI (pointwise mutual information) matrix, linking Word2Vec to count-based methods like GloVe.

Embeddings absorb biases in their training text (gender and occupation associations, for example). Audit them before using them in decisions about people.

Where this comes back

  • Week 19Attention produces contextual embeddings that solve the one-vector-per-word problem.

GloVe learns word vectors from global co-occurrence counts, factorising a word–word matrix so dot products approximate log co-occurrence. Different route, similar result to Word2Vec.

The intuition to keep: an embedding maps something discrete (a word, sentence, image, user) to a point in a continuous space where geometry encodes meaning. Similar things are close; cosine similarity measures closeness.

Today's sentence-embedding models (week 30) do this for whole passages. Retrieval in RAG is then: embed the question, find the nearest document embeddings. You are looking at the root of the whole AI engineer track.

Going deeper

Static embeddings give each word one vector; contextual embeddings (BERT and beyond) give a different vector per occurrence. Sentence embeddings pool or train for a single vector per passage, which is what retrieval needs.

Embedding spaces from different models aren't compatible: switching embedding models means re-embedding the whole corpus, which is a real migration cost to plan for.

Best resources for this lesson

Where this comes back

  • Week 30Sentence embeddings + nearest-neighbour search = vector databases.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    TF-IDF search engine

    Index 1,000 documents with TF-IDF, answer queries by cosine similarity, and compare with BM25 (rank_bm25) on 10 queries you label yourself.

  2. Core

    Word analogies

    Load pretrained GloVe vectors and test 20 analogies (king − man + woman). Find three that fail and explain why.

  3. Stretch

    Train Word2Vec

    Train Word2Vec with Gensim on a domain corpus and compare nearest neighbours for 10 domain terms against general-purpose GloVe.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 78Monday7 Dec2 h planned

    Text preprocessing: tokenisation, casing, regex cleaning

  2. Day 79Tuesday8 Dec2 h planned

    Stemming, lemmatisation, stopwords, spaCy pipeline

  3. Day 80Wednesday9 Dec2 h planned

    Bag-of-words and n-grams

  4. Day 81Thursday10 Dec2 h planned

    TF-IDF in depth

  5. Day 82Friday11 Dec2 h planned

    Word2Vec: CBOW and skip-gram

  6. Day 83Saturday12 Dec3 h planned

    GloVe + embedding intuition (this is the ancestor of vector search)

  7. Day 84Sunday13 DecReview

    Review the week, finish anything unfinished, rest

Watch

Word Embedding and Word2Vec

StatQuest

Stanford CS224N: NLP with Deep Learning

Stanford Online · playlist

Lectures 1–2 cover word vectors in depth.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Explain TF-IDF and why it beats raw counts.
  • TF rewards frequency in the document
  • IDF discounts words common across documents
  • Highlights distinctive terms; BM25 refines it
How does Word2Vec learn word meaning?
  • Predict context words (skip-gram) or centre word (CBOW)
  • Words in similar contexts get similar vectors
  • Negative sampling makes training cheap

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why does TF-IDF give 'the' a near-zero weight?

  2. 2.Bag-of-words treats 'dog bites man' and 'man bites dog' as…

  3. 3.Main limitation of Word2Vec embeddings?

  4. 4.Should you lowercase and strip punctuation before a transformer model?

  5. 5.Sparse vs dense representations: which is dense?