AI Engineer Path

Week 13

NLP (2/2): classic models + Project 2

Build classic NLP pipelines and a baseline you'll compare everything against.

Why this week matters

'We tried the simple thing first' is one of the strongest signals you can give an interviewer. Project 2 is the yardstick every later LLM approach must beat.

Done when

Project 2 exists: a TF-IDF + logistic regression classifier with recorded metrics. This is your permanent baseline.

Milestone: Project 2: TF-IDF text classifier (LLM baseline)

Concepts

6 lessons · tick each one once you could explain it

Multinomial Naive Bayes models each class as a distribution over words. To classify, it sums log-probabilities of the document's words under each class plus the class prior, and picks the largest.

Laplace smoothing (add-one) prevents a single unseen word from zeroing out a class probability. Work in log space to avoid numeric underflow.

It trains in one pass, works with little data, and is a classic spam filter. Logistic regression on TF-IDF usually edges it out on accuracy.

c^=arg⁡max⁡c[log⁡P(c)+∑w∈dlog⁡P(w∣c)]\hat c = \arg\max_c \Big[\log P(c) + \sum_{w\in d}\log P(w\mid c)\Big]

Going deeper

Complement Naive Bayes (ComplementNB) is more robust on imbalanced text classes. NB-SVM (Naive Bayes features into a linear SVM) was a famously strong sentiment baseline.

Naive Bayes trains incrementally (partial_fit), useful for streaming classification such as spam filtering.

A complete pipeline: clean text lightly (keep negations), vectorise with TF-IDF using unigrams and bigrams, train logistic regression, evaluate with per-class precision, recall and F1 plus a confusion matrix.

The most valuable step is error analysis: read 20–50 misclassified examples and group them. Typical buckets: sarcasm, negation scope ('not bad at all'), mixed sentiment, domain vocabulary, label noise.

Those buckets tell you whether better features, more data or a different model class (a transformer, in week 21) would help.

python
from sklearn.pipeline import make_pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

clf = make_pipeline(TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True),
                    LogisticRegression(max_iter=2000, C=4))
clf.fit(X_train, y_train)
print(classification_report(y_test, clf.predict(X_test)))

Going deeper

Aspect-based sentiment separates opinions about different things in one review ('great battery, awful screen'), usually what businesses actually need.

Label noise is common in sentiment data: if annotators agree only 80% of the time, no model can reliably exceed that agreement. Measure inter-annotator agreement to set a realistic ceiling.

Best resources for this lesson

Where this comes back

  • Week 36Error analysis by bucketing failures is the core of LLM eval work too.

Latent Dirichlet Allocation assumes each document is a mixture of topics and each topic is a distribution over words. Fitting it finds, say, 20 topics, each summarised by its top words, plus each document's topic proportions.

Choosing the number of topics is part art: use coherence scores and, more importantly, read the topics. Bag-of-words input means preprocessing (stopwords, lemmatisation) matters a lot here.

Modern alternative: embed documents, cluster the embeddings (BERTopic). Same goal, better semantics.

Going deeper

LDA's Dirichlet priors control sparsity: low α means each document has few topics; low β means each topic uses few words. Topic coherence scores help compare settings but don't replace reading the topics.

BERTopic replaces bag-of-words with sentence embeddings, UMAP and HDBSCAN, then labels clusters with class-based TF-IDF. It's usually more coherent on short texts.

Part-of-speech tagging labels tokens as noun, verb, adjective and so on. Named entity recognition finds spans like PERSON, ORG, GPE (places), DATE and MONEY. Both are sequence labelling tasks, usually with BIO tags (Begin, Inside, Outside).

spaCy ships pretrained models for both. NER is evaluated at the span level: a predicted entity counts only if its boundaries and type are exactly right.

Where it shows up later: extracting structured fields from documents. LLMs can now do this zero-shot, but a small NER model is far cheaper per call when the entity types are fixed.

python
doc = nlp("Anthropic raised money in San Francisco on 4 March.")
[(e.text, e.label_) for e in doc.ents]
# [('Anthropic', 'ORG'), ('San Francisco', 'GPE'), ('4 March', 'DATE')]

Going deeper

Classic NER used conditional random fields (CRFs) over hand-crafted features; modern NER fine-tunes transformers for token classification. GLiNER-style models do zero-shot NER for arbitrary entity types cheaply.

Evaluate with entity-level precision/recall/F1 (seqeval). Partial-span matches count as errors, which is often stricter than downstream needs.

Where this comes back

  • Week 28Project 4 does structured extraction with an LLM; compare its cost against a NER model.

A TF-IDF + logistic regression classifier trains in seconds, costs nothing per prediction and is often within a few points of far bigger models on narrow tasks. Without it, you can't say whether an LLM is worth 1,000× the cost.

Save the baseline's metrics to a file in the repo. Every later approach (fine-tuned BERT in week 21, prompted LLMs in week 27, fine-tuned LLMs in week 33) reports its delta against these numbers on the same test set.

Report cost and latency next to accuracy. 'Two points better at 400× the cost' is a real tradeoff worth stating.

Going deeper

Make the baseline a first-class artefact: a script that reproduces it, a frozen test set, and a metrics file in the repo that CI can compare against.

Always report the comparison with cost. Plenty of production teams run a cheap classical model for 90% of traffic and route only uncertain cases to an LLM.

Backend engineer tip: This is the same instinct as benchmarking before optimising.

Best resources for this lesson

Where this comes back

  • Week 21Fine-tuned BERT must beat this baseline, and you report by how much.
  • Week 33Your QLoRA fine-tune is compared against prompting and this baseline.

BLEU (machine translation) measures n-gram precision against reference translations with a brevity penalty. ROUGE (summarisation) measures n-gram and longest-common-subsequence recall against reference summaries. Both reward surface overlap.

BERTScore compares contextual embeddings token by token, so paraphrases get credit. It correlates better with human judgement but still needs references and can't check facts.

For open-ended LLM output, these metrics are weak: a correct answer worded differently scores low, and a fluent wrong answer can score high. That gap is why LLM evaluation moved toward task-specific checks, rubric-based LLM judges and human review.

BLEU=BP⋅exp⁡(∑n=14wnlog⁡pn)\text{BLEU} = \text{BP}\cdot\exp\Big(\sum_{n=1}^{4} w_n \log p_n\Big)

Common pitfalls

  • Optimising ROUGE until summaries become extractive copies.
  • Comparing BLEU scores computed with different tokenisation.

Where this comes back

  • Week 36LLM evals replace overlap metrics with task-specific checks and judges.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Baseline in an afternoon

    Build Project 2's TF-IDF + logistic regression classifier and save metrics.json with per-class precision/recall/F1.

  2. Core

    Error analysis

    Read 50 misclassified examples and bucket them (negation, sarcasm, label noise…). Write which buckets a transformer might fix.

  3. Stretch

    NER comparison

    Compare spaCy's NER with a zero-shot LLM prompt on 30 sentences: entity-level F1, cost and latency.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 85Monday14 Dec2 h planned

    Naive Bayes for text classification

  2. Day 86Tuesday15 Dec2 h planned

    Sentiment analysis pipeline

  3. Day 87Wednesday16 Dec2 h planned

    Topic modelling with LDA

  4. Day 88Thursday17 Dec2 h planned

    Named entity recognition and POS tagging

  5. Day 89Friday18 Dec2 h planned

    PROJECT 2: TF-IDF + logistic regression text classifier

  6. Day 90Saturday19 Dec3 h planned

    PROJECT 2: evaluate and write up - this is your LLM baseline

  7. Day 91Sunday20 DecReview

    Review the week, finish anything unfinished, rest

Watch

Naive Bayes, Clearly Explained

StatQuest

Stanford CS224N: NLP with Deep Learning

Stanford Online · playlist

Lectures 1–2 cover word vectors in depth.

Project 2

TF-IDF text classifier (the LLM baseline)

A TF-IDF + logistic regression classifier with recorded metrics. This becomes your permanent baseline: every later LLM approach gets compared against it.

  • Reproducible preprocessing + TF-IDF + logistic regression pipeline
  • Precision / recall / F1 per class and a confusion matrix
  • Error analysis: 20 misclassified examples, categorised
  • Saved metrics file the later projects read and compare against
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why build a TF-IDF baseline before trying an LLM?
  • Cheap, fast, strong on narrow tasks
  • Quantifies the value of expensive approaches
  • Fallback and routing option in production
How would you evaluate a summarisation system?
  • ROUGE/BERTScore as smoke tests only
  • Rubric-based judging: faithfulness, coverage, concision
  • Human review on a sample; check factual consistency

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why is Laplace smoothing needed in Naive Bayes?

  2. 2.The most valuable step after training a text classifier is…

  3. 3.NER is evaluated…

  4. 4.Why keep the TF-IDF baseline after you have LLMs?

  5. 5.LDA represents a document as…