AI Engineer Path

Week 21

NLP with deep learning: Hugging Face

Use and fine-tune pretrained transformers with the Hugging Face stack.

Why this week matters

You'll beat your own week 13 baseline with a fine-tuned transformer, and be able to say by exactly how much and at what cost.

Done when

A fine-tuned BERT-family classifier beating your week 13 TF-IDF baseline, and you know by how much.

Milestone: Fine-tuned BERT classifier vs your baseline

Concepts

5 lessons · tick each one once you could explain it

The Hugging Face Hub hosts hundreds of thousands of pretrained models. pipeline('sentiment-analysis') or pipeline('zero-shot-classification') gives a working model in one line; the Auto* classes (AutoTokenizer, AutoModelForSequenceClassification) load any architecture by name.

Families: encoder-only (BERT, RoBERTa, DeBERTa) for classification, NER and embeddings; decoder-only (GPT, LLaMA, Mistral) for generation; encoder–decoder (T5, BART) for translation and summarisation.

Read the model card: training data, intended use, licence, known limitations.

python
from transformers import pipeline
clf = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")
clf("The refund never arrived", candidate_labels=["billing", "shipping", "account"])

Going deeper

The Auto classes read config.json to instantiate the right architecture; AutoModelForX adds the right task head. device_map='auto' spreads big models across available devices.

Check a model's licence before use: many open-weight models restrict commercial use or require attribution.

Transformers use sub-word tokenization: common words are single tokens, rare words split into pieces ('tokenization' → 'token', '##ization' in WordPiece). Algorithms: WordPiece (BERT), BPE (GPT), Unigram/SentencePiece (T5, LLaMA).

A tokenizer returns input_ids, an attention_mask (1 for real tokens, 0 for padding) and special tokens ([CLS], [SEP], <eos>). Always use the tokenizer that matches the model.

Token counts drive cost, latency and context limits for every LLM you'll use. Non-English text and code often need more tokens per word.

python
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("bert-base-uncased")
enc = tok(["Tokenization is subtle"], padding=True, truncation=True, max_length=128)
tok.convert_ids_to_tokens(enc["input_ids"][0])
# ['[CLS]', 'token', '##ization', 'is', 'subtle', '[SEP]']

Going deeper

Fast tokenizers (Rust-backed) return offset mappings from tokens back to character spans, essential for NER and extractive QA, where labels live on characters.

Training a new tokenizer for a specialised domain (code, chemistry, a new language) can cut token counts substantially, but you then can't reuse a pretrained model's embeddings directly.

Best resources for this lesson

Where this comes back

  • Week 25You'll implement BPE by hand.

load_dataset('imdb') downloads and caches a dataset; local CSV/JSON/Parquet files load the same way. Data is stored as memory-mapped Apache Arrow, so datasets bigger than RAM are fine.

dataset.map(fn, batched=True) applies tokenization at speed and caches the result. train_test_split, filter, select and shuffle cover most needs.

Use the same splits as your week 13 baseline so the comparison is fair.

Going deeper

map with batched=True and num_proc parallelises preprocessing; results are cached by fingerprint so re-runs are instant. set_format('torch') hands tensors straight to PyTorch.

Dataset cards document provenance, licence and known biases. Read them before training, and write one for any dataset you publish.

Best resources for this lesson

AutoModelForSequenceClassification puts a fresh linear head on top of the encoder's [CLS] representation. Fine-tune all weights for 2–4 epochs with a small learning rate (2e-5 to 5e-5), warmup and weight decay. Large rates destroy the pretrained knowledge ('catastrophic forgetting').

The Trainer API handles the loop, evaluation, checkpointing and mixed precision. Pass a compute_metrics function so you track F1, not just loss.

Report the result against your TF-IDF baseline on the same test set, with cost: training time, inference latency per example, model size.

python
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
model = AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased", num_labels=2)
args = TrainingArguments("out", learning_rate=3e-5, num_train_epochs=3, per_device_train_batch_size=16,
                         warmup_ratio=0.1, weight_decay=0.01, eval_strategy="epoch", bf16=True)
trainer = Trainer(model=model, args=args, train_dataset=tok_train, eval_dataset=tok_val,
                  compute_metrics=compute_f1)
trainer.train()

Going deeper

Run several seeds: fine-tuning BERT-class models on small datasets is notoriously seed-sensitive, and a single run can over- or under-state the gain.

Smaller distilled models (DistilBERT, MiniLM) often lose little accuracy while running several times faster, so try them before the large variants.

Where this comes back

  • Week 13Compare against the TF-IDF baseline you recorded.
  • Week 33LoRA makes this affordable for billion-parameter models.

trainer.push_to_hub() uploads weights, tokenizer and config. Write a model card: data, metrics, intended use, limitations. Version models like code.

Token classification (NER), extractive question answering (predict answer start/end spans), summarisation and translation (seq2seq) all follow the same load → tokenize → fine-tune → evaluate pattern with different heads and metrics.

Many of these tasks are now done by prompting an LLM. The engineering question is when a small fine-tuned model is cheaper, faster and good enough. Often it is.

Going deeper

Model cards (Mitchell et al.) standardise what to disclose: intended use, evaluation across groups, limitations and ethical considerations. Regulators increasingly expect this documentation.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Tokenizer comparison

    Tokenize the same 5 sentences (English, Hindi, code, numbers, emoji) with BERT, GPT-2 and a modern LLM tokenizer. Count tokens and explain the differences.

  2. Core

    Beat your baseline

    Fine-tune DistilBERT on the Project 2 task, run 3 seeds, and report mean ± std F1 vs the TF-IDF baseline, plus latency and cost.

  3. Stretch

    Push to the Hub

    Publish the fine-tuned model with a proper model card (data, metrics, limitations).

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 141Monday8 Feb2 h planned

    HF LLM Course ch1-2: transformer models and pipelines

  2. Day 142Tuesday9 Feb2 h planned

    ch3: tokenizers in depth

  3. Day 143Wednesday10 Feb2 h planned

    ch4: datasets library

  4. Day 144Thursday11 Feb2 h planned

    ch5: fine-tuning a pretrained model

  5. Day 145Friday12 Feb2 h planned

    ch6-7: sharing models, classic NLP tasks

  6. Day 146Saturday13 Feb3 h planned

    Fine-tune a BERT-family model on a classification task

  7. Day 147Sunday14 FebReview

    Review the week, finish anything unfinished, rest

Watch

Hugging Face CoursePrimary

Hugging Face · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Why do transformers use sub-word tokenization?
  • Small vocabulary, no out-of-vocabulary words
  • Frequent words single tokens; rare words split
  • Trade-off: token count varies by language and domain
How would you fine-tune BERT for a classification task?
  • Pretrained encoder + new classification head
  • Small LR (2e-5–5e-5), warmup, few epochs
  • Evaluate vs baseline on same split; multiple seeds

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Typical learning rate for fine-tuning BERT?

  2. 2.What is the attention_mask for?

  3. 3.Which family suits text classification and NER best?

  4. 4.Your fine-tuned model scores +3 F1 over TF-IDF. What else should you report?

  5. 5.'tokenization' → 'token', '##ization' is an example of…