AI Engineer Path

Week 30

Vector databases

Choose embeddings, distance metrics and indexes with data, not vibes.

Why this week matters

pgvector is frequently the right production answer, and many teams already run Postgres. Don't reach for a specialised vector database reflexively.

Done when

You've benchmarked recall@k across three embedding models and chosen with data.

Milestone: Embedding model benchmark

Concepts

7 lessons · tick each one once you could explain it

A sentence embedding model maps text to a dense vector so semantically similar texts are close. Similarity can be measured by cosine (angle only), dot product (angle and magnitude) or Euclidean (L2) distance.

For unit-normalised vectors, all three produce the same ranking: dot product equals cosine, and squared L2 distance equals 2 − 2·cosine. Many systems normalise at insert time and use the fast inner product.

Use whatever the model card specifies. Mixing metrics silently degrades retrieval.

∥a−b∥2=∥a∥2+∥b∥2−2 a⋅b=2−2cos⁡θ(∥a∥=∥b∥=1)\|a - b\|^2 = \|a\|^2 + \|b\|^2 - 2\,a\cdot b = 2 - 2\cos\theta \quad (\|a\|=\|b\|=1)

Going deeper

Matryoshka embeddings are trained so the first k dimensions are themselves a good embedding, letting you truncate 1024-d vectors to 256-d for a cheap first pass and re-score with full vectors.

Binary and int8 quantised embeddings cut storage 4–32× with a rescoring step to recover most of the accuracy.

Where this comes back

  • Week 4The dot product and cosine similarity, at scale.

Sentence-transformer models (bi-encoders) embed queries and documents independently, so documents are embedded once, offline. The MTEB leaderboard compares models across tasks: look at the retrieval scores, not the overall average.

Trade-offs: dimension (storage and speed), max input length, multilinguality, licence, hosted vs self-hosted, cost per million tokens. Some models need instruction prefixes (like query: and passage: ); forgetting them quietly hurts quality.

Benchmarks are a shortlist, not a decision. Your domain (legal, medical, code, your company's jargon) can reorder the rankings, which is why you benchmark on your own data this week.

python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
D = model.encode(passages, normalize_embeddings=True, batch_size=64)
q = model.encode(["how do refunds work?"], normalize_embeddings=True)
scores = D @ q.T

Going deeper

Bi-encoders (embed separately) scale; cross-encoders (score pairs jointly) are accurate. Late-interaction models like ColBERT sit in between: they keep one vector per token and match token by token.

Asymmetric retrieval (short query, long passage) benefits from models trained for it; check whether the model expects query/passage prefixes or instructions.

HNSW builds a layered proximity graph. Build parameters M (links per node) and ef_construction; query parameter ef_search. Higher values raise recall, memory and latency. It is the default in most vector databases and in pgvector.

IVF clusters vectors with k-means and searches only the nearest nprobe clusters. PQ compresses each vector into short codes, cutting memory 10–30× at some recall cost. IVF-PQ suits billion-scale collections.

Always measure recall@k of the ANN index against exact (brute-force) search on a sample, at the latency you need. Filtering (metadata WHERE clauses) interacts with the index and can quietly reduce recall.

Going deeper

Filtered search is the hard part in practice: post-filtering can return too few results, and pre-filtering can break HNSW's graph connectivity. Databases solve it differently (filter-aware HNSW, partitioning by tenant). Test recall with your real filters.

Where this comes back

  • Week 3The graph and clustering intuition from DSA week.

Chroma runs in-process (or as a server), stores embeddings with documents and metadata, and can compute embeddings for you. Ideal for notebooks and prototypes.

Core operations: create a collection with a distance function, add ids + documents + metadata, query with text or vectors and an optional metadata where filter.

Move to a production store once you need concurrency, replication, backups and access control.

python
import chromadb
client = chromadb.PersistentClient(path="./db")
col = client.get_or_create_collection("docs", metadata={"hnsw:space": "cosine"})
col.add(ids=ids, documents=chunks, metadatas=[{"source": s} for s in sources])
col.query(query_texts=["refund policy"], n_results=5, where={"source": "handbook.pdf"})

Going deeper

Treat the vector store as derived data: keep source documents and chunking code as the source of truth so you can rebuild the index when you change embedding models or chunking.

Best resources for this lesson

pgvector adds a vector type and HNSW/IVFFlat indexes to Postgres. Embeddings live next to your relational data: one backup story, transactions, joins with metadata, and access control you already have. It handles millions of vectors well.

Qdrant (and Weaviate, Milvus) are purpose-built: rich filtering integrated with the index, quantisation, sharding, hybrid search, and better behaviour at very large scale or very high query rates.

Decide on: scale, filter complexity, operational overhead, and whether a second database is worth it. 'We used pgvector because we already run Postgres, and here's the recall and latency we measured' is a strong answer.

sql
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (id bigserial PRIMARY KEY, doc_id text, body text, embedding vector(384));
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
SELECT id, body, 1 - (embedding <=> $1) AS cosine
FROM chunks WHERE doc_id = ANY($2)
ORDER BY embedding <=> $1 LIMIT 5;

Going deeper

pgvector tips: build the HNSW index after bulk loading, raise maintenance_work_mem for faster builds, set hnsw.ef_search per query for recall, and use iterative index scans (recent versions) to make filtered queries return enough rows.

Supabase and other managed Postgres services include pgvector, so there's often no new infrastructure to run at all.

Backend engineer tip: Fewer moving parts is a feature. Reach for a new database only when measurements say you must.

Create 50–200 realistic queries, each labelled with the chunk(s) that answer it. You can bootstrap by having an LLM write questions for sampled chunks, then review them by hand. Synthetic questions are often easier than real ones.

Metrics: recall@k (did any relevant chunk appear in the top k?), MRR (how high was the first relevant one?), nDCG for graded relevance. Measure for each embedding model, plus latency and cost.

Choose with data, record the numbers in your README, and keep the query set: it becomes the retrieval half of your RAG eval suite.

Recall@k=1∣Q∣∑q∈Q1[rel(q)∩topk(q)≠∅],MRR=1∣Q∣∑q1rankq\text{Recall@}k = \frac{1}{|Q|}\sum_{q\in Q}\mathbb{1}\big[\text{rel}(q)\cap \text{top}_k(q) \neq \varnothing\big],\qquad \text{MRR} = \frac{1}{|Q|}\sum_q \frac{1}{\text{rank}_q}

Going deeper

Synthetic queries skew easy because they reuse the chunk's wording. Mix in real user queries, paraphrases and multi-hop questions, and track metrics by query type.

Best resources for this lesson

Where this comes back

  • Week 32Retrieval metrics are half of the Flagship 1 eval suite.

General embedding models struggle with jargon, internal product names and domain-specific notions of 'similar'. Fine-tuning a bi-encoder on your own (query, positive passage) pairs, with in-batch negatives or mined hard negatives, adapts the space to your domain.

Training data can be bootstrapped by having an LLM write questions for your chunks, then filtering them, the same synthetic-query trick used for evaluation. Keep a held-out set to prove the gain.

Fine-tuning means re-embedding the corpus and versioning the model. Measure recall@k before and after, and try a reranker first, since it is often a cheaper win.

python
from sentence_transformers import SentenceTransformer, losses
from sentence_transformers.trainer import SentenceTransformerTrainer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
loss = losses.MultipleNegativesRankingLoss(model)   # in-batch negatives
trainer = SentenceTransformerTrainer(model=model, train_dataset=pairs, loss=loss)
trainer.train()

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Three-model benchmark

    Build a 100-query labelled set and measure recall@5 and MRR for three embedding models in pgvector or Chroma, plus embedding cost and latency.

  2. Stretch

    Tune HNSW

    Vary ef_search and M; plot recall vs latency against exact search.

  3. Warm-up

    Metric sanity check

    Show numerically that cosine, dot product and L2 give identical rankings on normalised vectors, and different ones on unnormalised.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 204Monday12 Apr2 h planned

    Embeddings: cosine vs dot vs Euclidean

  2. Day 205Tuesday13 Apr2 h planned

    Sentence Transformers; choosing models via the MTEB leaderboard

  3. Day 206Wednesday14 Apr2 h planned

    ANN indexes: HNSW and IVF-PQ

  4. Day 207Thursday15 Apr2 h planned

    Chroma hands-on

  5. Day 208Friday16 Apr2 h planned

    Qdrant and pgvector hands-on

  6. Day 209Saturday17 Apr3 h planned

    MILESTONE: benchmark recall@k across three embedding models

  7. Day 210Sunday18 AprReview

    Review the week, finish anything unfinished, rest

Watch

Generative AI using LangChainPrimary

CampusX · playlist

HNSW for Vector Search, Explained and Implemented

James Briggs

Cosine Similarity, Clearly Explained

StatQuest

Vector databases are so hot right now

Fireship

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How would you choose an embedding model?
  • Shortlist from MTEB retrieval scores
  • Benchmark recall@k on your own labelled queries
  • Weigh dimension/cost, max length, language, licence
pgvector or a dedicated vector database?
  • pgvector: one system, transactions, joins, ACLs; fine for millions
  • Dedicated: scale, rich filtering, quantisation, high QPS
  • Decide on measured recall/latency and ops cost

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.For unit-normalised vectors, cosine similarity equals…

  2. 2.On the MTEB leaderboard, which column matters most for RAG?

  3. 3.Raising HNSW's ef_search typically…

  4. 4.You run Postgres and have 3M chunks with relational filters. Sensible first choice?

  5. 5.Recall@5 = 0.82 means…