AI Engineer Path

Week 39

Positioning and job prep

Reposition yourself as an AI engineer and start the interview drills.

Why this week matters

Start applying in week 39, not week 40. You'll be more ready than you feel, and early interviews are diagnostic: they tell you what to fix while you still have runway.

Done when

Résumé, GitHub and portfolio are live, one technical post is published, and the first applications are out.

Milestone: Résumé, GitHub and portfolio live

Concepts

6 lessons · tick each one once you could explain it

Put an 'AI engineering projects' section first, with each flagship as a bullet group: what it does, the stack, and numbers (eval scores, latency, cost reductions).

Recast previous engineering experience in terms AI teams value: reliability, scaling, observability, API design, distributed systems, security. These are exactly the skills LLMOps and agent roles struggle to hire for.

Mirror the vocabulary of the roles you target (RAG, evals, agents, LLMOps) honestly. Every keyword should be defensible in an interview.

Going deeper

Use quantified bullets: action + system + measurable result ('Built hybrid-search RAG over 12k docs; raised recall@5 from 0.71 to 0.83 and cut cost/query 26%'). One page for under ~10 years of experience; keep the AI projects above the fold.

Backend engineer tip: Your backend years are not a past life; they are the production half of 'AI engineer'. Say so explicitly.

Best resources for this lesson

GitHub: a profile README with a one-line positioning statement and links to the two flagships and the from-scratch GPT. Pin six repositories. Make sure each pinned repository runs from its README instructions.

Write one substantial post: e.g. 'What actually improved my RAG system: measured' or 'Hardening a RAG app against indirect prompt injection'. Specific, with numbers and code. Share it on LinkedIn.

A post with real measurements is a conversation starter in interviews and evidence of communication skills.

Going deeper

The best technical posts teach one thing with evidence: a problem, what you tried, the numbers, and what you'd do next. Publish where the audience is (your blog plus LinkedIn), and link it from your résumé and pinned repos.

Practise the chain: text → tokens (BPE) → embeddings → transformer blocks (attention moves information between positions, MLPs compute per position) → logits → sampling.

Then cost: prefill processes the prompt in parallel (time to first token); decode generates one token per step using the KV cache (tokens per second); the KV cache grows with context, consuming memory; attention compute grows with context length; long prompts are re-sent and re-billed every turn unless cached.

Rehearse aloud and time yourself. Interviewers listen for whether you understand why, not whether you know the terms.

Going deeper

Have numbers ready: KV cache per token ≈ 2 × layers × KV heads × head dim × bytes; decode is memory-bandwidth-bound; prefill is compute-bound; GQA cuts the KV cache; attention compute is O(n²) in context length.

Best resources for this lesson

Where this comes back

  • Week 19Attention and the transformer block.
  • Week 26Prefill, decode and caching.

Clarify: query volume, freshness needs, document types, access control, latency target (time to first token vs full answer), languages. Estimate: 10M docs × ~10 chunks = ~100M vectors; at 768 dimensions float32 that's ~300 GB raw, so quantisation or PQ, sharding, and a careful index choice matter.

Ingestion: queue-based pipeline, incremental updates, deduplication, chunking with metadata and ACLs. Query path: query rewriting (cached), hybrid retrieval with permission filters, sharded ANN, cross-encoder rerank on top 50, streaming generation, caching of frequent queries. Budget the latency per stage.

Operations: retrieval and generation evals, tracing, index rebuild strategy, embedding-model migration plan, cost per query.

108 vectors×768×4 B≈307 GB10^8 \text{ vectors} \times 768 \times 4\text{ B} \approx 307\text{ GB}

Going deeper

Latency budget example for 'sub-second': query rewrite 0 ms (skip or cached), embedding 20 ms, hybrid ANN 30–60 ms, rerank top-50 80–150 ms, time-to-first-token 300–500 ms, then stream. Say which stage you'd cut first if over budget.

No ground truth? Start with error analysis on real traces; write binary criteria per failure mode; label a small set yourself; use LLM-as-judge validated against those labels; use reference-free metrics (faithfulness to retrieved context, schema validity); collect user feedback; run paired comparisons between versions.

RAG vs fine-tuning: RAG for changing, citable, large or permissioned knowledge; fine-tuning for format, style, behaviour, latency or cost (distilling into a smaller model). Usually both. Start with prompting + RAG and fine-tune for a measured gap.

Have one concrete story from your own projects for each.

Going deeper

Bring a story with a number: 'I found 40% of failures were retrieval misses on product codes; adding BM25 fixed most of them, recall@5 went from 0.71 to 0.83.'

Best resources for this lesson

Where this comes back

1. Clarify: users, inputs and outputs, scale (requests per second, corpus size), latency target, quality bar, constraints (privacy, residency, budget). 2. Estimate: tokens per request, cost per request and per month, storage for embeddings, GPU needs if self-hosting.

3. Baseline design: the simplest thing that could work (often prompt + RAG), drawn end to end with offline and online paths. 4. Deep dive: pick the riskiest component (retrieval quality, agent safety, latency) and go deep with alternatives and their trade-offs.

5. Close with operations: how you'd evaluate (offline set, online signals), observe (tracing, cost), secure (injection, PII, permissions) and iterate. Interviewers weigh this last part heavily for AI roles because it's what separates demos from products.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Résumé rewrite

    Rewrite your résumé with AI projects first and every bullet quantified. Have someone in the field review it.

  2. Stretch

    Publish one post

    Write and publish one technical post with real numbers from a flagship.

  3. Warm-up

    Timed explanations

    Record yourself explaining attention + KV cache in 3 minutes and RAG design in 10. Re-record until both fit.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 267Monday14 Jun2 h planned

    GitHub profile, pinned repos, portfolio site

  2. Day 268Tuesday15 Jun2 h planned

    Resume rewrite: Java backend -> AI engineer framing

  3. Day 269Wednesday16 Jun2 h planned

    LinkedIn update + write one technical post about a project

  4. Day 270Thursday17 Jun2 h planned

    Interview drill: attention, KV cache, why context costs what it does

  5. Day 271Friday18 Jun2 h planned

    Interview drill: design RAG for 10M docs with sub-second latency

  6. Day 272Saturday19 Jun3 h planned

    Interview drill: evaluating with no ground truth; RAG vs fine-tuning

  7. Day 273Sunday20 JunReview

    Review the week, finish anything unfinished, rest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Design a customer-support assistant over 10M documents with sub-second time to first token.
  • Clarify scale, latency, ACLs, freshness
  • Sharded hybrid ANN + rerank, permission filters, caching
  • Streaming, citations, abstain; evals, tracing, cost per query
  • Injection hardening and human escalation
How would you evaluate a system with no ground-truth labels?
  • Error analysis on real traces → failure taxonomy
  • Binary judges validated against human labels
  • Reference-free checks, user feedback, paired comparisons

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Best time to start applying?

  2. 2.Rough raw size of 100M 768-dimensional float32 vectors?

  3. 3.First step when asked to evaluate a system with no labelled data?

  4. 4.A strong technical post is…

  5. 5.How should prior backend experience appear on an AI engineer résumé?