AI Engineer Path

Week 40

Interview and launch

Drill the classic questions, run mock interviews, and get applications out.

Why this week matters

Interview performance is a skill you practise, not a by-product of knowledge. This week converts 39 weeks of work into offers.

Done when

Applications are out, tracked, and you have a continuous-learning cadence for after the path ends.

Milestone: Applications out

Concepts

5 lessons · tick each one once you could explain it

Open honestly: models can't reliably separate instructions from data, so injection is mitigated, not solved.

Then layers: treat all retrieved and tool content as untrusted and delimit it; least privilege and scoped credentials; human approval for side effects; don't combine private data, untrusted input and an exfiltration path in one agent; output filtering (URL allow-lists, no auto-rendered images, schema validation); input and output classifiers; monitoring and an injection eval set in CI.

Close with your Flagship 1 example: the planted-instruction test and what it caught.

Going deeper

Name the frameworks: OWASP LLM01, the lethal trifecta, and the design-pattern paper (dual LLM, plan-then-execute, CaMeL). Interviewers want to hear that you know it's unsolved and design around it.

Where this comes back

  • Week 37Prompt injection and guardrails.

Measure: cost by feature, model, prompt section and user; find the top drivers (usually long repeated prompts, agent loops, or using a frontier model for easy traffic).

Levers: prompt caching for stable prefixes; trim context (fewer, reranked chunks; compaction; truncated tool results); route easy requests to smaller models; semantic and exact caching; batch APIs for offline work; cap output length; distil or fine-tune a small model for a high-volume narrow task; self-host at very high volume.

Guard: run the eval suite on every change so quality is proven unchanged, and roll out gradually.

Going deeper

Order levers by savings per unit of risk: caching and context trimming (low risk), routing and output caps (medium, needs evals), distillation or self-hosting (high effort). Quote a plausible savings range for each.

Where this comes back

Be ready to explain: data leakage with concrete examples and the pipeline fix; why accuracy fails under imbalance and what to use; precision vs recall in terms of business cost; bias–variance and how to diagnose it with learning curves; why gradient boosting dominates tabular data; time-series splitting.

Interviewers use these to check whether your LLM evaluation instincts rest on solid foundations.

Revisit the week 6–11 quizzes and aim for full marks without notes.

Going deeper

Expect 'design an ML system' questions too: problem framing, labels, features, baseline, metrics, offline/online evaluation, deployment, monitoring. The same skeleton as LLM system design.

Where this comes back

Do at least three mocks: one system design (RAG at scale, an agent platform, an eval pipeline), one coding (Python, data manipulation, a small LLM integration), one project deep-dive (expect 'what would you do differently?' and 'how do you know it works?').

Use a structure for design: requirements → estimates → high-level design → deep dive on the riskiest component → evals, cost and failure modes → trade-offs.

Record yourself or practise with a peer. Fix the two weakest answers before the real ones.

Going deeper

Record mocks and score yourself on a rubric: clarified requirements? stated assumptions? estimated? went deep on the riskiest part? discussed evals, cost, failure modes and trade-offs?

Track every application: company, role, date, contact, stage, next action, notes on what was asked. Review weekly; patterns in rejections tell you what to fix. Referrals and direct outreach to engineers beat cold applications.

After the path: pick three ongoing inputs (not twenty), such as Simon Willison's blog, Hugging Face daily papers and one newsletter. Read one paper a month, and ship a small project or post every month or two.

The field moves fast, but the foundations you built (evaluation discipline, attention, retrieval, production engineering) change slowly. That is what lasts.

Going deeper

A sustainable cadence: one deep read per week (paper or long post), one small build per month, one write-up per quarter. Prune your inputs ruthlessly to three good sources.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Three mocks

    Do one system design, one coding and one project deep-dive mock; score each against a rubric and fix the weakest answer.

  2. Warm-up

    Application tracker

    Set up a tracker (company, role, contact, stage, next action, notes on questions asked) and send the first 10 applications or referrals.

  3. Stretch

    Learning plan after the path

    Pick three ongoing sources and one quarterly project; put them on your calendar.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 274Monday21 Jun2 h planned

    Interview drill: prompt injection defence

  2. Day 275Tuesday22 Jun2 h planned

    Interview drill: cut the LLM bill 60% without losing quality

  3. Day 276Wednesday23 Jun2 h planned

    Interview drill: classical ML (leakage, imbalance, metric choice)

  4. Day 277Thursday24 Jun2 h planned

    Mock interviews and system design practice

  5. Day 278Friday25 Jun2 h planned

    START APPLYING - set up an application tracker

  6. Day 279Saturday26 Jun3 h planned

    Buffer + set your continuous-learning cadence

  7. Day 280Sunday27 JunReview

    Review the week, finish anything unfinished, rest

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

Cut our LLM bill by 60% without losing quality.
  • Measure and attribute cost first
  • Caching, context trimming, routing, output caps, batch APIs
  • Distillation or self-hosting for high volume
  • Prove quality with evals; roll out gradually
How do you stay current in a field that changes monthly?
  • A few high-signal sources, read deeply
  • Build small things with new tools
  • Anchor on durable foundations: evals, retrieval, systems

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Your opening line on prompt injection should acknowledge…

  2. 2.First step to cut an LLM bill?

  3. 3.What protects quality while cutting costs?

  4. 4.A good system design answer structure ends with…

  5. 5.After the path, how many ongoing learning inputs?