Open honestly: models can't reliably separate instructions from data, so injection is mitigated, not solved.
Then layers: treat all retrieved and tool content as untrusted and delimit it; least privilege and scoped credentials; human approval for side effects; don't combine private data, untrusted input and an exfiltration path in one agent; output filtering (URL allow-lists, no auto-rendered images, schema validation); input and output classifiers; monitoring and an injection eval set in CI.
Close with your Flagship 1 example: the planted-instruction test and what it caught.
Going deeper
Name the frameworks: OWASP LLM01, the lethal trifecta, and the design-pattern paper (dual LLM, plan-then-execute, CaMeL). Interviewers want to hear that you know it's unsolved and design around it.
Best resources for this lesson
Where this comes back
- Week 37Prompt injection and guardrails.