AI Engineer Path

Week 27

Prompt engineering

Get reliable, structured, tool-using behaviour from models.

Why this week matters

Prompts are the first lever on every LLM feature, and structured output plus tool calling are the foundation of everything in the AI application weeks.

Done when

You have a reusable prompt library, not a folder of one-off strings.

Concepts

6 lessons · tick each one once you could explain it

Zero-shot: instructions only. Works for common tasks if the instructions are clear: say what to do, for whom, in what format, and what to do in edge cases. Few-shot: add 2–5 input→output examples. They teach format and subtle judgement better than descriptions, but choose diverse, representative examples, because models copy their quirks.

Use delimiters (XML-style tags like <document>…</document> work very well) to separate instructions from untrusted data. This improves accuracy and is a first layer of defence against injection.

The system prompt sets the role, rules and context that persist; the user turn carries the task. Explain why a rule exists; models follow reasoned instructions more robustly than bare commands.

text
System: You classify customer emails for a support team.
Categories: billing, shipping, account, other.
If an email fits several, pick the one the customer needs resolved first.

User: <email>
I was charged twice and my package still hasn't arrived.
</email>
Return only the category name.

Going deeper

Order matters for long prompts: put long documents first and the question and instructions at the end. Models often follow end-of-prompt instructions more reliably at long context.

Tell the model what to do instead of what not to do ('respond in plain prose paragraphs' beats 'don't use markdown'), and explain the reason for unusual constraints.

The CoT paper showed that including worked reasoning in few-shot examples, or simply asking the model to think step by step, substantially improves accuracy on arithmetic, logic and multi-step questions. The model gets 'scratch space' in tokens before committing to an answer.

Structure it: ask for reasoning inside <thinking> tags and the final answer inside <answer> tags, so code can parse only the answer.

Costs: more output tokens and latency. Reasoning models do this internally, so explicit CoT prompting matters less for them. Visible reasoning is also not a faithful trace of how the model reached its answer.

Going deeper

Zero-shot CoT ('let's think step by step') works surprisingly well; few-shot CoT with worked examples works better on specialised tasks. For reasoning models, explicit CoT instructions are mostly redundant.

Visible chain-of-thought isn't guaranteed to reflect the model's actual computation, so don't rely on it as an audit trail of 'why'.

Where this comes back

  • Week 34ReAct agents interleave reasoning with actions.

Self-consistency: sample several reasoning paths at moderate temperature and take the majority answer. It improves accuracy on problems with a single correct answer, at N× the cost.

Prompt chaining: break a complex task into steps, each a separate focused call whose output feeds the next (extract → validate → summarise → format). Each step is easier to get right, test and debug.

Chaining is the bridge to orchestration frameworks (week 29) and is usually better than one giant prompt.

Going deeper

Tree-of-thoughts generalises self-consistency into a search over partial reasoning steps with evaluation and backtracking, powerful but expensive, and largely subsumed by reasoning models.

Backend engineer tip: A prompt chain is a pipeline of small services: test each stage in isolation.

Levels of reliability: asking for JSON in the prompt (fragile); JSON mode (guarantees valid JSON, not your schema); structured outputs / constrained decoding with a JSON Schema (guarantees schema-conformant output on supporting APIs).

Define the schema once with Pydantic and generate the JSON Schema from it. Validate every response anyway, and on failure retry with the validation error included so the model can correct itself. Cap the retries.

Design schemas for the model: descriptive field names and descriptions, enums for categorical fields, optional/nullable fields where data may be missing (so the model isn't forced to invent values).

python
from pydantic import BaseModel, Field, ValidationError
from typing import Literal

class Ticket(BaseModel):
    category: Literal["billing", "shipping", "account", "other"]
    urgency: int = Field(ge=1, le=5, description="5 = customer is blocked")
    summary: str = Field(max_length=200)

def extract(email: str, attempts=3) -> Ticket:
    msg = f"<email>{email}</email>"
    for _ in range(attempts):
        raw = call_llm(msg, response_schema=Ticket.model_json_schema())
        try:
            return Ticket.model_validate_json(raw)
        except ValidationError as e:
            msg += f"\n\nYour last output was invalid: {e}. Return corrected JSON."
    raise RuntimeError("extraction failed")

Going deeper

Constrained decoding is enforced by the inference engine (grammar-based token masking), so schema validity is guaranteed while semantic correctness still isn't. Validate values (ranges, cross-field rules) after parsing.

Keep schemas shallow and fields well described. Very deep or very large schemas raise latency and error rates, so split extraction into steps when needed.

Common pitfalls

  • Required fields with no source data: the model will invent them.
  • Trusting output without validation because 'it's JSON mode'.

Best resources for this lesson

Where this comes back

  • Week 28Project 4 is built around exactly this loop.

You send tool definitions (name, description, JSON Schema for arguments). The model may respond with a tool call instead of text. Your code runs the function, then sends the result back as a tool-result message, and the model continues, possibly calling more tools.

The model never executes anything itself. You control execution, so validate arguments, enforce permissions and handle errors (return a useful error message to the model rather than crashing).

Tool descriptions are prompts: say when to use the tool, when not to, and what it returns. This loop is the entire foundation of agents.

python
tools = [{
  "name": "get_order_status",
  "description": "Look up the shipping status of one order. Use when the user asks where an order is.",
  "input_schema": {"type": "object",
                   "properties": {"order_id": {"type": "string", "pattern": "^ORD-\\d{6}$"}},
                   "required": ["order_id"]},
}]
# loop: response = llm(messages, tools)
#   if response has tool_use -> result = run_tool(...) -> append tool_result -> call llm again
#   else -> final answer

Going deeper

Tool definitions are part of the prompt: names, descriptions and parameter docs steer the model's choices. Anthropic's guidance on writing tools for agents recommends fewer, higher-level tools with meaningful, token-efficient outputs.

Parallel tool calls let the model request several independent lookups at once; execute them concurrently and return all results together.

Where this comes back

  • Week 34Agents are tool calling in a loop.

Store prompts as files or code objects with: a name and version, a template with typed variables, the model and parameters it was tuned for, example inputs, and a small regression set of input→expected-property cases.

Change prompts like code: diff, review, run the regression set, and record the result. A prompt tweak that fixes one case often breaks three others, so only measurement tells you.

Keep the instructions, the few-shot examples and the dynamic data visibly separate in the template.

Going deeper

Version prompts alongside code (or in a prompt registry), tag each production call with the prompt version, and keep a regression set per prompt. This makes 'which prompt produced this output?' answerable from any trace.

Backend engineer tip: Prompts are configuration that changes behaviour. Version them, test them, and roll them out like code.

Where this comes back

  • Week 36The regression set grows into a proper eval suite in CI.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Prompt teardown

    Take a weak prompt, improve it in 4 steps (role, structure with XML tags, examples, output format) and measure accuracy on 30 cases after each step.

  2. Core

    Validated extraction

    Extract a Pydantic schema from 50 emails with structured outputs, retries on validation failure, and a report of failure types.

  3. Stretch

    Tool-calling loop by hand

    Implement the tool-calling loop without a framework: 2 tools, argument validation, error messages returned to the model, a max-iterations guard.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 183Monday22 Mar2 h planned

    Zero-shot, few-shot, delimiters, role prompting

  2. Day 184Tuesday23 Mar2 h planned

    Chain-of-thought + read the CoT paper

  3. Day 185Wednesday24 Mar2 h planned

    Self-consistency and prompt chaining

  4. Day 186Thursday25 Mar2 h planned

    Structured output, JSON mode, schema enforcement

  5. Day 187Friday26 Mar2 h planned

    Function and tool calling basics

  6. Day 188Saturday27 Mar3 h planned

    Anthropic prompt engineering course + build your own prompt library

  7. Day 189Sunday28 MarReview

    Review the week, finish anything unfinished, rest

Watch

Prompting 101Primary

Anthropic

How I use LLMs

Andrej Karpathy

Paper of the week

Chain-of-Thought Prompting Elicits Reasoning in LLMs

Why asking for intermediate steps improves answers, and the seed of today's reasoning models.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How do you make LLM outputs reliably machine-readable?
  • Schema-constrained structured outputs or tool calling
  • Validate with Pydantic; retry with the error
  • Nullable fields for missing data; low temperature
How does function/tool calling work?
  • Tools described with JSON Schema
  • Model returns a tool call; your code executes it
  • Result sent back; loop until a final answer
How do you manage prompts in production?
  • Versioned like code with tests
  • Regression set per prompt; evals on change
  • Trace every call with prompt version

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Why wrap user-provided documents in tags like <document>?

  2. 2.JSON mode guarantees…

  3. 3.In tool calling, who executes the tool?

  4. 4.Self-consistency improves accuracy by…

  5. 5.Your schema has a required 'invoice_number' but many documents lack one. Risk?