AI Engineer Path

Week 34

Agentic AI (1/2): planning, tools and state

Understand agents as loops, and engineer their tools like production APIs.

  • Never cut

Why this week matters

An agent calling tools is a distributed system with a non-deterministic scheduler. Timeouts, retries with backoff, circuit breakers, dead-letter queues and compensating transactions all apply directly, and most people building agents have never encountered them.

Done when

You can explain when to use a workflow vs an agent, and your tools are idempotent, permissioned and safe to retry.

Concepts

7 lessons · tick each one once you could explain it

A workflow follows code paths you define, with LLM calls at fixed steps. An agent lets the model decide the sequence: it chooses a tool, sees the result, decides what to do next, and repeats until it judges the task done or hits a limit.

Agents suit open-ended tasks where the steps can't be known in advance (research, debugging, multi-system operations). They cost more, are slower and less predictable. Start with the simplest thing (a single call, then a workflow) and move to an agent only when needed.

Core pieces: a model, tools, instructions, memory/state, and guardrails: step limits, cost ceilings, approval gates.

Going deeper

Anthropic's distinction is the cleanest: workflows orchestrate LLMs through predefined code paths; agents dynamically direct their own process and tool use. Their advice is to find the simplest solution and add agency only when it demonstrably improves outcomes.

Chip Huyen frames agents by their tools (read-only knowledge tools vs write actions), planning capability and failure modes, a useful structure for design reviews.

Best resources for this lesson

Where this comes back

  • Week 33Workflows are the simpler alternative to reach for first.

ReAct interleaves reasoning traces with actions. The model writes a thought ('I need the order date first'), emits an action (a tool call), receives an observation (the tool result), and continues reasoning with that new information.

Reasoning keeps the plan coherent; acting grounds it in real data and reduces hallucination. Modern tool-calling APIs implement this loop natively; the 'thoughts' may be internal reasoning.

Failure modes to design against: loops (calling the same tool repeatedly), premature stopping, hallucinated tool arguments, and over-trusting a bad observation. Step through a run, including a failing tool and a retry, in the simulation.

python
def run_agent(task, tools, max_steps=10, budget_usd=0.50):
    messages, spent = [user(task)], 0.0
    for step in range(max_steps):
        resp = llm(messages, tools=tools)
        spent += resp.cost
        if spent > budget_usd: return stop("cost ceiling reached")
        if not resp.tool_calls: return resp.text            # done
        messages.append(resp.message)
        for call in resp.tool_calls:
            result = execute_safely(call)                     # validate, permission, timeout
            messages.append(tool_result(call.id, result))
    return stop("step limit reached")

Going deeper

A 'think' tool (a no-op tool where the model writes its reasoning mid-task) measurably helps in long tool-use sequences with policies to follow, a cheap trick from Anthropic's engineering notes.

ReAct's weakness is myopia: it decides one step at a time. Plan-and-execute variants generate an explicit plan first and re-plan on surprises.

Where this comes back

  • Week 27Tool calling is the primitive; ReAct is the loop around it.

Decomposition / plan-and-execute: first produce a plan of sub-tasks, then execute them, re-planning when results change things. It keeps long tasks on track and makes progress inspectable.

Reflection: after a draft or a failed attempt, have the model critique it against explicit criteria (or test results) and revise. It works best with an external signal (failing tests, a validator, a checklist) rather than pure self-judgement.

Orchestrator–workers: a lead agent splits work and delegates to specialised sub-agents with focused contexts. Powerful, but multiplies cost and failure points.

Going deeper

Reflexion stores verbal feedback from failed attempts in memory and retries with it, effective when there's an external success signal (tests, an environment reward).

Evaluator–optimiser loops (one call generates, another critiques against criteria) work best when the criteria are explicit and checkable.

Design tools for the model as user: few, well-named tools with clear descriptions of when to use them; tight argument schemas (enums, patterns, required fields); outputs trimmed to what's useful. Prefer one tool that does a complete job over several the model must chain correctly.

Models can emit parallel tool calls for independent lookups: execute them concurrently.

Errors are information: return a clear, actionable error ('order_id must look like ORD-123456; got 123456') instead of a stack trace, so the model can correct itself. Distinguish retryable errors from permanent ones.

Going deeper

Design tools like a good API for a new teammate: consolidate common workflows into one tool, return meaningful identifiers rather than opaque UUIDs, paginate and truncate by default, and evaluate tools with realistic agent tasks, not just unit tests.

Common pitfalls

  • Twenty overlapping tools that confuse the model.
  • Tools returning huge raw payloads into the context.

Idempotency: agents and frameworks retry. A tool that charges a card or sends an email must accept an idempotency key so a retry doesn't do it twice. Least privilege: give each agent only the tools and scopes its task needs, with credentials scoped per user, not a shared admin key.

Approval gates: side-effectful or irreversible actions (payments, deletes, external messages) require human confirmation. Sandboxing: run code execution and file access in isolated containers with no network or secrets by default.

Compensating actions for multi-step operations that fail halfway (the saga pattern), and an audit log of every tool call with arguments and results.

Going deeper

Stripe's idempotency design is the reference: clients send a unique key per logical operation, and the server stores the first result and returns it for repeats. Apply the same contract to every agent tool with side effects.

The saga pattern sequences local transactions with compensating actions for rollback, the right model for multi-step agent workflows that touch several systems.

Backend engineer tip: This is where backend experience is worth the most: it's distributed-systems hygiene with a non-deterministic caller. Say exactly this in interviews.

Where this comes back

  • Week 37Excessive agency is on the OWASP LLM Top 10.

An agent run is a long-lived process. Represent its state explicitly (messages, plan, intermediate results, counters) and checkpoint it after each step to durable storage.

Checkpointing enables resuming after crashes, pausing for human approval and continuing later, inspecting exactly what happened ('time travel' debugging), and replaying from a step with a fix.

LangGraph implements this with checkpointers keyed by a thread ID. The same idea is a workflow engine or a state machine with persisted transitions.

Going deeper

Durable execution engines (Temporal, and LangGraph's checkpointers) record each step's inputs and outputs so a crashed run resumes exactly where it stopped, without re-running side effects.

Backend engineer tip: Think durable workflow engine: explicit state, persisted transitions, resumable execution.

Best resources for this lesson

Where this comes back

  • Week 35LangGraph's checkpointers and interrupts build on this.

In an orchestrator–worker design, a lead agent breaks a task down, spawns sub-agents that each work in a fresh, focused context (often in parallel), and combines their condensed results. Anthropic's multi-agent research system found this helps most on breadth-first tasks, like researching many independent directions at once.

The costs are real: multi-agent runs use many times more tokens than single-agent chats, coordination failures (duplicated work, conflicting assumptions) are common, and debugging and evaluation get harder.

Rule of thumb: start with one agent and good tools; add sub-agents when the task naturally parallelises, a single context window is the bottleneck, and the value justifies the cost.

Common pitfalls

  • Multi-agent by default, for tasks one agent handles fine.
  • Vague delegation: sub-agents need clear objectives, output formats and boundaries.

Where this comes back

  • Week 35LangGraph sub-graphs implement orchestrator–worker designs.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    A ReAct loop by hand

    Implement a ReAct agent without a framework: 3 tools, schema validation, error messages back to the model, a step limit and a cost ceiling.

  2. Warm-up

    Idempotent side effects

    Give a 'send email' tool an idempotency key and prove that retried calls send only once.

  3. Stretch

    Trajectory evals

    Write 10 agent tasks with expected tool sequences and end states; score your agent on success, steps and cost.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 232Monday10 May2 h planned

    HF Agents Course unit 1: what an agent actually is

  2. Day 233Tuesday11 May2 h planned

    Planning & Reasoning: the ReAct loop + read the ReAct paper

  3. Day 234Wednesday12 May2 h planned

    Reflection, self-critique, task decomposition

  4. Day 235Thursday13 May2 h planned

    Tool Use: schema design, parallel calls, error handling

  5. Day 236Friday14 May2 h planned

    Tool Use: idempotency, permissions, sandboxing (apply your backend instincts)

  6. Day 237Saturday15 May3 h planned

    State & Memory: LangGraph state, checkpointing, resumability

  7. Day 238Sunday16 MayReview

    Review the week, finish anything unfinished, rest

Watch

How We Build Effective AgentsPrimary

Barry Zhang, Anthropic · AI Engineer

Agentic AI using LangGraphPrimary

CampusX · playlist

What's next for AI agentic workflows

Andrew Ng · Sequoia

Tips for building AI agents

Anthropic

Paper of the week

ReAct: Synergizing Reasoning and Acting in Language Models

The loop underneath nearly every agent framework you will use.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How do you make agent tool calls safe?
  • Least privilege, scoped per-user credentials
  • Idempotency keys, timeouts, retries only for transient errors
  • Human approval for side effects; sandbox code execution; audit logs
Workflow vs agent: how do you choose?
  • Known steps → workflow (cheaper, predictable)
  • Open-ended, unknown steps → agent
  • Start simple; add agency where it measurably helps
What is the lethal trifecta?
  • Private data + untrusted content + external communication
  • Together enable data exfiltration via injection
  • Break at least one leg by design

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.When should you choose an agent over a workflow?

  2. 2.A payment tool gets called twice because of a retry. The fix?

  3. 3.A tool fails with bad arguments. Best response to the model?

  4. 4.Why checkpoint agent state after every step?

  5. 5.In ReAct, observations…