AI Engineer Path

Week 35

Agentic AI (2/2): LangGraph, MCP + Flagship 2

Build a stateful agent with approval gates, tracing and a cost ceiling, plus an MCP server.

  • Never cut
  • Flagship project

Why this week matters

This is Flagship 2: an agent that is safe to leave running. MCP is how tools are increasingly shared across agents and apps.

Done when

Flagship 2: a multi-step agent with tools, persistent memory, human approval gates on side-effectful actions, full tracing and a hard cost ceiling. Plus one MCP server exposing a real tool.

Milestone: Flagship 2: agent + MCP server

Concepts

6 lessons · tick each one once you could explain it

In LangGraph you define a typed state (a TypedDict or Pydantic model), nodes (functions that take state and return updates) and edges. Conditional edges route based on state: 'if the last message has tool calls, go to the tool node; otherwise end'. Cycles make loops natural.

State updates can use reducers: for example messages appends rather than overwrites.

Compile the graph with a checkpointer and every step is persisted per thread. The graph is also a diagram you can show in your README.

python
from langgraph.graph import StateGraph, START, END, MessagesState
from langgraph.prebuilt import ToolNode, tools_condition
from langgraph.checkpoint.memory import MemorySaver

def agent(state: MessagesState):
    return {"messages": [llm_with_tools.invoke(state["messages"])]}

g = StateGraph(MessagesState)
g.add_node("agent", agent)
g.add_node("tools", ToolNode(tools))
g.add_edge(START, "agent")
g.add_conditional_edges("agent", tools_condition)   # tools or END
g.add_edge("tools", "agent")
app = g.compile(checkpointer=MemorySaver())
app.invoke({"messages": [("user", "Refund order ORD-123456")]}, {"configurable": {"thread_id": "t1"}})

Going deeper

LangGraph 1.0 is the runtime underneath LangChain's create_agent. Use the graph API directly when you need explicit control flow: parallel branches (fan-out/fan-in), sub-graphs per agent, or deterministic steps mixed with agentic ones.

The Send API dynamically fans out work (one branch per item, map-reduce style), useful for processing many documents in parallel within one graph.

LangGraph can interrupt before a node (or inside one with interrupt()), persisting state. Your UI shows the pending action and its arguments; a human approves, edits or rejects; the graph resumes from the checkpoint with that decision.

Gate every side-effectful or irreversible action: sending messages, payments, deletes, writes to production systems. Show the exact arguments, not a model-written summary of them.

Approval gates turn 'the agent might do something bad' into 'the agent proposes, a human disposes'.

Going deeper

Design approval UX for speed and safety: show the exact tool and arguments, the reason the agent gives, and the predicted effect, with approve/edit/reject. Log the decision and the reviewer for audit.

Common pitfalls

  • Approving a summary while the actual tool arguments differ.

Best resources for this lesson

Beyond per-thread checkpoints, LangGraph stores hold long-term memory across threads (user profiles, learned facts), namespaced per user.

Streaming modes expose node updates and tokens as they happen, so UIs can show 'searching… found 3 results… drafting'.

Sub-graphs compose: a research sub-agent becomes a node in a larger graph. The LangChain Academy course covers these patterns with exercises; work through it this week.

Going deeper

Long-term memory in LangGraph lives in a store with namespaces (per user, per org); the agent reads and writes through tools or explicit nodes, keeping memory separate from the per-thread checkpoint.

Best resources for this lesson

Without a standard, every app integrates every tool separately (M × N integrations). MCP defines a client–server protocol: an MCP server exposes tools (functions the model can call), resources (data it can read) and prompts (templates). Any MCP-capable client (desktop assistants, IDEs, agent frameworks) can connect to any server.

Transports: stdio for local servers launched as subprocesses, and Streamable HTTP for remote servers. Messages are JSON-RPC. As of the 2026-07-28 specification the protocol core is stateless: there is no initialize handshake or session ID, every request carries its protocol version and client capabilities, and clients can call server/discover to learn a server's versions and capabilities up front. Sampling and roots are deprecated in favour of calling model APIs directly and passing paths as tool parameters.

Security matters: an MCP server runs with real permissions, and tool descriptions themselves can carry prompt injection. Install only servers you trust and scope their credentials.

Going deeper

The 2026-07-28 specification also replaced the HTTP GET notification endpoint with subscriptions/listen, introduced multi-round-trip requests in place of server-initiated sampling and elicitation, and moved client registration from dynamic client registration to Client ID Metadata Documents. Read the changelog before building on older tutorials.

Anthropic's 'code execution with MCP' pattern lets agents call MCP tools from generated code instead of loading every tool definition into context, cutting token use dramatically for large tool sets.

Best resources for this lesson

Where this comes back

  • Week 37Tool poisoning and supply-chain risks apply to MCP servers.

With the standalone FastMCP library, decorate functions with @mcp.tool. Type hints and docstrings become the tool schema and description. Add resources with @mcp.resource(uri). (The official mcp SDK bundled an older FastMCP; in its v2 the class is renamed MCPServer, so check which package and version a tutorial uses.)

Build one that exposes a real tool you'd use: query your Flagship 1 RAG system, search your notes, or check a service's status. Test it with the MCP Inspector, then connect it to a client.

Apply week 34 hygiene: validate inputs, scope credentials, return clear errors, log calls.

python
from fastmcp import FastMCP   # pip install fastmcp
mcp = FastMCP("docs-search")

@mcp.tool
def search_docs(query: str, k: int = 5) -> list[dict]:
    """Search the engineering handbook. Use for questions about internal processes."""
    return [{"id": h.id, "title": h.title, "snippet": h.text[:300]} for h in rag.search(query, k)]

if __name__ == "__main__":
    mcp.run()   # stdio by default; mcp.run(transport="http") for remote

Going deeper

Test servers with the MCP Inspector before connecting a client. For remote servers, put authentication (OAuth) and per-user authorisation in front of every tool, because a remote MCP server is a public API.

Framework options: LangGraph (explicit graphs, durable state), OpenAI Agents SDK (lightweight agents and handoffs), CrewAI (role-based multi-agent), smolagents (minimal, code-writing agents). The concepts transfer; pick by control needs.

Non-negotiables for Flagship 2: tracing of every LLM call and tool call (Langfuse or LangSmith) with tokens and cost; a hard cost ceiling per run enforced in code; loop guards (max steps, detecting repeated identical calls); timeouts on every tool.

Show sample traces in the README: one successful run, one where a guard stopped a bad run. That second trace is more convincing than the first.

Going deeper

Evaluating agents needs more than final-answer checks: score the trajectory (right tools, sensible order, no unnecessary calls), the end state (did the database actually change correctly?) and cost per task. Anthropic's guide to agent evals covers graders, environments and pitfalls.

Backend engineer tip: Budgets and circuit breakers belong in code, not in the prompt.

Where this comes back

  • Week 36Traces and cost data feed observability dashboards and evals.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    MCP server

    Build an MCP server with FastMCP exposing a real tool (your RAG search or notes), test it with the MCP Inspector, and connect it to a client.

  2. Warm-up

    Approval gate

    Add a LangGraph interrupt before any side-effectful tool; resume with approve, edit and reject paths and confirm each works.

  3. Stretch

    Flagship 2

    Ship the agent with persistent memory, tracing in Langfuse, a hard cost ceiling, loop guards, and README traces of one success and one stopped run.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 239Monday17 May2 h planned

    LangGraph: graphs, conditional edges, human-in-the-loop interrupts

  2. Day 240Tuesday18 May2 h planned

    LangChain Academy LangGraph course

  3. Day 241Wednesday19 May2 h planned

    MCP: protocol concepts, clients and servers

  4. Day 242Thursday20 May2 h planned

    MILESTONE: build an MCP server exposing a real tool

  5. Day 243Friday21 May2 h planned

    FLAGSHIP 2: multi-step agent - tools, memory, approval gates

  6. Day 244Saturday22 May3 h planned

    FLAGSHIP 2: tracing, hard cost ceiling, loop guards

  7. Day 245Sunday23 MayReview

    Review the week, finish anything unfinished, rest

Watch

Agentic AI using LangGraphPrimary

CampusX · playlist

Model Context Protocol

CampusX · playlist

Building Agents with MCP: Full Workshop

Mahesh Murag, Anthropic · AI Engineer

Flagship project

Flagship 2: agent + MCP server

A multi-step agent that is safe to leave running: tools, memory, approval gates and a hard cost ceiling.

  • LangGraph agent with tools and persistent memory
  • Human approval gate on every side-effectful action
  • Full tracing (Langfuse or LangSmith) and loop guards
  • Hard cost ceiling per run
  • An MCP server exposing a real tool
Track it on the Projects page

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

What is MCP and what changed in the 2026 spec?
  • Standard protocol for exposing tools, resources, prompts to any client
  • JSON-RPC over stdio or Streamable HTTP
  • 2026-07-28: stateless core (no handshake/sessions), server/discover, sampling and roots deprecated
How would you stop an agent running up a huge bill?
  • Hard per-run cost and step limits in code
  • Loop detection on repeated identical calls
  • Timeouts; tracing and alerts on spend

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.In LangGraph, conditional edges…

  2. 2.What problem does MCP solve?

  3. 3.An approval gate should show the human…

  4. 4.Where should an agent's cost ceiling be enforced?

  5. 5.Which MCP transport launches a local server as a subprocess?