AI Engineer Path

Week 29

Orchestration

Compose LLM calls, tools and data into reliable pipelines.

Why this week matters

Learn the patterns and stay loose on the framework. Most of this is service composition you'll recognise: chains, routers, fallbacks, streams.

Done when

You can build the same small pipeline with LangChain and with plain Python, and say what the framework bought you.

Concepts

7 lessons · tick each one once you could explain it

LangChain's core abstraction is the Runnable: prompts, models, parsers, retrievers and plain functions all expose invoke, batch, stream and async versions. LCEL composes them with |: prompt | model | parser is itself a Runnable.

You get batching, streaming, async, retries and tracing hooks for free on any composed chain. RunnableParallel fans out to several branches; RunnablePassthrough carries inputs forward; RunnableLambda wraps any function.

The trade-off: another layer of abstraction to debug, and frequent API churn. Use it where it saves real work; drop to plain code where it obscures.

python
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

prompt = ChatPromptTemplate.from_messages([
    ("system", "Summarise for a busy engineer in 3 bullets."),
    ("user", "{text}"),
])
chain = prompt | llm | StrOutputParser()
chain.invoke({"text": doc})
chain.batch([{"text": d} for d in docs], config={"max_concurrency": 5})

Going deeper

Current state (LangChain 1.0, released October 2025): LCEL is not deprecated, and it remains the recommended way to compose custom chains and RAG pipelines. For agents, the recommended entry point is now create_agent with middleware, which runs on LangGraph under the hood. Older tutorials that build agents with LCEL or AgentExecutor are outdated.

The langchain-core package holds the stable abstractions (Runnables, messages, prompts); integrations live in separate provider packages. Pin versions: the ecosystem moves fast.

Best resources for this lesson

Sequential chains pass outputs to the next step. Routing sends a request to a specialised chain based on a classification (by a cheap model or rules): billing questions to one prompt, technical ones to another. Routing to a cheaper model for easy requests is one of the biggest cost levers.

Fallbacks (chain.with_fallbacks([backup])) retry with a different model or provider when the primary errors or times out: availability engineering for LLMs.

Keep routes observable: log which route handled each request so you can evaluate each one separately.

Going deeper

Model routing can be learned: train a small classifier on logged requests labelled by which model handled them well enough, then route new requests to the cheapest model predicted to succeed.

Fallbacks should change something meaningful: a different provider (outage), a larger model (capability), or a simpler deterministic response (graceful degradation). Retrying the same model is a retry, not a fallback.

Backend engineer tip: Router + fallback is a load balancer with a circuit breaker. The same failure thinking applies.

Where this comes back

  • Week 36Model routing and semantic caching are core cost tools in production.

Output parsers convert raw model output into Python objects: strings, lists, JSON, Pydantic models. With modern APIs, prefer llm.with_structured_output(MyModel), which uses the provider's native schema enforcement or tool calling under the hood.

Parsing failures still happen. Wrap with retry logic that feeds the error back (the same loop as Project 4).

Keep schemas small and focused per step; one giant schema for a whole document is harder for the model and harder to validate.

python
class Sentiment(BaseModel):
    label: Literal["positive", "negative", "neutral"]
    confidence: float = Field(ge=0, le=1)

structured = llm.with_structured_output(Sentiment)
structured.invoke("The update fixed my crash but the UI is slower.")

Going deeper

In LangChain 1.0, structured output for agents is configured with a response_format (a Pydantic model or schema), using provider-native structured output where available and tool calling otherwise.

Best resources for this lesson

Where this comes back

  • Week 27Same principles as structured output in prompting.

Streaming shows text as it's generated, cutting perceived latency dramatically even though total time is unchanged. chain.stream() or astream() yields chunks; astream_events exposes intermediate steps (retrieval done, tool called) for richer UIs.

Streaming complicates things: you can't validate a full structured response until it's complete, errors can arrive mid-stream, and you need a protocol to the client (Server-Sent Events is the usual choice).

Callbacks fire on events (LLM start/end, tool start/end, errors), and they are how tracing tools like LangSmith and Langfuse capture each step.

Going deeper

Stream semantic events, not just tokens: 'retrieving', 'calling tool X', 'drafting'. Users tolerate latency far better when they can see progress.

OpenTelemetry's GenAI semantic conventions standardise span attributes for LLM calls (model, tokens, operation), so traces are portable across observability tools.

Where this comes back

  • Week 36SSE streaming is a standard deployment pattern.

LlamaIndex focuses on the data side of LLM apps: connectors for many sources, node parsers (chunkers), indexes, retrievers, query engines and response synthesisers. It is quick for document Q&A and has strong advanced-retrieval building blocks.

LangChain is broader and more general-purpose (chains, tools, agents via LangGraph). The two interoperate, and many teams use pieces of both, or neither.

Choose on fit, team familiarity and how much you're willing to depend on fast-moving APIs.

Going deeper

LlamaIndex's strengths are ingestion (LlamaParse for complex PDFs), node parsers and advanced retrieval recipes; its Workflows abstraction covers event-driven multi-step apps.

Best resources for this lesson

Every orchestration pattern is a few dozen lines without a framework: a prompt template is an f-string; a chain is function composition; routing is an if/else on a classifier; fallback is try/except; streaming is an async generator.

Frameworks earn their place with integrations, tracing and standard abstractions. They cost you in debuggability, version churn and hidden behaviour (hidden prompts, hidden retries).

A strong interview answer: 'I can build it either way; here's why I chose this one for this system.'

Going deeper

A common production pattern is a thin internal layer (your own client with retries, tracing and cost accounting) plus framework components only where they clearly help, so a framework upgrade never blocks you.

Backend engineer tip: Same judgement as choosing an ORM versus raw SQL: understand what's generated underneath.

Best resources for this lesson

LangChain 1.0 consolidated agent-building into one function, create_agent(model, tools, system_prompt=...), which runs a tool-calling loop on LangGraph. You get persistence, streaming and human-in-the-loop support from the LangGraph runtime without writing a graph yourself.

Middleware hooks into the agent loop (before and after model calls, around tool calls) to add cross-cutting behaviour: summarising long histories, human approval for certain tools, PII redaction, retries, model fallbacks, call limits. It's the same idea as web-framework middleware.

When you need custom control flow (branches, parallel steps, multi-agent hand-offs), drop down to LangGraph directly (week 35). LCEL remains the tool for deterministic chains.

python
from langchain.agents import create_agent

def get_order_status(order_id: str) -> str:
    """Look up an order's shipping status."""
    return orders.status(order_id)

agent = create_agent(
    model="anthropic:claude-sonnet-5-5",   # any supported "provider:model" string
    tools=[get_order_status],
    system_prompt="You are a concise support assistant.",
)
agent.invoke({"messages": [{"role": "user", "content": "Where is ORD-123456?"}]})

Common pitfalls

  • Following pre-2025 tutorials that use AgentExecutor or LCEL-built agents.
  • Unpinned versions: APIs move quickly; pin and read the migration guide.

Where this comes back

  • Week 35LangGraph gives full control when create_agent isn't enough.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Core

    Same pipeline, two ways

    Build summarise → classify → route with LCEL and again in plain Python with asyncio. Compare lines of code, debuggability and tracing.

  2. Stretch

    Fallback and routing

    Add a cheap-model router plus a fallback provider; simulate an outage and confirm requests still succeed.

  3. Warm-up

    Stream to a browser

    Expose a chain via FastAPI with Server-Sent Events and render tokens live in a minimal HTML page.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 197Monday5 Apr2 h planned

    LangChain core: runnables and LCEL

  2. Day 198Tuesday6 Apr2 h planned

    Chains, routing, fallbacks

  3. Day 199Wednesday7 Apr2 h planned

    Output parsers and structured output

  4. Day 200Thursday8 Apr2 h planned

    Streaming and callbacks

  5. Day 201Friday9 Apr2 h planned

    LlamaIndex basics and how it compares

  6. Day 202Saturday10 Apr3 h planned

    CampusX GenAI with LangChain catch-up

  7. Day 203Sunday11 AprReview

    Review the week, finish anything unfinished, rest

Watch

Generative AI using LangChainPrimary

CampusX · playlist

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

When would you use a framework like LangChain, and when would you avoid it?
  • Use: integrations, tracing, standard agent runtime
  • Avoid: simple pipelines, when abstractions hide behaviour
  • Either way: know the underlying pattern
How do you reduce LLM cost with routing?
  • Classify request difficulty cheaply
  • Send easy traffic to small models, hard to large
  • Measure quality per route with evals

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.In LCEL, `prompt | llm | parser` produces…

  2. 2.Routing simple requests to a cheaper model mainly improves…

  3. 3.Streaming improves…

  4. 4.with_fallbacks() is best described as…

  5. 5.A main cost of using a heavy orchestration framework?