AI Engineer Path

Week 36

LLMOps (1/2): UX, deployment, observability and evals

Ship LLM features like production services, and prove changes with numbers.

  • Never cut

Why this week matters

Of the four job profiles, LLMOps Engineer is where backend engineers are most obviously qualified and least contested. 'How do you know it got better?' is the question that separates candidates.

Done when

Evals run in CI as a deploy gate.

Milestone: Evals wired into CI as a deploy gate

Concepts

7 lessons · tick each one once you could explain it

Latency: stream tokens, show progress for multi-step work ('searching 3 sources…'), and use optimistic UI where possible. Uncertainty: show sources and citations, signal low confidence, and make 'I couldn't find that' a first-class answer.

Control: let users edit, regenerate, give feedback (thumbs plus a reason) and undo. Keep the human in charge of consequential actions.

Failure states: timeouts, refusals, rate limits and bad outputs all need designed responses, not raw errors. Feedback signals double as eval data.

Going deeper

Set expectations at the point of use (what the feature can and can't do), make errors recoverable (edit and regenerate), and collect feedback with a reason, not just thumbs. Google's People + AI Guidebook has worked patterns for each.

Interactive: an async API (FastAPI) streaming tokens to the browser over Server-Sent Events, with timeouts, cancellation when the client disconnects, and per-user rate limits.

Background: long jobs (document processing, agent runs) go through a queue and workers, with status polling or webhooks. Offline batch jobs can use providers' batch APIs at a discount, or self-hosted batched inference.

Plus the usual: config and secrets management, gradual rollouts, feature flags for prompt and model versions, and the ability to roll back a prompt as easily as code.

python
from fastapi.responses import StreamingResponse

@app.post("/chat")
async def chat(req: ChatRequest):
    async def events():
        async for chunk in llm.astream(req.messages):
            yield f"data: {json.dumps({'delta': chunk.text})}\n\n"
        yield "data: [DONE]\n\n"
    return StreamingResponse(events(), media_type="text/event-stream")

Going deeper

SSE works through most proxies and load balancers but needs buffering disabled and keep-alive comments for long gaps. WebSockets suit bidirectional interaction (voice, collaborative editing).

Detect client disconnects and cancel the upstream LLM request, or you pay for tokens nobody reads.

Best resources for this lesson

vLLM serves open models at high throughput using continuous batching and PagedAttention (efficient KV-cache memory). Ollama and llama.cpp run quantised models locally: great for development, privacy and offline use. Self-hosting trades API cost for GPU and operations cost; do the maths at your volume.

Exact caching returns stored responses for identical requests. Semantic caching returns a stored response when a new query's embedding is very close to a previous one: high savings for repetitive traffic, with a risk of wrong hits (tune the threshold and scope per user).

Model routing sends each request to the cheapest model that handles it well, often via a small classifier or rules. Together with prompt caching, these are the main levers for cutting an LLM bill.

Going deeper

Self-hosting break-even: compare GPU hourly cost at realistic utilisation (often 30–60%) against API token prices for your traffic. Low or spiky volume almost always favours APIs; steady high volume, privacy or latency requirements can favour self-hosting.

Where this comes back

  • Week 40'Cut the LLM bill by 60% without losing quality' is a classic interview question.

A trace captures one request end to end: retrieval with the retrieved chunks, each LLM call with prompt, response, model, tokens, cost and latency, each tool call, and the final output, all linked. Langfuse, LangSmith and Phoenix provide this, mostly on OpenTelemetry concepts.

Dashboards to build: cost per request, feature and user; p50/p95 latency and time-to-first-token; error, refusal and validation-failure rates; user feedback rates.

Traces are also your best eval data source: sample production traces, label the bad ones, and add them to the eval set.

Going deeper

Sample traces for human review every week: production traces are where new failure modes and eval cases come from. Tag traces with prompt version, model, user segment and feature for slicing.

Backend engineer tip: Distributed tracing for LLM calls. Same spans and attributes mindset, with tokens and cost as first-class metrics.

Start with error analysis: read real outputs, categorise failures, and write evals for the failure modes that matter. Build a golden dataset of inputs with expected outputs or explicit pass/fail criteria. Include edge cases and adversarial inputs. Version it.

Graders: code-based checks wherever possible (schema valid, contains citation, exact match, regex, unit tests); LLM-as-judge for subjective qualities, using a clear rubric and preferably binary pass/fail per criterion rather than 1–10 scores.

Judges have failure modes: position bias, verbosity bias, self-preference, leniency. Validate the judge against human labels on a sample and measure agreement before trusting it. Use the statistics from week 5: confidence intervals, paired comparisons.

Going deeper

Hamel Husain and Shreya Shankar's method: error analysis first. Review real traces and write free-form notes (open coding), group them into a failure taxonomy (axial coding), then build evals only for failure modes that matter. Prefer binary pass/fail judgements per failure mode over Likert scales.

Validate LLM judges like classifiers: measure true positive and true negative rates against human labels on a held-out set, and iterate on the judge prompt until agreement is good enough to trust.

Common pitfalls

  • Optimising for an unvalidated judge.
  • Tiny eval sets that make noise look like progress.
  • Evaluating only on easy cases.

Where this comes back

  • Week 5A/B testing logic decides whether a change really helped.
  • Week 32Ragas metrics are LLM-judged evals for RAG.

Treat prompts, model versions, retrieval settings and tool definitions as code. On every pull request, run a fast eval subset (DeepEval or pytest-style assertions); before release, run the full suite.

Gate on thresholds and on regressions versus the last release: e.g. faithfulness must not drop more than 2 points and the schema-valid rate must stay at 100%. Post the results table on the PR.

Exercise: retrofit this onto Flagship 1, swap the embedding model, and let the gate prove or block the change with numbers.

python
# tests/test_rag_evals.py, run in CI
import pytest
from evals import load_golden, run_pipeline, judge_faithfulness

CASES = load_golden("evals/golden_v3.jsonl")

@pytest.mark.parametrize("case", CASES, ids=lambda c: c["id"])
def test_answer_is_grounded(case):
    out = run_pipeline(case["question"])
    if case["answerable"]:
        assert out.citations, "answer must cite sources"
        assert judge_faithfulness(out) == "pass"
    else:
        assert out.abstained, "should say it doesn't know"

Going deeper

Keep two tiers: a fast, cheap subset (deterministic checks plus a few judged cases) on every pull request, and the full suite nightly or before release. Track metrics over time on a dashboard so slow drifts are visible.

Inspect (UK AI Security Institute) and promptfoo provide eval harnesses with CI integration; DeepEval brings pytest-style assertions.

Backend engineer tip: It's a regression test suite whose assertions are partly probabilistic. Track pass rates over time, not just pass/fail.

Offline evals use fixed datasets before deployment. Online evaluation watches live traffic: user feedback (thumbs, edits, regenerations, abandonment), implicit signals (did they copy the answer? escalate to a human?), and automated judges run on a sample of production traces.

Roll out changes as A/B tests or canaries with guardrail metrics (error rate, cost, latency) and outcome metrics (task success). Pre-register what 'better' means to avoid chasing noise.

Close the loop: sample failing production traces weekly, label them, add them to the offline eval set, and fix. That flywheel is the core of LLMOps.

Best resources for this lesson

Where this comes back

  • Week 5A/B testing statistics apply directly.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Trace everything

    Instrument Flagship 1 with Langfuse: retrieval spans, LLM spans with tokens and cost, user feedback. Build a cost-per-request chart.

  2. Core

    Error analysis session

    Review 50 traces, write open-coded notes, group them into a failure taxonomy, and write one binary eval per top failure mode.

  3. Stretch

    Evals as a deploy gate

    Wire the eval suite into CI; swap the embedding model and let the gate pass or block it with numbers.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 246Monday24 May2 h planned

    UX & Product Design: streaming, uncertainty, citations, failure states

  2. Day 247Tuesday25 May2 h planned

    Deployment: serving patterns, SSE streaming, async batching

  3. Day 248Wednesday26 May2 h planned

    Deployment: vLLM, Ollama, semantic caching, model routing

  4. Day 249Thursday27 May2 h planned

    Observability: tracing with Langfuse; token and cost accounting

  5. Day 250Friday28 May2 h planned

    Evals: golden datasets, LLM-as-judge and its failure modes

  6. Day 251Saturday29 May3 h planned

    Evals: DeepEval regression suite; wire evals into CI as a deploy gate

  7. Day 252Sunday30 MayReview

    Review the week, finish anything unfinished, rest

Watch

LLM Evals: Common MistakesPrimary

Hamel Husain & Shreya Shankar

p-values: What they are and how to interpret them

StatQuest

LLM Eval Office Hours: Start With Error Analysis

Hamel Husain

How to Build AI Evals, Step by Step

Aakash Gupta, with Hamel Husain & Shreya Shankar

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How do you know an LLM feature got better?
  • Error analysis → failure taxonomy → targeted evals
  • Fixed eval set, paired comparison, confidence intervals
  • Online signals and A/B tests in production
How do you trust an LLM-as-judge?
  • Binary, rubric-based criteria
  • Measure agreement with human labels (TPR/TNR)
  • Watch for position, verbosity and self-preference bias
Design observability for an LLM app.
  • Traces with retrieval, LLM and tool spans
  • Tokens, cost, latency (TTFT, p95), errors per feature
  • Prompt/model version tags; feedback capture; alerting

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Semantic caching risks…

  2. 2.Before trusting an LLM-as-judge, you should…

  3. 3.What does vLLM's continuous batching improve?

  4. 4.A good eval-gate rule in CI is…

  5. 5.The best source of new eval cases is…