AI Engineer Path
Phase 1 · Prerequisites28 Sept – 4 Oct

Week 2

Python for services

The Python equivalents of a production backend toolkit.

Why this week matters

Pydantic and asyncio are the two you will use hardest later: Pydantic for structured LLM outputs, asyncio for concurrent model calls.

Done when

A small service of yours runs as FastAPI, containerised, with passing tests.

Milestone: Port a small service to FastAPI + pytest + Docker

Concepts

8 lessons · tick each one once you could explain it

Annotations like def score(doc: str, k: int = 5) -> list[float]: are not enforced by the interpreter. Static checkers (mypy, pyright) read them and flag mismatches before you run anything. Your editor uses them for autocompletion.

Useful types: list[str], dict[str, int], X | None (optional), Literal['low', 'high'], Callable[[str], int], TypedDict for dict shapes, and Protocol for structural interfaces (duck typing with checking).

In AI code, hints pay off twice: they make data flowing between pipeline stages explicit, and libraries like Pydantic and FastAPI use them at runtime to validate and to generate JSON Schema.

python
from typing import Literal, Protocol

class Embedder(Protocol):
    def embed(self, texts: list[str]) -> list[list[float]]: ...

def search(q: str, embedder: Embedder, k: int = 5,
           mode: Literal["dense", "hybrid"] = "dense") -> list[str]:
    ...

Going deeper

Generics let containers stay precise: def first[T](xs: list[T]) -> T (3.12 syntax) or TypeVar. TypedDict describes JSON-like dicts; Literal and Enum describe fixed choices; Annotated[int, Field(gt=0)] attaches metadata that Pydantic and FastAPI read.

Run mypy or pyright in CI with --strict on new code. Gradual typing means you can start with your public interfaces and tighten over time.

Common pitfalls

  • Assuming hints validate input at runtime. They don't, unless a library like Pydantic reads them.
Backend engineer tip: `Protocol` is Python's answer to Java interfaces, but structural: any class with matching methods qualifies.

Best resources for this lesson

A Pydantic BaseModel turns type hints into runtime validation. Invoice.model_validate(data) either returns a typed object or raises a ValidationError that lists exactly which fields failed and why.

Add constraints with Field(gt=0, max_length=200), custom logic with @field_validator, and cross-field checks with @model_validator. model_json_schema() emits a JSON Schema: exactly what LLM APIs accept for structured output and tool definitions.

pydantic-settings loads configuration from environment variables and .env files into a typed settings object, which keeps API keys out of code.

python
from pydantic import BaseModel, Field, field_validator

class LineItem(BaseModel):
    description: str
    amount: float = Field(gt=0)

class Invoice(BaseModel):
    vendor: str
    currency: str = Field(pattern=r"^[A-Z]{3}$")
    items: list[LineItem]

    @field_validator("vendor")
    @classmethod
    def strip(cls, v: str) -> str:
        return v.strip()

schema = Invoice.model_json_schema()  # hand this to an LLM

Going deeper

Pydantic v2's core is written in Rust, so validation is fast enough to sit in hot paths. Strict mode (model_config = ConfigDict(strict=True)) disables coercion when you want '12' to be an error rather than 12.

Discriminated unions (Field(discriminator='type')) model 'one of several shapes': perfect for tool-call outputs or event types. TypeAdapter validates plain types (like list[Invoice]) without a wrapper model.

Common pitfalls

  • Mixing Pydantic v1 and v2 APIs (.dict() vs .model_dump()).
  • Over-permissive coercion: a string '12' becomes int 12 unless you use strict mode.

Best resources for this lesson

Where this comes back

  • Week 27Structured output and tool calling are defined with JSON Schema from Pydantic models.
  • Week 28Project 4 retries an LLM call when Pydantic validation fails.

An LLM call spends almost all its time waiting on the network. asyncio lets one thread start many such calls and switch between them while each waits. async def defines a coroutine; await pauses it until the awaited thing is ready, letting other tasks run.

asyncio.gather(*coros) runs coroutines concurrently and collects results. asyncio.TaskGroup (3.11+) does the same with structured error handling. An asyncio.Semaphore caps how many run at once, which is how you respect a provider's rate limit.

async is cooperative: a CPU-heavy function or a blocking call (time.sleep, requests.get) inside a coroutine freezes everything. Use async libraries (httpx, asyncpg) or push blocking work to a thread with asyncio.to_thread.

python
import asyncio

sem = asyncio.Semaphore(8)   # at most 8 calls in flight

async def classify(text: str) -> str:
    async with sem:
        await asyncio.sleep(0.5)        # stands in for an API call
        return "positive"

async def main(texts):
    return await asyncio.gather(*(classify(t) for t in texts))

labels = asyncio.run(main(["a", "b", "c"]))

Going deeper

The event loop runs one coroutine at a time and switches only at await points. That's why shared state rarely needs locks in async code, and why one blocking call stalls everything.

Cancellation is cooperative: task.cancel() raises CancelledError at the next await. Always re-raise it after cleanup. asyncio.timeout(10) (3.11+) wraps a block with a deadline, and TaskGroup cancels sibling tasks if one fails, which is usually what you want.

Common pitfalls

  • Calling a blocking library inside async code.
  • Unbounded gather over 10,000 items: you hit rate limits instantly. Always use a semaphore.
Backend engineer tip: Think CompletableFuture without threads: an event loop multiplexes waits instead of a thread pool.

Best resources for this lesson

Where this comes back

  • Week 36Async batching and concurrent calls are core serving patterns.

httpx offers sync and async clients with the same API. Reuse one client (it pools connections), and always set a timeout: the default of waiting forever is how services hang.

Retry only transient failures: timeouts, HTTP 429 (rate limited) and 5xx. Never retry 400-class errors that mean your request is wrong. Use exponential backoff with jitter (wait 0.5s, 1s, 2s… plus randomness) so many clients don't retry in lockstep. Honour a Retry-After header when present.

Libraries like tenacity express this declaratively. LLM SDKs ship their own retry logic: know what it does so you don't stack retries on retries.

waitn=min⁡(cap, base⋅2 n)+jitter\text{wait}_n = \min\left(\text{cap},\ \text{base}\cdot 2^{\,n}\right) + \text{jitter}
python
import httpx
from tenacity import retry, stop_after_attempt, wait_random_exponential, retry_if_exception

def transient(e):
    return isinstance(e, httpx.TimeoutException) or (
        isinstance(e, httpx.HTTPStatusError) and e.response.status_code in (429, 500, 502, 503))

@retry(retry=retry_if_exception(transient), wait=wait_random_exponential(max=20), stop=stop_after_attempt(5))
async def post(client: httpx.AsyncClient, url, payload):
    r = await client.post(url, json=payload, timeout=30)
    r.raise_for_status()
    return r.json()

Going deeper

Retries amplify load during an outage. Production systems add a retry budget (retry at most ~10% of requests) and a circuit breaker that stops calling a failing dependency for a cooldown period, then probes with a few requests before reopening.

Set separate connect, read and pool timeouts. For streaming LLM responses the read timeout is per chunk, not the whole response, so a slow but steady stream won't trip it.

Common pitfalls

  • Retrying non-idempotent requests can double-charge or double-write.
  • Stacked retries (SDK + yours + gateway) multiply load during an outage.

Best resources for this lesson

Where this comes back

  • Week 28Project 4 adds rate limiting, timeouts and backoff around LLM calls.
  • Week 34Agent tool calls need the same discipline, plus idempotency.

FastAPI reads your function signature: path and query parameters come from arguments, request bodies from Pydantic models, and the return model (response_model) shapes and validates the output. Interactive docs appear at /docs.

Dependency injection via Depends() supplies shared resources (a DB session, an API client, the current user) to endpoints and makes them easy to override in tests. Use async def endpoints when the work is awaitable I/O.

Lifespan events (@asynccontextmanager passed as lifespan=) are where you create long-lived clients and load models once at startup rather than per request.

python
from fastapi import FastAPI, Depends
from pydantic import BaseModel

app = FastAPI()

class Query(BaseModel):
    question: str
    k: int = 5

class Answer(BaseModel):
    answer: str
    sources: list[str]

def get_retriever():
    return MyRetriever()

@app.post("/ask", response_model=Answer)
async def ask(q: Query, retriever=Depends(get_retriever)):
    docs = await retriever.search(q.question, q.k)
    return Answer(answer="...", sources=[d.id for d in docs])

Going deeper

Sync (def) endpoints run in a thread pool; async (async def) endpoints run on the event loop. Mixing a blocking library into an async def endpoint freezes the whole server: use def, or an async client.

BackgroundTasks handle small fire-and-forget work after the response; anything long or important belongs in a real queue (Celery, RQ, Arq) with retries and visibility.

Backend engineer tip: `Depends` is constructor injection, function-scoped. Lifespan is your `@PostConstruct`.

Best resources for this lesson

Where this comes back

  • Week 11Project 1 is deployed behind FastAPI.

pytest discovers test_*.py files and test_* functions. Plain assert statements give detailed failure diffs. Fixtures (@pytest.fixture) provide reusable setup and are injected by argument name. @pytest.mark.parametrize runs one test over many inputs.

For FastAPI, TestClient (or httpx.AsyncClient with an ASGI transport) exercises endpoints in-process, and app.dependency_overrides swaps real dependencies for fakes. Mock the network boundary, not your own logic.

The habit matters more than the tool: in weeks 32 and 36 you will write evals as tests, and gate deploys on them.

python
import pytest
from fastapi.testclient import TestClient
from app import app, get_retriever

class FakeRetriever:
    async def search(self, q, k): return []

@pytest.fixture
def client():
    app.dependency_overrides[get_retriever] = FakeRetriever
    yield TestClient(app)
    app.dependency_overrides.clear()

@pytest.mark.parametrize("k", [1, 5, 20])
def test_ask(client, k):
    r = client.post("/ask", json={"question": "hi", "k": k})
    assert r.status_code == 200

Going deeper

Fixture scopes (function, module, session) control how often setup runs; conftest.py shares fixtures across files. monkeypatch swaps environment variables and attributes safely; tmp_path gives a clean directory per test.

Snapshot or golden-file testing (store expected outputs, diff on change) is the bridge to LLM evals, where exact matches give way to scored comparisons.

Best resources for this lesson

Where this comes back

  • Week 36DeepEval brings pytest-style tests to LLM outputs; evals run in CI.

uv replaces pip, venv and pip-tools with one fast tool: uv init, uv add fastapi, uv run pytest. It writes pyproject.toml and a uv.lock lockfile so builds are reproducible. Ruff lints and formats in milliseconds.

A container image packages your code with its exact runtime. Keep images small (a python:3.12-slim base), install dependencies in a layer before copying source so rebuilds are cached, run as a non-root user, and pass secrets as environment variables, never baked into the image.

This is the deployment unit for every project from here on: Project 1, both flagships and any model you serve.

dockerfile
FROM python:3.12-slim
COPY --from=ghcr.io/astral-sh/uv:latest /uv /bin/uv
WORKDIR /app
COPY pyproject.toml uv.lock ./
RUN uv sync --frozen --no-dev
COPY . .
USER nobody
CMD ["uv", "run", "uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

Going deeper

Multi-stage builds keep images small: install and compile in a builder stage, copy only the virtualenv and code into a slim runtime stage. Pin the base image digest for reproducibility, and add a .dockerignore so .git, data and secrets never enter the build context.

For GPU workloads, start from NVIDIA CUDA or PyTorch base images that match your driver, and expect multi-gigabyte images. Model weights usually belong in a volume or object storage, not the image.

Common pitfalls

  • Copying source before installing dependencies invalidates the cache on every change.
  • Committing .env files with API keys.

Best resources for this lesson

Git stores snapshots of your project as commits linked into a history. A branch is a movable pointer to a commit; you develop on a branch, then merge it back. git status, add, commit, switch -c, push, pull and log --oneline --graph cover most days.

Work in small, focused commits with messages that say why. Open a pull request for review, let CI run tests, then merge. Rebase rewrites your branch onto the latest main for a clean history; merge preserves exactly what happened. Never rewrite history that others have pulled.

For ML projects, keep large data and model files out of Git (use .gitignore, object storage, or tools like DVC and Git LFS), and never commit secrets: a leaked API key in history stays leaked even after you delete the file.

bash
git switch -c feat/retry-client        # new branch
git add -p                             # stage hunks selectively
git commit -m "Add retry with jitter to LLM client"
git fetch origin && git rebase origin/main
git push -u origin feat/retry-client   # then open a pull request

Common pitfalls

  • Committing .env files or API keys.
  • Giant commits mixing refactors, features and formatting.
  • Force-pushing over a shared branch.

Best resources for this lesson

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Validate messy JSON

    Write Pydantic models for a nested API response with optional fields, enums and a discriminated union. Feed it 10 deliberately broken payloads and read the error messages.

  2. Core

    Concurrent, polite fetcher

    Fetch 200 URLs with httpx.AsyncClient, at most 10 in flight, a timeout on each, retries with jittered backoff for 429/5xx only. Log successes, retries and failures.

  3. Stretch

    Service port, end to end

    Port a small service to FastAPI with dependency injection, pytest coverage of the API contract, a multi-stage Dockerfile and a GitHub Actions workflow that runs tests on every push.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 8Monday28 Sept2 h planned

    Type hints, mypy basics

  2. Day 9Tuesday29 Sept2 h planned

    Pydantic: models, validators, settings

  3. Day 10Wednesday30 Sept2 h planned

    asyncio: coroutines, tasks, gather, semaphores

  4. Day 11Thursday1 Oct2 h planned

    httpx async client: retries, timeouts, backoff

  5. Day 12Friday2 Oct2 h planned

    FastAPI: routing, dependency injection, response models

  6. Day 13Saturday3 Oct3 h planned

    PROJECT: port a small Spring Boot service to FastAPI + pytest + Docker

  7. Day 14Sunday4 OctReview

    Review the week, finish anything unfinished, rest

Watch

Pydantic: Complete Data Validation Course

Corey Schafer

Asyncio in Python: Full Tutorial

Tech With Tim

FastAPI Course for Beginners

freeCodeCamp

Docker Tutorial for Beginners: Full Course

freeCodeCamp

Learn Docker in 7 Easy Steps

Fireship

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

How does asyncio achieve concurrency on one thread, and when is it the wrong tool?
  • Event loop switches tasks at await points while they wait on I/O
  • Great for many concurrent network calls
  • Wrong for CPU-bound work (use processes) or blocking libraries (use threads or async clients)
Design the retry policy for a client calling an LLM API.
  • Retry only transient errors: timeouts, 429, 5xx
  • Exponential backoff with jitter, honour Retry-After
  • Cap attempts and total time; idempotency for side effects
  • Retry budgets / circuit breaker to avoid amplifying outages
Why use Pydantic models at API boundaries?
  • Runtime validation with precise errors
  • Single source for types, docs and JSON Schema
  • Prevents malformed data propagating inward

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.Do Python type hints stop a wrong-typed argument at runtime?

  2. 2.You need 500 LLM calls with at most 10 in flight. What do you reach for?

  3. 3.Which response should NOT be retried?

  4. 4.Why add jitter to exponential backoff?

  5. 5.What does Pydantic's `model_json_schema()` give you that matters for LLM work?