AI Engineer Path
Phase 1 · Prerequisites21 Sept – 27 Sept

Week 1

Python core

Read and write Python without friction.

Why this week matters

If you already program, you are learning a dialect, not programming. Skip explanations of what a loop is and focus on what is genuinely Pythonic: the data model, comprehensions, generators, decorators and context managers.

Done when

You can read an unfamiliar Python file and predict what it does.

Concepts

7 lessons · tick each one once you could explain it

In Python every value is an object with an identity, a type and a value. A variable is a name bound to an object. a = [1, 2]; b = a does not copy the list: both names now point at the same object, so b.append(3) changes what a sees too.

Objects are either mutable (list, dict, set, most class instances) or immutable (int, float, str, tuple, frozenset). Immutable objects can be dictionary keys because their hash never changes. Mutability is the source of most surprising Python bugs.

The classic trap is a mutable default argument: def f(x, acc=[]) creates the list once, when the function is defined, and every call shares it. Use acc=None and create the list inside the function.

Names are sticky notes. Assignment moves a sticky note onto an object; it never photocopies the object.

python
a = [1, 2]
b = a            # same object, two names
b.append(3)
print(a)         # [1, 2, 3]
print(a is b)    # True: identity
c = list(a)      # explicit (shallow) copy
print(c is a)    # False

Going deeper

Function arguments are passed by assignment ('call by object reference'): the parameter becomes a new name for the caller's object. Mutating it inside the function is visible outside; rebinding it (x = something_else) is not. That single rule explains every 'why did my list change?' surprise.

Small integers and some strings are cached by CPython, so a is b can be True for equal values by accident. Never rely on it. sys.getrefcount and id() let you observe identity while learning, and copy.deepcopy handles nested structures, including cycles.

Common pitfalls

  • Mutable default arguments shared across calls.
  • is checks identity, == checks equality. Use is only for None, True, False.
  • copy.copy is shallow: nested lists are still shared. Use copy.deepcopy when you need a full copy.
Backend engineer tip: Like Java object references, minus the primitives: in Python even an `int` is an object.

Best resources for this lesson

list is an ordered, mutable array. tuple is an ordered, immutable record. dict is a hash map that preserves insertion order. set is a hash set for membership tests and deduplication. Lookups in dict and set are O(1) on average; x in some_list is O(n).

Slicing seq[start:stop:step] works on any sequence and returns a new object. seq[::-1] reverses; seq[-3:] takes the last three.

A comprehension expresses map + filter declaratively: [f(x) for x in xs if cond(x)]. There are dict ({k: v for ...}) and set ({x for ...}) versions, and a generator expression (x for x in xs) that produces items lazily without building a list.

python
words = ["rag", "agent", "llm", "agent"]
lengths = {w: len(w) for w in words}          # dict comprehension
unique = {w.upper() for w in words}           # set comprehension
long_ones = [w for w in words if len(w) > 3]  # filter
total = sum(len(w) for w in words)            # generator expression, no list built

Going deeper

Know the complexity table: list append and index are O(1), insert/delete at the front are O(n); dict and set operations are O(1) on average. collections adds deque (O(1) at both ends), Counter (frequency counts), defaultdict (no missing-key checks) and namedtuple.

Comprehensions have their own scope, so the loop variable doesn't leak. Generator expressions inside function calls need no extra parentheses: sum(x * x for x in xs). For large pipelines, prefer generators to avoid materialising intermediate lists.

Common pitfalls

  • Nested comprehensions over two levels become unreadable: write a loop.
  • Using a list for membership checks inside a loop turns O(n) into O(n²); use a set.

Best resources for this lesson

Where this comes back

  • Week 5Pandas and NumPy replace most loops over data with vectorised operations.

Functions can be passed around, returned and stored like any value. That is what makes decorators, callbacks and higher-order functions natural in Python.

*args collects extra positional arguments into a tuple; **kwargs collects extra keyword arguments into a dict. The same syntax unpacks at call sites: f(*items, **options). Keyword-only parameters come after a bare *: def call(model, *, temperature=0.0) forces callers to name the argument.

Name lookup follows LEGB: local, enclosing function, global (module), built-in. A nested function that captures a variable from its enclosing function is a closure. Assigning to a captured name needs nonlocal.

python
def make_counter():
    count = 0
    def inc(step=1):
        nonlocal count
        count += step
        return count
    return inc

c = make_counter()
c(); c(5)   # -> 6

def log_call(fn, *args, **kwargs):
    print(fn.__name__, args, kwargs)
    return fn(*args, **kwargs)

Going deeper

Parameter kinds in full: positional-only (before /), positional-or-keyword, *args, keyword-only (after *), **kwargs. Library APIs use positional-only parameters to keep names private and keyword-only parameters to keep call sites readable: client.create(model, *, temperature=0, max_tokens=512).

Closures capture variables, not values, which causes the late-binding trap: [lambda: i for i in range(3)] all return 2. Fix with a default argument (lambda i=i: i) or functools.partial.

Common pitfalls

  • Forgetting nonlocal raises UnboundLocalError when you assign to a captured variable.
  • Late binding in closures: lambdas created in a loop all see the loop variable's final value.

Best resources for this lesson

A class bundles state and behaviour. Methods receive the instance explicitly as self. There is no access control: a leading underscore _name is a convention meaning 'internal'.

Dunder (double-underscore) methods let objects work with built-in syntax: __init__ (construction), __repr__ (debug display), __eq__ and __hash__ (equality, dict keys), __len__, __iter__, __getitem__ (sequence behaviour), __call__ (make instances callable), __enter__/__exit__ (with-blocks).

@dataclass generates __init__, __repr__ and __eq__ from type-annotated fields. @dataclass(frozen=True) makes instances immutable and hashable. Prefer composition over deep inheritance hierarchies.

python
from dataclasses import dataclass, field

@dataclass(frozen=True)
class Chunk:
    doc_id: str
    text: str
    score: float = 0.0

@dataclass
class Batch:
    chunks: list[Chunk] = field(default_factory=list)
    def __len__(self):
        return len(self.chunks)

Going deeper

__repr__ should be unambiguous (ideally valid code to recreate the object); __str__ is for humans. Implement __eq__ and __hash__ together, and only make objects hashable if they're effectively immutable.

Use @property for computed attributes, __slots__ to cut memory for millions of small objects, and abc.ABC or typing.Protocol to define interfaces. Dataclass options worth knowing: frozen, slots=True, kw_only=True, order=True.

Common pitfalls

  • Using a mutable default in a dataclass field: use field(default_factory=list).
  • Defining __eq__ without __hash__ makes instances unhashable.

Best resources for this lesson

Where this comes back

  • Week 2Pydantic models look like dataclasses but validate data at runtime.

Python style is EAFP ('easier to ask forgiveness than permission'): try the operation and catch the specific exception, rather than checking every precondition. Catch the narrowest exception type you can; a bare except: also swallows KeyboardInterrupt.

try / except / else / finally: else runs only if no exception occurred; finally always runs. Raise your own exceptions by subclassing Exception, and chain them with raise NewError(...) from err to keep the original traceback.

A context manager guarantees setup and teardown: with open(path) as f: closes the file even if an error occurs. You can write your own with @contextlib.contextmanager and a single yield. Use pathlib.Path for file paths.

python
from contextlib import contextmanager
from pathlib import Path
import time

@contextmanager
def timed(label):
    t0 = time.perf_counter()
    try:
        yield
    finally:
        print(f"{label}: {time.perf_counter() - t0:.3f}s")

with timed("read"):
    text = Path("notes.md").read_text(encoding="utf-8")

Going deeper

Exception groups (ExceptionGroup, except*, Python 3.11+) let concurrent code report several failures at once. asyncio.TaskGroup raises them when multiple tasks fail.

contextlib has more than @contextmanager: suppress (ignore specific exceptions), ExitStack (manage a dynamic number of contexts), closing, and asynccontextmanager for async resources such as HTTP clients and database pools.

Common pitfalls

  • Bare except: hides bugs.
  • Forgetting encoding='utf-8' when reading text on Windows.
Backend engineer tip: `with` is Python's try-with-resources.

Best resources for this lesson

Anything you can loop over is an iterable; looping calls iter() to get an iterator, then next() until StopIteration. A function containing yield becomes a generator: calling it returns an iterator, and execution pauses at each yield and resumes on the next request.

Generators make streaming pipelines memory-efficient: you can process a 50 GB file line by line, or stream tokens from an LLM as they arrive, without holding everything in memory.

yield from other_iterable delegates to another generator. itertools (islice, chain, groupby, batched in 3.12+) composes iterators cleanly.

python
def read_chunks(path, size=500):
    buf = []
    with open(path, encoding="utf-8") as f:
        for line in f:
            buf.append(line)
            if sum(map(len, buf)) >= size:
                yield "".join(buf)
                buf = []
    if buf:
        yield "".join(buf)

for chunk in read_chunks("corpus.txt"):
    ...  # one chunk in memory at a time

Going deeper

Generators can receive values (gen.send(x)) and clean up in finally when closed, the basis of coroutine-style code before async def. yield from delegates and also forwards sends and return values.

Async generators (async def with yield) stream from network sources: async for chunk in client.stream(...). Every LLM streaming API in Python is consumed this way.

Common pitfalls

  • A generator can only be consumed once.
  • Calling list() on a huge generator defeats the purpose.

Best resources for this lesson

Where this comes back

  • Week 29Streaming LLM responses are consumed as (async) generators.
  • Week 31Chunking documents for RAG is a generator pipeline.

@decorator above def f is shorthand for f = decorator(f). A decorator is a function that takes a function and returns a new function, usually a closure that calls the original.

Always wrap the inner function with @functools.wraps(fn) so the name and docstring survive. Decorators that take arguments add one more level: @retry(times=3) calls retry(times=3), which returns the actual decorator.

Built-in examples you will use constantly: @functools.cache, @property, @staticmethod, @classmethod, @dataclass. Frameworks lean on them heavily: FastAPI's @app.get, pytest's @pytest.fixture, LangChain's @tool.

python
import functools, time

def retry(times=3, delay=0.5):
    def decorator(fn):
        @functools.wraps(fn)
        def wrapper(*args, **kwargs):
            for attempt in range(1, times + 1):
                try:
                    return fn(*args, **kwargs)
                except Exception:
                    if attempt == times:
                        raise
                    time.sleep(delay * 2 ** (attempt - 1))
        return wrapper
    return decorator

@retry(times=4)
def call_model(prompt): ...

Going deeper

Decorators can be classes (implement __call__), can decorate classes (like @dataclass), and can be stacked. Order matters: the decorator closest to def is applied first.

For async functions, the wrapper must itself be async def and await the original. A common bug is a sync retry decorator wrapped around an async function, which returns an un-awaited coroutine and never retries.

Common pitfalls

  • Forgetting functools.wraps breaks introspection and debugging.
  • Retrying on every exception, including bugs; catch only the retryable ones.

Best resources for this lesson

Where this comes back

  • Week 34Agent frameworks register tools with decorators like @tool.

Practice

Hands-on work that makes the lessons stick. Warm-ups take minutes; stretch goals are optional.

  1. Warm-up

    Predict, then run

    Take 10 short snippets about mutability, defaults, closures and scope. Write down the output, then run them. Every surprise is a gap closed.

  2. Core

    A streaming log analyser

    Read a large log file with a generator pipeline (read → parse → filter → count) and report the top 10 error types with collections.Counter. Memory use must stay flat as the file grows.

  3. Stretch

    Retry and timing decorators

    Write @timed and @retry(times, exceptions) with functools.wraps, test both, then make async versions that work on async def functions.

This week, day by day

Dates follow your pace from Settings. Open the notebook icon to log hours and notes.

  1. Day 1Monday21 Sept2 h planned

    Setup: uv, VS Code, Jupyter, git repo. Python syntax, types, control flow

  2. Day 2Tuesday22 Sept2 h planned

    Data structures: list / dict / set / tuple, slicing, comprehensions

  3. Day 3Wednesday23 Sept2 h planned

    Functions, *args/**kwargs, scope, modules and packages

  4. Day 4Thursday24 Sept2 h planned

    OOP in Python: classes, dunder methods, dataclasses

  5. Day 5Friday25 Sept2 h planned

    File I/O, exceptions, context managers

  6. Day 6Saturday26 Sept3 h planned

    Iterators, generators, decorators + CampusX catch-up

  7. Day 7Sunday27 SeptReview

    Review the week, finish anything unfinished, rest

Watch

100 Days of Python ProgrammingPrimary

CampusX · playlist

Python Programming Beginner Tutorials

Corey Schafer · playlist

Short, precise videos. Good for filling specific gaps.

Python OOP Tutorials

Corey Schafer · playlist

Decorators

Corey Schafer

Generators

Corey Schafer

Learn Python: Full Course for Beginners

freeCodeCamp

Only if you are new to programming.

Read and use

Interview prep

Questions this week's material gets asked as. Answer out loud first, then open the outline.

What's the difference between `is` and `==`?
  • is compares identity (same object); == compares value via __eq__
  • Use is only for singletons like None
  • Small-int and string caching can make is look like it works by accident
Explain mutable default arguments and how to avoid the bug.
  • Defaults are evaluated once at function definition
  • A mutable default is shared across calls
  • Use None and create the object inside the function
When would you use a generator instead of a list?
  • Large or infinite sequences, streaming data
  • Constant memory; values produced lazily
  • Single pass only; can't index or reuse

Check yourself

Five questions. The done-when test above is the real bar; this is a quick self-check.

  1. 1.`b = a` where `a` is a list, then `b.append(1)`. What happens to `a`?

  2. 2.Why is `def f(items=[])` dangerous?

  3. 3.What does a function containing `yield` return when called?

  4. 4.What is `@retry(times=3)` above `def f` equivalent to?

  5. 5.Which collection gives O(1) average membership checks?