micrograd wraps a scalar in a Value that stores its data, its gradient, its parent Values and a small _backward function holding its local derivative rule. Overloading +, *, ** and tanh builds the computation graph as you write normal Python.
backward() topologically sorts the graph, sets the output's gradient to 1, and calls each node's _backward in reverse order, accumulating (+=) gradients into parents. That accumulation is the multivariable chain rule's sum over paths.
Then a Neuron, Layer and MLP on top, a training loop, and you have a working neural-network library in about 150 lines. PyTorch is the same idea over tensors.
class Value:
def __init__(self, data, _children=()):
self.data, self.grad = data, 0.0
self._prev, self._backward = set(_children), lambda: None
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data * other.data, (self, other))
def _backward():
self.grad += other.data * out.grad
other.grad += self.data * out.grad
out._backward = _backward
return out
def backward(self):
topo, seen = [], set()
def build(v):
if v not in seen:
seen.add(v); [build(c) for c in v._prev]; topo.append(v)
build(self); self.grad = 1.0
for v in reversed(topo): v._backward()Going deeper
Extending micrograd to tensors is the leap to PyTorch: each op stores a backward function over arrays, and broadcasting in the forward pass becomes summation over the broadcast dimensions in the backward pass.
Topological sorting guarantees each node's gradient is complete (all downstream contributions accumulated) before it propagates further. Without it, gradients would be partial.
Common pitfalls
- Using
=instead of+=for gradients breaks any node used twice.