In plain words
An autoencoder squeezes its input into a small code and then tries to rebuild the original from that code. Because the code is small, the network must keep only what matters most. If something rebuilds badly, it's unusual, which makes autoencoders useful for spotting anomalies.
Summarising a book in one paragraph, then trying to rewrite the book from your summary: what survives is what mattered.
In detail
An autoencoder has an encoder that maps input to a low-dimensional latent code and a decoder that reconstructs the input from it. Training minimises reconstruction error, so the bottleneck forces the code to capture the most important structure: a non-linear generalisation of PCA.
Variants: denoising autoencoders reconstruct clean input from corrupted input; sparse autoencoders penalise active units.
Uses: anomaly detection (high reconstruction error = unusual), compression, pretraining. Sparse autoencoders are now a key tool for interpreting what LLM features represent.
- the input
- the encoder, with weights
- the latent code: the compressed representation
- the decoder, with weights
- the reconstruction
- the reconstruction error
Worked example
Compressing digits and catching anomalies
- MNIST digits: 784 pixels → encoder → 32 numbers → decoder → 784 pixels, a 24.5× compression.
- Train to minimise reconstruction error between input and output.
- On normal digits, the average error is around 0.01.
- Feed it a letter 'A' it never saw: the error jumps to about 0.08, because the code has no way to represent it.
- Set the anomaly threshold at, say, the 99th percentile of errors on normal validation data.
A narrow bottleneck forces the code to capture structure; high reconstruction error flags the unfamiliar.
Common mistakes
- A bottleneck as wide as the input: the network just learns to copy and captures nothing useful.
- Trusting the anomaly threshold without checking it on labelled examples of real anomalies.
Check yourself
How is an autoencoder related to PCA?Show answer
A linear autoencoder with squared error learns the same subspace as PCA. Non-linear layers let it capture curved structure PCA can't.
What's a denoising autoencoder?Show answer
One trained to reconstruct the clean input from a corrupted version, which forces it to learn robust structure rather than copy pixels.
Going deeper
Sparse autoencoders trained on an LLM's internal activations decompose them into many interpretable 'features' (concepts the model represents), now a central tool in mechanistic interpretability.