Concept Lab · Week 26 · AI Engineer Track
Temperature and top-p
Reshape a next-token distribution and sample from it.
The idea: Decoding: temperature, top-p, top-k
At each step the model emits logits over the vocabulary. Temperature divides the logits before softmax: below 1 sharpens the distribution (more deterministic), above 1 flattens it (more random); 0 means greedy. Top-k keeps only the k most likely tokens; top-p (nucleus) keeps the smallest set whose probabilities sum to p.
Guidance: low temperature (0–0.3) for extraction, classification, code and anything you'll parse; moderate (0.7–1.0) for creative writing and brainstorming. Change temperature or top-p, not both at once.
Even at temperature 0, outputs are not perfectly deterministic across calls (batching and floating-point effects). Design for variation: validate outputs, don't assume byte-identical repeats. Reshape a distribution yourself in the simulation.
Next simulation: Prefill, decode and the KV cache