The laboratory
Temperature: changing probabilities
The same scores, different probabilities. Explore temperature, top-k and top-p with four tokens.
Constructed experiment · no model running
The situation · Understand variation in answers
Your assistant sometimes phrases answers differently to the same request. You consider lowering temperature to make it more reliable. First, you want to understand what this setting actually changes.
What you will learn
Distinguish the probability of selecting a token from the probability that an answer is true.
Your mission
Use four candidates as a magnifying glass on a single choice. Change temperature and filters, then compare expected probabilities with observed draws.
By the end of this activity. You will be able to explain temperature, top-k and top-p, including the limits of a demonstration that does not run a complete model.
Your turn to experiment
Greedy mode selects the highest score. Temperature and filters are disabled.
The scores stay the same. You change how they become probabilities and how they are filtered.
Constructed scores: 2 · 1 · 0.5 · 0.1. The French words are fixed labels for this experiment.
Show the values
| Token | Before filtering | After filtering |
|---|---|---|
| chat | 57.5 % | 57.5 % |
| chien | 21.1 % | 21.1 % |
| souris | 12.8 % | 12.8 % |
| oiseau | 8.6 % | 8.6 % |
Fixed seed 42; observed frequencies are not the exact probabilities.
Understand the values before changing the settings
The four starting numbers, 2, 1, 0.5 and 0.1, are constructed scores. They rank candidates but are not yet probabilities. Softmax transforms them into positive values that sum to one. This makes it possible to draw a candidate randomly according to the resulting distribution.
At temperature 1, “chat” receives about 57.5% probability. Across many independent draws from exactly this distribution, its frequency would tend toward that value. This does not mean the word is true or answers a question well. The capsule contains no factual question against which relevance could be judged.
A guided experiment in three stages
Keep all four candidates and set top-p to 1. Lower temperature from 1 to 0.5 and observe the first bar. The already favored candidate gains probability because dividing by a smaller temperature amplifies score differences. The ranking stays the same; concentration changes.
Return to the initial temperature, then set top-k to 2. Two candidates are excluded. The survivors’ probabilities are recalculated so they still sum to one. Finally, lower top-p: once the first candidate’s cumulative probability reaches the threshold, the second is removed too.
Read the draws without confusing randomness with error
The draw button converts the distribution into one hundred observations. Counts can differ from expected percentages because finite samples fluctuate. The capsule restarts with the same pseudorandom seed to support comparison. A study of variability would use several seeds and more repetitions.
Before reading the conclusion, make a prediction: with top-k equal to 1, can another candidate appear? No, because its final probability is exactly zero. That answer follows from filtering, not from the quality of a model’s knowledge.
Remember
A low temperature concentrates probability on the highest scores. It provides no guarantee of truth.
Order of operations
Temperature → stable softmax → top-k → renormalization → top-p → final renormalization. Top-p keeps the smallest prefix reaching the threshold; ties follow the initial token order.
p(i) = exp((z(i) − max(z)) / T) / Σ exp((z(j) − max(z)) / T)
What this example does not show
Four constructed scores do not represent a full vocabulary. This capsule does not generate text or run a model. Libraries may use different filtering conventions.