Articles · Mustapha Alouani
Temperature does not control truth
Understand what decoding changes, and what it cannot guarantee.
A dial that does not know what is true
Temperature transforms a probability distribution before a token is chosen. It consults neither a document database nor a truth judge. Lowering it generally concentrates the distribution; this can make an error more repeatable just as it can a correct answer.
Start with four constructed scores: 2, 1, 0.5 and 0.1. Divide them by a positive temperature, then apply softmax. The highest score keeps its rank, but probability gaps change. The capsule displays the calculated values so you can see this directly.
Three settings, three operations
Temperature changes concentration. Top-k retains the k highest-ranked candidates. Top-p keeps the smallest leading set whose cumulative probability reaches a threshold. After candidates are removed, the remaining probabilities must be renormalized.
Our experiment applies temperature, softmax, top-k, renormalization, top-p, then renormalization. This order is part of the protocol. Other implementations may use different conventions, so comparing only the numbers in two interfaces is insufficient.
Why one hundred draws do not exactly match probabilities
A probability of 20% does not require exactly twenty occurrences in a hundred draws. Frequencies fluctuate. The capsule uses a fixed seed for reproducible comparisons; it does not show variation across seeds. Probability bars and observed counts are different objects.
Greedy mode chooses the highest-ranked candidate. It does not calculate softmax with T = 0, which would divide by zero. The capsule handles this mode separately.
A practical protocol
For a factual task, keep a question set with reference answers. Compare settings on the same inputs, record correct answers and variations, then evaluate the chosen setting on reserved questions. A stable answer is not necessarily correct; a varied answer is not necessarily creative in a way that helps the task.
Start with scores to understand probabilities
An output score, often called a logit, is not yet a probability. It can be negative, and the scores need not add up to one. Softmax converts them into positive numbers that sum to one. For each candidate, exponentiate its score divided by temperature, then divide by the sum of those exponentials over all candidates.
With our capsule’s scores and T = 1, the first candidate receives about 57.5% probability. The second receives about 21.1%, the third 12.8% and the fourth 8.6%. At T = 0.5, the first rises to about 82.8%. No additional document has been read: we have only amplified differences between existing scores.
A precise way to understand this effect is to compare two candidates. Their probability ratio equals exp((score A − score B) / T). When A has the higher score, lowering T amplifies its advantage. This explains concentration without vaguely describing the model as “more intelligent” or “less imaginative”.
Follow a filtering operation
Now let top-k keep only the first two candidates, still with T = 1. Their original probabilities total approximately 78.6%. To obtain a new distribution summing to one, renormalize them: the first becomes about 73.1% and the second 26.9%. The other two become zero and can no longer be drawn.
Applying top-p = 0.7 next leaves only the first candidate, which already exceeds the threshold in the renormalized distribution. With top-p = 0.9, both candidates must remain. The threshold p is not a number of candidates, nor does it mean the selected token will be correct with probability p.
Connect the calculation to a complete text
In an autoregressive model, this choice repeats for every next token. The chosen token is added to the context and new scores are calculated. A small early difference can therefore change subsequent distributions. Our capsule holds the four scores fixed: it isolates a single choice to make the mechanism visible, but does not reproduce an entire generation.
This limitation suggests a useful learning method: understand one step first, then identify what repeats and what changes in the complete system. The capsule’s calculations are exact for the constructed example; conclusions about a real assistant must subsequently be tested under its own protocol.
Exercise and answer
With top-k = 1, does changing a positive temperature change the final token? Not in this capsule: the ranking stays the same and only one candidate survives. This does not mean all parameters in all libraries are interchangeable.