THE LEARNING PATH · Intermediate
The LLM garage: patch an activation to understand
Is information visible inside a model actually used? Build a contrast, measure restoration and learn how to test an explanation.
The question: a useful part, or merely a witness?
Imagine two identical cars: one starts, the other does not. You move a part from the first into the second. If it now starts, you have learned through intervention, beyond observation alone. Chapter 9 of Volume II uses this analogy to introduce activation patching. It guides the operation without implying that a neural network consists of independent mechanical parts.
Our objective is specific: understand how to test the role of an internal state in an output preference. First distinguish weights, retained between requests, from activations, values computed for a particular input. We replace an activation during a run; we do not retrain the weights.
Build two comparable computations
Consider “James and Sarah go to the store. James gives a book to…”. The expected recipient is Sarah. If the second subject becomes Sarah, James becomes the expected recipient. The clean condition supplies a donor state; the corrupted condition receives it. These names describe the protocol, not the linguistic quality of the sentences.
Run both inputs. At a chosen site — a layer and a token position — save the clean state. Then rerun the corrupted computation, replacing only that site’s state with the saved state. Subsequent layers continue computing. Positions must align in the actual tokenization: counting words is not enough.
Choose the measure before inspecting the result
A logit is a score before conversion into probabilities. Measure the difference between the logits of the two candidates at the same position. A positive difference favors the first; a negative difference favors the second. This measures a defined contrast, not the model’s overall quality.
Use the book’s invented teaching example: −4 in the corrupted condition, +3 in the clean condition and +1.5 after patching. The reference gap is 7. The intervention recovers 5.5 units of that gap: restoration is 5.5 / 7, approximately 0.79. It is not a 79% probability of truth.
Three checks that make the result meaningful
- Self-patching puts the corrupted state back in its own place. The score should remain unchanged within numerical tolerance. Otherwise, fix the setup first.
- Norm-matched control perturbations test whether any displacement of the same size could produce the effect. They help distinguish the intervention’s direction from its strength alone.
- Reverse patching transfers the corrupted state into the clean computation. It complements the contrast without requiring symmetry in a nonlinear system.
Site selection must also be separate from validation. If you explore twenty sites and report only the best on the same sentences, you have mainly found the winner of your search. Choose using discovery examples, then check different examples.
What you can now conclude
Restoration supported by controls provides evidence for a causal role of the specified intervention under the tested conditions. It does not show that the replaced vector contains only one concept, or that the complete circuit has been identified. A missing effect can also depend on the site, the measure or redundant pathways.
In the calculator below, find no effect, full restoration and overshoot in turn. Explain each case using the formula before opening the answer. This tool computes an exact ratio from selected values; it does not run an LLM.
Your turn: restore a preference
Two runs provide reference points: −4 for the corrupted condition and +3 for the clean condition. Move the post-patching result. The slider changes a numerical example, not the activations of a real model.
(patched − corrupted) / (clean − corrupted)
Normalized restoration :
Patching recovers part of the gap between the two conditions. This percentage is neither the probability of a correct answer nor a share of “understanding”.
A site restores 80% of the gap. Have we found the complete mechanism? — See the reasoning
No. We have established an effect of this intervention on this measure. Check the setup with self-patching, compare control perturbations and test examples not used to select the site. Replacing a whole vector can move several kinds of information at once.