Articles · Mustapha Alouani
Precision and recall: choose what to protect
A needless alert and a missed incident may have very different costs.
The detector that always raises an alarm
A service receives one hundred messages, ten of which report an incident. A detector raises an alert for every message. It finds all ten incidents: recall is 100%. But only ten of its hundred alerts are justified: precision is 10%. These are constructed numbers, not results from the book.
Recall alone gives it an excellent grade. Precision alone tells another story. Before choosing a metric, identify the errors your application needs to avoid.
Reconstruct the counts
A true positive is a correctly flagged incident. A false positive is an alert on a normal message. A false negative is a missed incident. A true negative is a normal message correctly left alone.
Precision divides true positives by all alerts. Recall divides them by all actual incidents. F1 combines precision and recall using a harmonic mean; it does not account for true negatives or an explicit business cost.
Why overall accuracy can mislead
On the same batch, never raising an alert gives 90% correct answers but detects no incidents. That number can look reassuring if the rarity of incidents is ignored. Keep counts and class proportions alongside every percentage.
Observed precision also depends on the population evaluated. An artificially balanced dataset has a different incident prevalence from the real stream. Do not mechanically transfer its numbers to production.
Choose a threshold without cheating
If the system produces a score, compare thresholds on a calibration set. Specify the criterion: human capacity to inspect alerts, the cost of missed incidents or a minimum recall. Once the threshold is selected, use a reserved test set to measure performance. Do not retune the threshold on that test.
With a small sample, a few cases can change the result substantially. Report counts and, where the protocol supports it, an uncertainty estimate. A two-point difference is not automatically robust progress.
Read a confusion matrix as a story
Metric vocabulary becomes easier when connected to a concrete situation. Imagine opening each alert and asking, “Was there really an incident?” You are viewing the system from the perspective of the person receiving alerts. That is the perspective of precision. Now imagine starting with the list of all actual incidents and checking which ones triggered an alert. That is the perspective of recall.
Both readings use the same decisions but begin with different populations. This starting point explains their denominators. It also explains why high precision is insufficient: a system can issue just one correct alert while missing nine other incidents. Its precision is then 100%, but recall is 10%.
Understand what F1 keeps and what it hides
F1 is useful when we want to summarize the balance between precision and recall. The harmonic mean penalizes a large imbalance. With 100% precision and 10% recall, the arithmetic mean would be 55%, whereas F1 is about 18.2%. The latter makes the fact that most incidents are missed more visible.
This summary contains no built-in business judgment. Two systems with the same F1 can produce different alert volumes, and the same score may be acceptable for exploration but inadequate in a process where every mistake is expensive. Before comparing systems, specify what predictions will be used for and who will handle their consequences.
Connect the metric to incident prevalence
Consider a fictional detector that finds 80% of incidents and falsely alerts on 10% of normal messages. In one hundred messages containing ten incidents, it produces an expected eight justified alerts and nine false alerts. Expected precision is therefore 8/17, about 47.1%. In one hundred messages containing fifty incidents, the same conditional rates yield forty justified alerts and five false alerts, or about 88.9% precision.
The detector has not changed in this thought experiment. Incident prevalence has changed. This is why precision measured on a balanced corpus cannot be uncritically interpreted as precision on a real stream. These values illustrate the calculation; in a real system, the conditional rates themselves can also change as the data shifts.
Make an explainable decision
A defensible choice links the metric to a concrete constraint. For example, seek the highest recall compatible with reviewing twenty alerts each day. This makes the trade-off visible and open to discussion. It is more informative than simply asking for higher F1 because it connects the calculation to the work of the people who will use the system.
Exercise and answer
With 8 true positives, 2 false positives and 2 false negatives, both metrics are 80%. Double false negatives and change nothing else. Precision stays at 80%; recall becomes 8/12, about 66.7%. Use the capsule to verify the calculation.