The laboratory

The overenthusiastic detector

Catch more cases at the cost of false alarms? Explore precision, recall and F1.

Constructed experiment · no model running

The situation · Evaluate the service provided

Your team receives the assistant’s alerts. Some identify genuine incidents; others waste time. Some incidents are missed altogether. What evaluation would help decide whether the tool is useful?

What you will learn

Read precision from the perspective of received alerts and recall from the perspective of actual incidents.

Your mission

Start with eight detected incidents. Increase false alarms, then missed incidents. Describe the effect on the team’s work before looking at percentages.

By the end of this activity. You will be able to calculate and interpret the three metrics without confusing a summary score with a business decision.

Your turn

Challenge: keep 8 true positives and 2 missed incidents. Increase false alarms from 2 to 8. Which measure stays the same?

Precision = TP / (TP + FP). Recall = TP / (TP + FN). F1 = 2TP / (2TP + FP + FN). A zero denominator is shown as “undefined”. True negatives are not used here.

Explanation

Recall stays at 80%. Precision drops from 80% to 50%, and F1 from 80% to about 61.5%. One metric alone cannot select a system.

Read the counts from the user’s perspective

Imagine reviewing alerts from a monitoring tool. With the initial values, you receive ten alerts: eight correspond to real incidents and two are unnecessary. Precision is therefore eight out of ten. To calculate recall, leave the alert inbox and examine all actual incidents: there were ten, including two that were missed. Recall is also eight out of ten, but for a different reason.

This initial equality can make the metrics seem interchangeable. The experiment is designed to dispel that impression. Add six false alarms: you now review sixteen alerts to find the same eight incidents. Your verification workload grows and precision falls to 50%. Yet no additional incidents have been missed, so recall stays at 80%.

Why a metric cannot decide for you

F1 combines the two ratios and penalizes a strong imbalance. It does not know your organization. If a false alarm costs one minute and a missed incident causes hours of downtime, system selection should account for that difference. Metrics describe decisions; consequences provide the context in which to interpret them.

Freely adjusting the counts explores properties of the formulas. It does not demonstrate that a real classifier can attain every combination. In a real system, moving a threshold generally changes several counts together. Measure those changes on representative examples instead of independently choosing convenient numbers.

Two different questions

Precision asks: of the alerts raised, how many were justified? Recall asks: of the actual incidents, how many did we find? Their denominators are different.

A thought experiment

With 8 detected incidents, 2 false alarms and 2 missed incidents, precision and recall both equal 80%. Increasing only false alarms reduces precision without changing recall. Increasing only missed incidents does the opposite.

A calculation is not a policy

These counts are freely adjustable: the capsule does not model the threshold curve of a real classifier. To choose a threshold, measure predictions on a calibration set, then evaluate the fixed choice on a separate test set. F1 represents neither error costs nor the frequency of all negative cases.

Bonus challenge

Set all counts to zero. An undefined score is not zero performance: there are insufficient observations to calculate the ratio.