The laboratory
The classifier taking a shortcut
A perfect score… until “thanks” switches sides. Find the shortcut.
Constructed experiment · no model running
The situation · Understand the request
Your team needs to identify incidents in incoming messages. In a demo, the detector recognizes every failure. You suspect it may instead rely on the politeness phrase added to the examples.
What you will learn
Distinguish success on a dataset from the ability to recognize the problem that matters.
Your mission
Keep the meaning of all four messages and change only the presence of “thanks”. Predict what will happen to the score, then test your hypothesis.
By the end of this activity. You will be able to propose a paired test for nuisance-cue dependence without claiming to have explained the entire model.
Your turn
Toy rule: if a message contains “thanks”, predict an incident. The rule is fixed; no training is simulated.
| Message | Reference | Prediction |
|---|
Politeness is not the business problem. A score measures a specific dataset, not a universal skill. Reversing this cue is a constructed stress test, not an estimate of real-world performance.
Separate what we want to recognize from what we observe
The task is to recognize an incident from a message’s meaning. A server outage and a failed payment are incidents with or without a politeness phrase. The Reference column represents this task definition and remains unchanged when you toggle the switch.
The prediction rule does not read the described problem. It only checks whether “thanks” has been added. Initially, we place the phrase on both incidents. The rule appears perfect because two different properties coincide: being an incident and containing “thanks”.
What the reversal demonstrates
Moving the phrase to normal messages breaks that coincidence. Neither incident triggers an alert, while both ordinary messages do. The score falls to zero without changing the rule. What changes is the relationship between the cue and the correct answer.
Calling this “cheating” is a playful description, not a claim about intention. We wrote a mechanical rule whose behavior is fully known. Finding an analogous sensitivity in a learned model would require testing; explaining its internal mechanism would require additional observations.
Keep this separation in mind when examining your results. An experiment can show that a system depends on a nuisance cue under certain conditions without showing that it depends only on that cue in every context. A good explanation states that boundary precisely.
The trap
You have four messages and a seemingly unbeatable rule. Before checking the box, predict its score after the reversal. Incidents remain incidents; only the politeness cue moves.
What changes, what stays fixed
The task content and reference labels stay fixed. We move “thanks” from incidents to normal messages. The rule only sees this cue, so it goes from 4 correct answers out of 4 to zero. This is a constructed teaching example, not a measurement of a trained model.
Try it in your project
Create pairs of messages with the same meaning but different politeness, signatures or layouts. A changed prediction reveals a sensitivity worth investigating. Then check fresh examples: repairing a known test is not evidence of generalization.