Articles · Mustapha Alouani

A perfect score deserves an investigation

Build a test that distinguishes useful ability from an accidental shortcut.

The easy cue

Imagine a corpus where every incident ends with “thanks” and ordinary messages do not. A rule based only on that word scores 100%. It knows nothing about incidents. This corpus and rule are constructed to expose the problem; this is not a measurement of a real model.

A shortcut emerges when a cue correlated with the target is easier to exploit than the phenomenon we wanted to learn. A signature, department name, template or text length can play this role.

Change a detail without changing the answer

Create two versions of a message. Keep its business meaning but add or remove the politeness phrase. If the reference remains “incident” while the prediction changes, you have identified a sensitivity. You have not proved its internal mechanism: this is a behavioral observation.

Counterexamples must remain plausible. A transformation that removes an essential word can change the meaning and therefore the correct answer. Review ambiguous pairs before concluding that the system is at fault.

Separate dataset roles

  • Training: fit the model.
  • Development: analyze errors and change the method.
  • Calibration: set a threshold or abstention rule when needed.
  • Test: evaluate an already fixed method on reserved examples.

After repeated fixes inspired by a test set, that set effectively becomes development data. It remains useful but is no longer independent evidence of generalization.

Keep an error table

For each error, record the text, reference, prediction, perturbation family and a short hypothesis. Show counts by family rather than only an overall average. Ten paraphrases of one message do not necessarily represent ten independent situations.

Understand what a score allows us to say

A percentage does not summarize a model alone. It summarizes the interaction between a decision rule, a set of examples and a counting procedure. Changing any one of these can change the result. Saying “my model is 95% reliable” without describing the task and data therefore omits much of the useful information.

Our rule succeeds because the evaluation set reproduces an association between politeness and incidents. That success is real on these four messages: the arithmetic is correct. What would be wrong is inferring that the rule recognizes a server failure or a failed payment. The error lies in interpreting the result, not counting correct answers.

Imagine a student memorizing the positions of correct answers on a quiz. The student can score well as long as answer order stays fixed. Shuffling the answers tests the memorization hypothesis. This analogy clarifies the experiment; it does not imply that a model has the intentions or awareness of a cheating student.

Build an investigation, not a collection of traps

Start with a testable hypothesis: “The politeness phrase influences the decision even when the business meaning stays the same.” Create several pairs that isolate this transformation. Include messages where changing politeness should not cause difficulty, rather than selecting only spectacular failures.

Record predictions separately before and after transformation. A high change rate deserves attention, but inspect the direction of those changes. Some transformations may correct an initial error; others may introduce one. The number of different answers alone does not tell you whether the system improves or deteriorates.

Keep a record of the choices made during the investigation. If you explored ten perturbations and publish only the one producing the largest difference, readers need to know. Such selection can help discover a flaw, but it does not by itself estimate representative real-world risk.

What a repair must demonstrate

A convincing correction should succeed on fresh cases, preserve previous successes and avoid shifting the problem to another message family. Present improvements and regressions together. The lesson is not that scores are useless: it is that we must report the conditions that make them meaningful.

Exercise and answer

Your score returns to 100% after adding the four counterexamples to training. Can you announce that the shortcut is gone? No. You have shown success on now-familiar cases. Reserve new messages and formulations, then check whether the improvement holds.

Open the interactive experiment