Articles · Mustapha Alouani
Three checks before trusting an extraction
Clean JSON can contain the wrong information. Separate form from meaning with a concrete protocol.
The meeting in the wrong city
Your document says: “The headquarters is in Paris. The meeting will take place in Lyon.” You ask for the meeting city. The system returns an object containing Paris. Nothing crashes, the application displays a map, and the demo looks successful. It nevertheless sends you to the wrong place.
This error passes two common checks. The JSON is readable and its value is a string. The problem is the role of the selected city, not its representation.
Define the contract before the prompt
Specify the expected field, its type and what should happen when information is missing. Here, meeting_city contains a string or null, with no additional fields. null means the meeting city is absent. It does not mean the model is uncertain or the city is missing from an allowed list.
Then create a reference answer for each document. If two annotators cannot agree on the city, the task is not yet defined clearly enough for a reliable evaluation.
Measure three separate successes
- Syntax: can a parser read the output?
- Schema: do the fields and types satisfy the contract?
- Fidelity: does the answer match the information and role requested?
Report each rate with an explicit denominator. “90% correct among valid outputs” does not mean “90% of documents handled correctly”. Invalid outputs must not silently disappear from the overall result.
Build a useful small test set
Include four families: one city, several cities with different roles, no meeting city, and a city outside the usual catalogue. Vary sentence order too. Keep fresh documents for the final test after making improvements.
Constrained decoding can reduce certain format errors. It does not replace reference answers or difficult cases. Volume 2 illustrates this distinction under a specific protocol; its numbers are not universal performance claims about all LLMs.
Follow one document from start to finish
Return to our document as a reader checking the result. We are looking for a relationship: which city is associated with the meeting? “Paris” satisfies the condition “a city appearing in the text”, but that condition is too weak. Its sentence describes the headquarters. The second sentence explicitly connects the meeting with Lyon. The reference must therefore come from that relationship, not from the first city encountered.
This explains why extraction is not always a simple word search. Recognizing an entity, assigning its role and placing it in the right field are conceptually different operations. Even when one model performs them in a single generation, separating them helps us understand its errors.
Now imagine three outputs: plain text saying “Lyon”, an object whose meeting_city is Paris, and an object whose meeting_city is Lyon. The first contains the correct information but breaks the application’s interface. The second satisfies the interface but communicates the wrong information. Only the third satisfies both requirements. Operational success needs both correct content and a usable representation.
Move from an overall result to a diagnosis
In another constructed example, suppose you evaluate one hundred documents. Ten outputs cannot be parsed, twenty parse but violate the schema, and fifty of the seventy conforming outputs contain the correct city. End-to-end success is 50 out of 100. Fidelity among conforming outputs is 50 out of 70, approximately 71.4%. Both numbers are useful, but they answer different questions. The first describes what users receive; the second helps isolate semantic errors after format errors.
Improving format can therefore increase the number of usable outputs without increasing their correctness. Conversely, better understanding can remain invisible to an application that cannot parse the outputs. Select the next correction based on the dominant error category, then repeat the measurement on reserved documents.
A simple rule for difficult cases
If a document mentions two meetings in two cities, our single-value contract becomes insufficient. Specify which meeting matters, allow multiple values or signal ambiguity. Asking the model to be “more precise” does not repair an ill-defined question. Excellent extraction begins with a task that can be explained and checked without the model.
Exercise and answer
“The headquarters is in Lyon. The meeting will be in Marseille.” Is null correct because Marseille is outside your list? No. The information is present. Provide a separate state for out-of-catalogue values, or allow their extraction, instead of confusing presence with absence.