The laboratory

The right citation, the wrong answer

Investigate a fictional shop: which policy actually supports an answer to the customer?

Constructed experiment · no model running

The situation · Find the applicable policy

A customer wants to return a personalized keyboard received twelve days ago. The assistant finds a page offering thirty-day returns. Is that enough to accept the request?

What you will learn

Distinguish an apparently relevant document from sufficient evidence for a specific situation.

Your mission

Read the documents, select the evidence supporting your decision, then answer the customer. Repeat when the exception policy becomes unavailable.

By the end of this activity. You will be able to explain why a genuine citation does not guarantee a correct answer and identify when verification is necessary.

The customer has changed their mind and requests a return. The product has no defect. What can you say using the available documents?

Documents proposed by retrieval

Select the documents that support your answer.

These shop policies are fictional. Document order is constructed for the exercise, not measured from a search engine.

Your answer to the customer

This experiment assesses your evidence selection in a constructed corpus. No LLM generates answers and no real search engine is evaluated.

Why start with a customer-service decision

A document assistant must do more than retrieve a passage resembling a question. It must connect that passage to the facts of the case. Here, the twelve-day delay attracts attention because it fits D1’s thirty-day window. But product category also matters: the customer has a personalized keyboard, whereas D1 only describes standard items.

Your task is to distinguish these conditions. The general policy is not false; it is insufficient to decide this case. Quoting an authentic sentence cannot repair reasoning that applies it outside its scope.

How to investigate

In the first case, read D1 and D2 before answering. D2 directly addresses personalized items and changes of mind. There is no defect, so the separate provision for defects does not alter this exercise’s decision. Select the document that establishes your conclusion. D3 contains “personalized keyboard” but concerns preparation and delivery, not the fictional shop’s return policy.

Then switch to the standard-keyboard case. The time since receipt stays the same, but product category changes. D1 becomes applicable. This variation shows that a good answer is not a memorized sentence: it depends on matching case facts to document conditions.

When the decisive evidence is missing

The third case removes D2 from the accessible documents. As a reader who has seen the first case, you know an exception exists. The assistant must nevertheless justify its answer using only the currently available corpus. Inspect D1 and D3 and choose verification rather than an unsupported decision. Missing evidence proves neither permission nor prohibition.

The feedback distinguishes the chosen answer from its sources. This avoids rewarding a correct decision supported by an irrelevant justification. In a real system, the distinction can reveal answers that are correct by chance or based on information outside the requested corpus.

The connection with RAG

RAG combines document retrieval with answer generation. Here we isolate the transition from retrieved documents to a decision: retrieval is represented by a constructed list and you act as the reader checking evidence. No similarity scores or model results are invented.

To apply this method, prepare questions where the first source looks relevant but a condition changes the answer. Evaluate separately whether the right source was retrieved, whether the answer is correct and whether the citation supports it. These three successes are not equivalent.