The safety net answered 32 of 150 questions from the wrong document
If the search found nothing in the right document, my system searched everything, so a user always got an answer. That fallback turned a visible failure into a confident wrong answer.
The safety net I built into my AI system answered 32 of 150 test questions from the wrong document.
Why it matters
The system answers questions from 888 technical documents, with citations. I had added a fallback: if the search found nothing in the right document, it searched everything, so users always got an answer.
That fallback turned a visible failure into a confident wrong answer.
In a company, the wrong document is another client’s contract. The answer still reads well and still cites a source, which is exactly what makes it easy to trust.
My scores showed the cost. In the same run, wrong-document answers scored 0.125 on correctness; right-document answers scored 0.468. In the 4 contaminated answers I reviewed with Claude Opus 5, the confidence check stayed quiet.
A second flaw skewed the average: I tested questions in dataset order, and the first 15 came from 6 documents, 7 of them from a single one.
Prompt tuning would not have fixed this. The problem was retrieval.
How it works
The old search ranked passages across all 888 documents, kept the target document’s survivors, and searched unfiltered when none survived. Filtering after ranking is common in vector search, and this is the point where it breaks: by the time the filter runs, the ranking has already spent its slots on the rest of the corpus.
The fix restricts the dense search (a FAISS ID selector) and the keyword search (BM25) to the target document before ranking. The fallback is gone: an empty result is recorded as a failure rather than answered from somewhere else. Questions are now sampled with a fixed seed across the full dataset, so the mix no longer depends on where the file happened to start.
Rerun correctness: 0.47, level with the right-document answers. The 32 contaminated answers were not a generation problem that better prompting could reach.
Two limits remain, and both are mine to close. The evaluation hands the system the right document, so choosing that document is still untested. An AI judge produces the scores, and I have not yet validated it against human labels.