Test the questions your documents cannot answer
An assistant tested only on questions with documented answers has not shown when it will withhold an unsupported answer. Test missing evidence and unnecessary abstention separately before approving its effect on the service queue.
A service owner is considering an internal policy assistant. Its demonstration answers the questions the team selected from approved guidance, links to the right documents and looks ready to absorb work from the help desk. The unanswered question concerns the requests for which the guidance contains no answer.
An employee can ask a sensible question about an appeal deadline when the policy only describes submitting the original claim. An assistant might borrow a nearby deadline and produce a plausible instruction. The demonstration never tested that boundary. A wrong instruction can create correction work, repeated contacts and an argument about who should have supplied the missing guidance.
Before approving the service, test relevant questions that its permitted sources cannot answer. Report unsupported answers separately from supported answers and unnecessary abstentions. Then examine the human work each outcome creates, rather than treating a combined answer score as a staffing plan.
Make the missing fact plausible
Our earlier article on the evaluation set already calls for refusal cases. The extension here is to define what makes a question answerable from a particular evidence set, and deliberately change that evidence while holding the question fixed. A general instruction to include awkward questions leaves this distinction to whoever assembled the test.
The SQuAD unanswerable-question paper constructed questions that remained relevant to their passages while having no supported answer. Plausible distractors made the task harder than recognising an unrelated question. The study concerns extractive reading comprehension and the models it examined. It does not establish a current error rate for an enterprise assistant or a recommended proportion of unsupported questions.
Apply that distinction to the proposed service. Have the policy owner confirm whether each expected answer is established by the approved material. Record the source revision, effective context and requester entitlement. A source that establishes that a benefit is unavailable supports a negative answer. That case should not be labelled unanswerable simply because the employee dislikes it.
Keep missing evidence separate from access denial, prohibited requests and conflicting guidance. They may lead to different messages and different owners. An assistant must not reveal restricted material while explaining its refusal. The retrieval access-control boundary remains a prerequisite for any answerability test.
Change the evidence and keep the question
Use a constructed policy exercise before the substantive acceptance run. A fictional passage says expense claims must arrive before the next payroll cut-off. Another describes permitted hotel categories. Neither establishes an appeal deadline or a meal allowance. These are invented test materials, not advice about an employer’s policy.
Ask when to submit an expense claim. The expected response follows the claim passage. Ask when to appeal a rejected claim. The expected response states that the approved material does not establish that deadline and follows the agreed handoff. Reusing the submission deadline would answer a different question. Ask about meals and check that the hotel passage does not become its authority.
Now remove the answering passage and repeat the original claim question. The assistant should stop giving the deadline. Restore the passage and repeat again. It should resume giving the supported answer. The paired change exposes both an unsupported answer and an assistant that has learned to decline regardless of available evidence.
This exercise specifies expected behaviour. A dictionary-presence rehearsal can check that the fixture marks the intended evidence as available or absent. It cannot measure a language model, validate retrieval or establish service quality. The real acceptance run must execute the actual system, preserve its retrieved passages and have the outputs adjudicated against those passages.
Separate a missing source from a retrieval miss
Keep a corpus-level label and a supplied-context label. The permitted policy collection may contain the answer while the retrieval step fails to return it. Alternatively, retrieval may work perfectly against a collection that never contained the fact. Both situations require the assistant to withhold an unsupported instruction, but their repairs differ.
The first directs attention to retrieval, filtering or document processing. The second needs a policy decision, new approved evidence or a service handoff. Buying a larger model does not supply the missing organisational decision. Record the relevant trace so the service owner can assign the work to the right team instead of asking for an undifferentiated accuracy improvement.
Check the actual support behind an answer. The ALCE citation research measures correctness and citation quality separately. Its distinction is useful here: a citation marker does not itself establish support for every instruction in a response. Its automated measures and experimental datasets do not provide a universal acceptance threshold for internal policy advice.
Give the service owner separate outcomes
Report supported correct answers, wrong answers despite available support, unsupported answers and unnecessary abstentions. Preserve the denominator for each group. An assistant that declines everything can suppress unsupported answers while transferring almost the entire workload to the help desk.
Deliberately difficult cases are useful for finding failures. Their share of an acceptance set need not match live demand. Before estimating queue effects, measure the frequency of relevant request types in the intended service and keep that operating view separate from the stress-test view. Mark contested interpretations for review rather than manufacturing a clean expected answer.
- Baseline
- Current resolution, repeat contacts and staff effort for the same request types.
- Outcome
- Supported resolutions and appropriate handoffs, reported by evidence availability.
- Guardrails
- Unsupported instructions, unnecessary abstention, queue delay and correction work.
- Decision rule
- Approve a bounded scope only when answer quality and handoff capacity meet the service owner's agreed conditions.
Agree who receives an unresolved request, what information accompanies it and what response the employee should expect. Measure review effort, repeated contacts and time to resolution in a bounded trial alongside answer support. A well-founded refusal still consumes operational capacity when somebody must resolve the underlying question.
The service owner can approve the request types with demonstrated support and an affordable handoff, assign missing guidance to its owner, and keep unsupported uses outside the release boundary. That decision makes the assistant’s limits part of the service plan before employees discover them through a wrong answer.