How to test an AI support knowledge base: separate retrieval from answers
A repeatable acceptance test for checking whether a support knowledge base retrieves the right evidence and whether the AI turns it into complete, faithful answers.
When a knowledge-base support bot gives a wrong answer, replacing the model is an easy first reaction. But the failure may have happened at either of two stages: the system did not retrieve the right material, or it retrieved useful evidence and then omitted, misunderstood, or embellished it while composing the answer.
If those stages are collapsed into one overall “answer quality” score, debugging becomes guesswork. A better acceptance test inspects retrieval first and the final answer second. Once the failing stage is visible, the team can decide whether to change the source documents, retrieval configuration, instructions, or model.
Why retrieval and generation need separate tests
Retrieval-augmented generation, or RAG, normally finds relevant passages in a knowledge base before sending those passages and the user's question to a language model. Microsoft's RAG architecture guidance treats search-result evaluation and end-to-end language-model evaluation as consecutive stages. It recommends establishing that relevant grounding evidence is returned before evaluating the generated response.
That distinction is especially useful in customer support:
- The evidence was never retrieved. The correct policy may exist in the help center but never reach the model.
- The evidence was only partially retrieved. The bot may see the refund conditions but miss an exception or deadline.
- The evidence was complete, but the answer was wrong. The model may confuse the subject, number, or condition.
- The evidence was insufficient, but the answer sounded fluent. The model may fill the gap with general knowledge or a guess. This is often the riskiest outcome.
A correct final answer therefore does not prove that retrieval is reliable. One lucky answer is not a substitute for a repeatable test.
Step 1: build questions with expected evidence
Do not store only a model answer. For every test question, record at least:
- A realistic customer phrasing, including abbreviations, conversational wording, and common misspellings.
- The document, section, or passage that should support the answer.
- The applicable version, region, plan, and effective date.
- The expected conclusion, including conditions and limitations that must appear.
- What the bot should do when the knowledge base cannot answer: decline to guess, request missing details, or hand off to a person.
Amazon Bedrock's RAG evaluation documentation similarly describes datasets that contain queries plus expected retrieved text or responses so the system can be compared with ground truth. A small team does not need an automated evaluation platform to begin; a well-structured spreadsheet can serve as the first benchmark set.
Cover more than the most common FAQ entries:
- Frequent single-fact questions, such as business hours or basic plan differences.
- Paraphrases of the same intent, to expose dependence on exact keywords.
- Questions that require evidence from two sections.
- Questions constrained by date, region, account type, or order status.
- Questions the knowledge base explicitly cannot answer.
- Cases where old and current policies coexist but only one is valid.
Negative cases matter. Microsoft's retrieval guidance recommends testing both positive and negative examples: answerable questions should retrieve the right material, while unanswerable questions should not return passages that merely look similar.
Step 2: inspect retrieval before judging the prose
Save the top retrieved passages for each question and temporarily hide the generated answer. Check:
- Relevance: Does the passage actually address the question, or does it only share vocabulary?
- Coverage: Are all material conditions present, or only one part of the answer?
- Ranking: Is the most useful source near the top, or is it displaced by secondary or outdated content?
- Scope: Does the evidence apply to the correct product, plan, region, and version?
- Conflict: When old and new policies appear together, do dates or status fields make the valid one distinguishable?
AWS summarizes common retrieval checks as context relevance and context coverage. Microsoft also documents methods such as Precision@K, Recall@K, and mean reciprocal rank. A manual acceptance test does not have to compute every formula, but it should preserve the same reasoning: how much of the returned material is truly relevant, whether required evidence is missing, and where the first useful result appears.
If the correct passage is absent, first confirm that the content exists, has been synchronized, and uses clear headings, terms, chunks, and metadata. If it appears far down the ranking, investigate the search method, filters, reranking, or the number of passages returned. Rewriting the answer prompt at this point usually treats the symptom rather than the cause.
Step 3: freeze the evidence, then inspect the answer
Once the evidence is usable, evaluate how the language model applies it. AWS separates this stage into dimensions including correctness, completeness, helpfulness, logical coherence, faithfulness, citation precision, and citation coverage.
For a support acceptance test, turn those dimensions into six questions:
- Is the conclusion correct, including names, numbers, dates, and responsible parties?
- Does the response address every part of the customer's question?
- Does it provide a clear next step instead of merely repeating the source?
- Are the conditions, exceptions, and steps internally consistent?
- Can every factual claim be traced to the retrieved passages?
- Does each citation or link actually support the nearby claim?
Faithfulness to the retrieved text and factual correctness are not interchangeable. Microsoft's end-to-end evaluation guidance notes that a response can follow its context while the source itself is outdated, or it can draw an incorrect conclusion from valid material. The test record should therefore preserve both the source version and the business-approved conclusion.
Use four outcomes to find the failing layer
Classifying each test into four outcomes is usually more actionable than producing one aggregate score:
- Good retrieval, good answer: keep the case as a regression benchmark.
- Bad retrieval, bad answer: fix knowledge coverage, synchronization, chunking, metadata, or search configuration first.
- Good retrieval, bad answer: inspect the instructions, answer template, model choice, and citation constraints.
- Bad retrieval, apparently correct answer: mark it high risk. The answer may come from model memory or coincidence and can break as soon as the knowledge base changes.
After a fix, rerunning only one or two questions is not enough. A chunking or search change can improve one case while degrading another. Rerun the complete benchmark and store the question, retrieved passages, final answer, configuration version, and human rating.
A minimum pre-launch checklist
Before connecting a knowledge base to live support traffic, verify that:
- Test questions come from real support work rather than marketing copy.
- Every question has expected evidence, scope, and an explicit unanswerable behavior.
- Retrieval and answers are scored separately, so a failure can be assigned to a stage.
- Obsolete material has been removed, archived, or clearly marked with an end date.
- Unanswerable, conflicting, and underspecified cases have been tested.
- High-impact topics still have a human review or handoff path.
- The regression set is rerun after document, model, prompt, or retrieval changes.
YourCopilot is designed to answer Telegram private chats, group chats, and Telegram Business messages using a configured knowledge base. This article describes a manual acceptance workflow for knowledge-base owners; it does not claim that the product includes the automated metrics or evaluation platforms discussed above. Check the current product page and management interface for available capabilities.
Sources
- Microsoft Azure Architecture Center, Develop a RAG Solution on Azure - Information-Retrieval Phase, last updated July 2, 2026.
- Microsoft Azure Architecture Center, Develop a RAG Solution on Azure - Large Language Model End-to-End Evaluation Phase, last updated November 21, 2025.
- Amazon Web Services, Use metrics to understand RAG system performance, continuously maintained documentation, accessed October 4, 2026.