Log in

Test your bot's answers

Build a set of test questions for your bot, each with checks its answer must pass, and run them whenever you change the prompt, model, settings or knowledge. Test runs show what passed, what failed and why.

ℹ️

Plan requirement

Bot tests are available on the Enterprise plan. On other plans, the Tests tab isn't shown.

Who can do this: team owners and admins.

Create test cases

  1. Open your bot and go to the Tests tab.
  2. Click Manage test cases, then New test case.
  3. In Question, enter a question a user might ask.
  4. Choose a check (see the table below) and fill in its fields: usually a Fact or statement and a Threshold.
  5. Click Add check to add more checks to the same question. A test case needs at least one check.
  6. Click Create test case. Repeat for each question.
    A test case with the question "What is Chat Thing?", a Factuality check and its fact or statement, and the Add check button

Check types

CheckPasses when
FactualityThe answer agrees with the fact or statement you provide.
SimilarityThe answer is similar enough to an example answer you provide.
RequirementsThe answer meets the requirements you list, such as "mentions the 30-day return window".
RelevanceThe answer is relevant to the question.
Power-up calledThe bot used a specific power-up while answering.
Power-up not calledThe bot did not use a specific power-up, for example because it should answer from its knowledge.
Power-up succeededThe bot used a power-up and every use succeeded.

Similarity, Requirements and Relevance checks have a Threshold from 0.1 to 0.9 that sets how strict the check is. The default of 0.7 is a good starting point; lower it if answers that look right are failing.

Run the tests

  1. On the Tests tab, click New test run.
  2. Give the run a Name (it defaults to "Test #" and a number) and click Start test.
  3. Open the run from the list to see its results.
    The Bot tests list showing two finished test runs with their success and failure percentages, and the Manage test cases and New test run buttons

To run a single test case, click its play button on the Test cases page.

Test runs use message tokens: for the bot's answers, and for grading the checks, which uses an AI model.

Read the results

Each run shows:

  • The percentage of checks passing, failing and pending.
    Results summary for a test run: a ring chart showing 50% of checks passing, with pass and fail counts
  • Duration, number of test cases, average latency and tokens used.
    Test run info showing duration, number of test cases, average latency and tokens used
  • Bot settings: a snapshot of the prompt, model and settings used for the run, so you can compare runs after changing them.
    The Bot settings snapshot for a test run, showing the prompt and the model, temperature, document relevance and context settings used
  • Each question with the bot's response and its checks. Click a result to see each check's Pass status and the Reason it passed or failed.
    Test case results listing each check's type, statement, threshold, Passed status and the reason it passed

Check it works

Run your tests once before changing anything to get a baseline. After each change, run them again and compare the pass rate and the failed checks with the baseline.

You can also create test cases and start runs from an AI assistant with the MCP tools.

Troubleshooting

Answers that look right are failing

Lower the check's Threshold, or reword the Fact or statement so it states only what matters. Open the result and read the Reason to see what the grader expected.

A question fails because the bot can't find the answer

The problem is usually retrieval rather than the prompt. See retrieval settings.

  • Write your bot's instructions

    Write the prompt that sets your bot's role, topics, tone, answer length, language, power-up use and handoff, with before-and-after examples.

  • Retrieval settings

    Control how much of your knowledge your bot reads for each message - document relevance, context documents, enhanced retrieval, sources and chunk size.

  • Improve your bot's answers

    Fix a bot that says "I don't know", gives wrong or made-up answers, goes off-topic, rambles, forgets context or replies in the wrong language.

Last updated