---
title: "Test your bot's answers"
description: "Create test cases with checks for facts, similarity, requirements, relevance and power-up use, then run them to catch wrong answers before users do."
canonical_url: "https://chatthing.ai/docs/bot-settings/test-your-bot"
last_updated: "2026-09-25"
---

# Test your bot's answers

**Available on:** Enterprise plans.

Build a set of test questions for your bot, each with checks its answer must pass, and run them whenever you change the prompt, model, settings or knowledge. Test runs show what passed, what failed and why.

> **Plan requirement**
>
> Bot tests are available on the Enterprise plan. On other plans, the **Tests** tab isn't shown.
>
> **Who can do this:** [team owners and admins](https://chatthing.ai/docs/account/teams#what-each-role-can-do).

## Create test cases

1. Open your bot and go to the **Tests** tab.
2. Click **Manage test cases**, then **New test case**.
3. In **Question**, enter a question a user might ask.
4. Choose a check (see the table below) and fill in its fields: usually a **Fact or statement** and a **Threshold**.
5. Click **Add check** to add more checks to the same question. A test case needs at least one check.
6. Click **Create test case**. Repeat for each question.
   ![A test case with the question "What is Chat Thing?", a Factuality check and its fact or statement, and the Add check button](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726763649/edit_test_case_85f18c8ebb.png)

### Check types

| Check | Passes when |
| --- | --- |
| **Factuality** | The answer agrees with the fact or statement you provide. |
| **Similarity** | The answer is similar enough to an example answer you provide. |
| **Requirements** | The answer meets the requirements you list, such as "mentions the 30-day return window". |
| **Relevance** | The answer is relevant to the question. |
| **Power-up called** | The bot used a specific power-up while answering. |
| **Power-up not called** | The bot did **not** use a specific power-up, for example because it should answer from its knowledge. |
| **Power-up succeeded** | The bot used a power-up and every use succeeded. |

Similarity, Requirements and Relevance checks have a **Threshold** from 0.1 to 0.9 that sets how strict the check is. The default of 0.7 is a good starting point; lower it if answers that look right are failing.

## Run the tests

1. On the **Tests** tab, click **New test run**.
2. Give the run a **Name** (it defaults to "Test #" and a number) and click **Start test**.
3. Open the run from the list to see its results.
   ![The Bot tests list showing two finished test runs with their success and failure percentages, and the Manage test cases and New test run buttons](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726763684/bot_test_runs_45ff7db818.png)

To run a single test case, click its play button on the **Test cases** page.

Test runs use message tokens: for the bot's answers, and for grading the checks, which uses an AI model.

## Read the results

Each run shows:

- The percentage of checks passing, failing and pending.
  ![Results summary for a test run: a ring chart showing 50% of checks passing, with pass and fail counts](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726763907/test_run_results_sumary_a2f94b89d8.png)
- **Duration**, number of test cases, average latency and tokens used.
  ![Test run info showing duration, number of test cases, average latency and tokens used](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726763907/test_run_info_977f6659dd.png)
- **Bot settings**: a snapshot of the prompt, model and settings used for the run, so you can compare runs after changing them.
  ![The Bot settings snapshot for a test run, showing the prompt and the model, temperature, document relevance and context settings used](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726763907/test_run_bot_settings_28f2ce3dbf.png)
- Each question with the bot's response and its checks. Click a result to see each check's **Pass** status and the **Reason** it passed or failed.
  ![Test case results listing each check's type, statement, threshold, Passed status and the reason it passed](https://res.cloudinary.com/djyjvrw5u/image/upload/v1726764094/test_case_results_b6098fefff.png)

## Check it works

Run your tests once before changing anything to get a baseline. After each change, run them again and compare the pass rate and the failed checks with the baseline.

You can also create test cases and start runs from an AI assistant with the [MCP tools](https://chatthing.ai/docs/mcp/tools).

## Troubleshooting

### Answers that look right are failing

Lower the check's **Threshold**, or reword the **Fact or statement** so it states only what matters. Open the result and read the **Reason** to see what the grader expected.

### A question fails because the bot can't find the answer

The problem is usually retrieval rather than the prompt. See [retrieval settings](https://chatthing.ai/docs/knowledge/retrieval-settings).

## Related

- [Write your bot's instructions](https://chatthing.ai/docs/bot-settings/instructions): Write the prompt that sets your bot's role, topics, tone, answer length, language, power-up use and handoff, with before-and-after examples.
- [Retrieval settings](https://chatthing.ai/docs/knowledge/retrieval-settings): Control how much of your knowledge your bot reads for each message - document relevance, context documents, enhanced retrieval, sources and chunk size.
- [Improve your bot's answers](https://chatthing.ai/docs/improve/improve-answers): Fix a bot that says "I don't know", gives wrong or made-up answers, goes off-topic, rambles, forgets context or replies in the wrong language.
