How can we help? 👋

Digital AI (LLM Chatbot) Test Runs

Test runs let you check how your Digital AI (LLM Chatbot) handles a set of customer scenarios before real customers do. You describe a scenario, the AI plays that customer and chats with your bot, and each test gets a grade for how well the bot handled it.

It's a safe way to test changes. You can run tests against a draft version of your bot, so you can see how an edit to your prompt or tools performs before you publish it.


🧪 What is a test run?

A test run is a batch of scenarios run against your bot in one go. Each scenario is a persona: a made-up customer with a situation and a goal. When you start the run, the AI acts as each persona, has a full conversation with your bot, and grades how well the bot did.

The three grades

Each test gets one of three grades.

Grade
What it usually means
What to do
Good
The bot handled the scenario well. It responded accurately and did what the customer needed.
Nothing. This is what you want to see.
Warning
Mostly fine, but something minor is worth a look, such as a slightly off or incomplete answer.
Open the test and check what happened.
Needs Attention
The bot likely got something wrong or handled the scenario poorly.
Review it, fix your prompt or Tools, and run the test again.

🛠️ Setting up a test run

You build a test run from a set of scenarios, then start it.

1. Open the Testing tab

Open the LLM Chatbot you want to test and go to the Testing tab, next to Prompt, Tools and Evaluations. Click New test run.

Notion image

2. Add your scenarios

Each scenario is a customer for the AI to play. You can add them two ways:

  • Generate scenarios: let AI create a set of scenarios for you.
  • Add scenario: write your own scenario.

Every scenario has a Title and a Prompt. The prompt describes who the customer is, how they feel and what they want, for example: "You are Priya, a 29-year-old marketing professional who finished a course last month. You are polite but anxious because you need your certificate for a job application due tomorrow."

Notion image

Your scenarios are saved between runs, so you can reuse them and edit, add or remove scenarios before each run.

3. Start the run

When your scenarios are ready, click Start test run at the bottom of the page. The AI plays each persona, has a conversation with your bot, and grades how well the bot handled it. Click Cancel if you want to discard the run instead.


💳 Projected usage

Running tests draws on your account's allowances. Before you start a run, click the cog icon ⚙️ on the New test run screen to see its Projected usage, an estimate of what the run will use and how much of each allowance you'll have left afterwards.

What each test uses

Every test in a run uses:

  • 1 evaluation
  • 1 AI interaction
  • roughly 150 AI tokens
 
Notion image

The run itself also uses 1 test run. So a run of 5 tests is projected to use 1 test run, 5 evaluations, 5 AI interactions and around 750 AI tokens.

The Projected usage panel shows these figures alongside how much of each allowance you have left, and how much will remain after the run.

ℹ️

Because each test uses one evaluation, large test runs draw on the same evaluation allowance as your Evaluations. Keep an eye on your remaining balance if you run a lot of tests.


📊 Reviewing a test run

Once a run finishes, you get a summary at the top and the full list of tests below.

The summary

At the top of the run you'll see:

  • Total tests: how many scenarios were run.
  • Pass rate: the percentage of tests that passed (graded Good).
  • The breakdown: how many tests were Good, Warning and Needs Attention.
  • The bot version that was tested (for example, Latest Draft) and when the run happened.
Notion image

The tests

Below the summary is every test, with its title, scenario and grade.

  • Use the Good, Warning and Needs Attention checkboxes to filter the list.
  • Use Expand all to open every test at once, or open one at a time with its arrow.
  • Expanding a test shows its Grades breakdown (the individual checks behind the grade, each with its own score, such as "Tools used correctly: 3/3") and the full Scenario it was tested against.
  • To see the conversation itself, click the chat icon on the test, or View in Interaction Logs inside the expanded test. Both open it in your Interaction Logs.
Notion image
ℹ️

All test runs

Click All test runs to see every run you've done. Each row shows:

  • Rerun: run the same set of scenarios again.
  • Title: the run's scenarios (the first scenario name, plus a count of the others). Click it to open the full run.
  • Results: how many tests were Good, Warning and Needs Attention.
  • Version: the bot version the run was tested against, such as Latest Draft.

You can also start a fresh run from here with New test run, and use Rerun to compare results after making changes to your bot.

Notion image

💡 Getting the most out of test runs

Test runs are most useful when you use them to catch problems before your customers do.

💡

Top tips:

  • Test before you publish. Run against your draft after changing your prompt or Tools, so you catch problems before they go live.
  • Cover the hard cases. Write scenarios for your trickiest and highest-stakes situations, not just the easy ones.
  • Re-run and compare. After a change, run the tests again and watch the pass rate and grades move.
  • Pair it with Evaluations. Use test runs to check changes before launch, and Evaluations to monitor real conversations afterwards.

❓ Frequently Asked Questions

Where do I find test runs?

Open the Digital AI (LLM Chatbot) you want to test and go to the Testing tab, next to Prompt, Tools and Evaluations.

How are tests graded?

Each test gets one of three grades: Good, Warning or Needs Attention. The pass rate is the percentage of tests graded Good.

Does running tests use up my allowances?

Yes. Each test uses one evaluation, one AI interaction and around 150 AI tokens, and each run uses one test run. Click the gear icon on the New test run screen to see the projected usage before you start.

What's the difference between test runs and Evaluations?

Test runs use scenarios you create and can be run any time, including against a draft version of your bot. Evaluations grade your real conversations after they have happened.

Do test runs affect my live chatbot or real customers?

No. The conversations in a test run are simulated, and running tests against a draft doesn't change your published bot.

 
Did this answer your question?
😞
😐
🤩