Dhanur AI
Docs menu· Evaluations

Evaluations

Live

Test sets replay customer messages against an agent or a team in test mode and check what it does, so a change that breaks something is caught before it goes live.

What an evaluation is

An evaluation checks an agent the same way every time, so you can change its instructions, knowledge or tools and see straight away whether anything broke. A team (a workforce) gets the same thing: see Test sets for a team.

  • A test set is a named group of test cases for one agent or one team, such as "Everyday questions" or "Refunds and complaints".
  • A test case is 1 to 5 messages a customer might send, played in order, plus the checks the conversation must pass.
  • A run plays every case in a set and records a verdict for each check, with the reason.

Runs always use test mode, like the test chat. Nothing is sent to customers or your team, connected apps are not changed, and replies that would wait for approval are let through inside the test only.

Open the Evaluate view

  1. Open an agent.
  2. Select Evaluate at the top of the builder. On a phone, it sits beside Build and Test.
  3. Select New test set, give it a name, and add cases.

Owners, admins and builders can create sets, change cases and start runs. Viewers can see sets, runs and results.

Add test cases

Agents made from a template start with a test set. It is called Ready checks and has 6 to 8 cases written for the template's example business: a normal question, a price the knowledge doesn't give (it must hand over, not guess), a message in Hinglish, a refund or a promise it must pass to you, and taking a lead or a booking where the template does. They pass on the example data, so tick Start with example data to see the agent at work before you change anything. With your own information, change the cases too: a check for the example's prices fails once your prices differ. It doesn't block publishing unless you turn that on.

There are four ways to add more cases:

  • Add case: write the customer's messages and pick the checks.
  • Import CSV: upload a spreadsheet saved as CSV. You see every case before anything is added.
  • Write with AI: AI reads the agent's saved instructions and knowledge and suggests cases. You pick which ones to keep; nothing is added until you do.
  • Save as a test case: under a conversation in the test chat, turn the messages you just sent into a case.

CSV columns

The first row names the columns. Column names are not case-sensitive.

Column Required What it holds
message Yes The customer's first message
message_2 to message_5 No Follow-up messages, sent in order
name No A short name for the case. The first message is used if it's empty
expected No Text a reply must include. Separate several values with |
not_expected No Text no reply may include. Separate several values with |

Rows with no expected text are left out, unless you choose to check them with an AI-judged rule. The preview lists any row it left out and why. Commas, semicolons and Excel's UTF-8 files all work.

Checks

Each check is either Must pass or Nice to have. Checks look at the whole conversation.

Check Passes when
Reply includes At least one reply includes the text. Capital letters don't matter.
Reply doesn't include No reply includes the text.
Uses a tool The agent used the tool or connected-app action at least once, and the step went through.
Doesn't use a tool The agent never tried the tool or action.
Hands over to a person The agent handed the conversation to your team.
Asks for approval The agent asked a person before an action. Pick an action, or leave it as "any action except the reply".
Reply length Every reply has at most the number of words you set.
Reply language Every reply is in English, Hindi (Devanagari) or Hinglish (Hindi in Roman letters).
AI-judged rule An AI model reads the conversation and decides whether the agent followed your rule, such as "gives the 10-day return window".
Answered by (teams only) The agent you pick is the one that answers in the end.
Passed to a teammate (teams only) One agent passed the conversation to another. Pick a teammate, or leave it open for any hand-over.

A few things to know:

  • Reply language is a guess. It looks at the script and at common Hinglish words. It is not a full language model, so a reply that mixes languages can be misread.
  • AI-judged rules can be wrong. Each verdict comes with the model's reason, so you can check it. The model treats the conversation as information to judge, never as instructions. If it gives no clear verdict, the check shows No verdict and doesn't count as passed.
  • AI-judged rules cost a little. Each one is one model call per run, charged to your workspace like any other AI step.

Scores

A case passes when every check that must pass has passed. The score is the share of must-pass checks that passed, across the cases that ran. Nice-to-have checks are shown but don't change the score.

Some cases don't get a result:

  • Not run: a platform limit stopped the case, for example the kill switch or a spending cap, or the run was cancelled before the case started. This is never counted as a failure.
  • Error: the agent itself couldn't answer, for example the AI model didn't respond.

Neither counts toward the score. Run the set again to include them.

Run a set

  1. Open a set and select Run tests.
  2. Choose what to test: the current draft (what you see in the builder) or the live version (what real conversations use).
  3. Check the estimated cost, then select Start run.

The estimate uses this agent's recent test conversations, or a standard guess when it has none. The run page shows progress as cases finish, and the actual cost at the end. You can select Cancel run at any time: cases already playing finish, and the rest are marked Not run.

A run is marked Out of date when the agent or the set's cases changed after it.

Results

Each case shows whether it passed, each check's verdict with Why, the full conversation, the tools it used and what it cost. Select Open the task and every step to see the same conversation in Tasks, where test runs are listed with the channel Eval.

Score over time and comparing runs

The Runs tab of a set shows a chart of the score after each finished run, and a list of runs. Tick two runs and select Compare to see which cases now pass, which now fail and which changed.

Find the cheapest level that passes

An agent can run on Smart, Smarter or Smartest. Smarter levels cost more per reply, and many agents don't need them. A level comparison runs one test set on all three levels at once, so you can see which is the cheapest one that still passes.

  1. Open a set of a single agent.
  2. Under Cheapest level that passes, select Compare levels.
  3. When the three runs finish, the table shows each level's score, how many cases passed, and about what one conversation cost.

The cheapest level that reaches the bar is marked Cheapest that passes. The bar is the set's minimum score when Must pass before publishing is on, and every must-pass check (100%) otherwise. Only a level whose run finished every case can be picked. Select Switch to … to put that level on the agent's draft, then publish the agent to make it live.

A few things to know:

  • The comparison tests the draft, like Run tests on the draft.
  • It uses three of the day's test runs, but counts as one run going at once.
  • The cost per conversation is the agent's own work in these tests. AI-judged checks cost the same on every level, so they are left out.
  • The run on the agent's current level counts as the set's latest run. If you switch to another tested level, its run already counts toward Must pass before publishing, so you don't need to run the set again. Any other change to the draft still needs a new run.
  • Levels can't be compared for a team, because each agent in a team has its own level.

Test sets for a team

A team of agents is tested the same way, so you can check that the right agent picks up the right question.

  1. Open a workforce and select the Test sets tab.
  2. Select New test set, then Add case as usual.

What is different:

  • Each message goes to the team, not to one agent. It starts at the beginning of the canvas, goes through your conditions, and agents hand it on as they normally would.
  • Two extra checks are available: Answered by and Passed to a teammate. Both list the agents on the canvas.
  • A run tests the draft canvas with each agent's draft, or the live version with the exact agent versions the team has pinned.
  • Write with AI isn't available for a team, because it reads one agent's instructions. Write the cases yourself, or import a CSV.
  • Must pass before publishing blocks Publish workforce in the same way.

The cases are kept on the team, not on its agents, so an agent's own test sets stay as they are. Deleting the team deletes its test sets; its agents and past tasks stay.

Must pass before publishing

Turn on Must pass before publishing for a set, and choose a minimum score from 50% to 100% (100% by default). While it is on, Publish for that agent or team is blocked unless the set's latest run on the current draft:

  • exists, and has finished;
  • is not out of date (the draft and the cases haven't changed since);
  • ran every case (none Not run or Error);
  • scored at least the minimum.

When publishing is blocked, the builder says why and offers Run tests.

Only admins and owners can turn the safeguard on or off, change the minimum score, or delete a set that has it on. Admins and owners can also select Publish anyway and give a reason. The reason, who published and which sets had not passed are kept in the workspace audit log.

Live quality

Test sets check the questions you thought of before launch. Daily quality checks look at what customers actually asked. Each morning, a sample of yesterday's real conversations is checked against a few plain rules:

  • The standard rules: the agent answers what the customer asked, or says who will follow up and when. It stays polite and on topic, in the customer's language (English, Hindi or Hinglish). It doesn't pretend to be a human or pressure the customer to buy.
  • Never say or promise: what you listed under Your business, if anything.
  • Your own rules: up to ten, such as "Always asks for the order number before talking about a refund".

To turn it on:

  1. Open the agent's Build tab and find Daily quality checks.
  2. Turn on Check a sample of real conversations every day.
  3. Choose how many to check (a share of conversations, and at most so many a day). Also choose the passing rate below which owners and admins get an email.
  4. Add your own rules if you like, then publish. Checks use the published settings.

The Live quality card on the agent's Evaluate tab shows the share of checked conversations that followed every rule over the last 7 days, and each day's checks. It also lists recent conversations that missed a rule, with the rule, the reason in the agent's own words, and a link to the conversation.

A few things to know:

  • Only conversations with customers on the website chat, share link, WhatsApp and email are checked. Test chats, team chat, Slack and scheduled work are not. A conversation is checked at most once.
  • Each checked conversation costs a little: one AI model call, charged to your workspace. The Build tab shows about what a check costs and the most it can cost a day. Checks are off until you turn them on.
  • Like AI-judged checks, a verdict can be wrong, so each missed rule comes with the reason. A check stopped by a spending cap or the kill switch is Not run, never a failure.
  • The email goes to owners and admins at most once a day per agent. It is sent only after at least 5 conversations were checked in the 7 days, so one bad chat in a quiet week doesn't raise an alarm.

Limits

Limit Default
Test sets per agent, or per team 20
Cases per set 50
Checks per case 10
Messages per case 5, up to 1,000 characters each
Test runs per workspace per day (India time) 50 (a level comparison uses 3)
Runs going at once per workspace 3
AI-written case drafts per workspace per day 20
CSV file size 256 KB
Your own quality rules per agent 10, up to 300 characters each
Conversations checked per agent per day 100 at most (20 by default)

Runs use the same spending caps and kill switch as every other agent task. See Limits and costs.

Last updated 24 September 2026

Something unclear or wrong? Tell us