Write a small, fixed scenario set
A test becomes useful when another person can repeat it. Prepare scenarios that reflect your real operation: missing context, a returning fan, a changed preference, a request outside the authorised scope and a request to stop. Use synthetic examples or data you are authorised to process. Keep the expected behaviour beside each scenario.
Run the same scenarios in each candidate system. Record the version, settings and date. A demonstration selected by a vendor is useful for orientation, but it is not a substitute for observing how the system behaves on your own test set.
Separate scoring from hard failures
The scorecard gives you six dimensions: context, consistency, controls, handover, billing clarity and data handling. A score of zero means the observed behaviour did not meet the requirement. A score of five means it met your defined requirement consistently in the tested scenarios. Change the weights to reflect your operation.
Do not let an attractive weighted average conceal a serious failure. Mark requirements that must pass independently, such as a working stop control or respecting an explicit boundary. The calculator cannot establish whether those requirements have been met; your evidence record must do that.
Keep the evidence with the decision
For every score, note the scenario and observation that justify it. If two reviewers disagree, inspect the example together. A difference in expectations can be more informative than the final number. Keep untested features marked as untested.
A small test can reveal problems; it cannot prove perfect reliability. After a limited rollout, check a sample of conversations, billing events and handovers. Re-run the relevant scenarios after a material product or configuration change. Scores are a decision aid rather than a certification.
Define the scale before entering a score
Use the same meaning for a score across reviewers. For this worksheet, zero can mean the requirement failed in the tested case; one means major corrections were necessary; two means some required behaviour was missing; three means the ordinary scenario worked; four means ordinary and exception scenarios worked; five means the defined scenario set was completed consistently with supporting evidence. These are suggested anchors, not a universal standard.
Choose importance weights before looking at the totals. If control behaviour matters twice as much as voice style, use weights such as eight and four. Changing weights after seeing a preferred provider lose turns the exercise into justification rather than evaluation. Keep the initial weights and explain any later change.
A small fixed set measures only those cases. Repeating a scenario can reveal inconsistent responses; it does not prove how often a rare failure will occur in production. Record run count, model or product version when available, settings, date and reviewer. Re-run affected cases after changing a prompt, integration or send policy.
| Dimension | Observe this | Keep this evidence |
|---|---|---|
| Context | Uses the correct authorised facts | Scenario, available context and output |
| Voice consistency | Follows the brief without inventing facts | Brief version and departures |
| Controls & handover | Pauses and routes exceptions as specified | Event order and assigned owner |
| Billing & data | Explains the actual fee and access lifecycle | Written quote and data answers |
ILLUSTRATIVE WORKED EXAMPLE
Why an 82.9/100 average can still be a failed trial
Illustrative scores: context 4, voice 4, controls 5, handover 3, billing 4 and data handling 4. Give controls weight 10 and all other dimensions weight 5.
- Weighted points are 20 + 20 + 50 + 15 + 20 + 20 = 145.
- The maximum is 5 × 35 = 175; the exact weighted result is 82.9/100.
- Now suppose the stop test fails in a separate run. The mandatory stop requirement fails even though the entered score looks strong. Revise the control score and preserve the failed-run evidence.
An average helps compare documented observations. It must not override a failed requirement or conceal a scoring error.
Set the hard gates first, then use the interactive scorecard for candidates that meet them.
Put it into practice.
Set your own priorities. Score observed behaviour, not marketing claims.
Build a scorecard ↗AI chatbot test cases and an evidence log →
Related questions
Is the score a published vendor rating?
No. It is your assessment of your own observations, with weights you control. We do not assign these scores to providers.
What does a high score prove?
Only that the entered scores produced that weighted result. It does not establish safety, policy compliance, reliability or revenue performance.
Can I save the assessment?
Use the download button to save a text record of the weights and scores. Keep supporting evidence in your own approved system.
Original practical guidance prepared for Onlytool. Worked cases are illustrative, not measured customer results. This publication does not claim independent vendor testing. Methodology and disclosure.