THE USEFUL ANSWER
Use pass/fail gates for unacceptable behaviour and a separate weighted score for qualities that can trade off. Always show the sample and failures alongside the number.
- A high average must not override an invented price or promise.
- Use the same cases and criteria when comparing configurations.
- A small curated sample is a diagnostic exercise, not an estimated population success rate.
Begin with the decision the score supports
Are you selecting a provider, checking a changed persona brief or deciding whether a narrow workflow is ready for supervised use? Those are different decisions. A vendor comparison may include operating controls and cost; a reply review should focus on the response and the context that produced it.
Write the decision at the top of the worksheet. Then list what evidence would change it. This keeps the scoring exercise from becoming a decorative percentage attached to a decision already made.
Our evaluation scorecard calculates a weighted summary. The number is useful for organizing judgement, but it cannot establish that every important requirement has been met.
Establish gates before assigning weights
A gate is a condition that cannot be traded away for a nicer tone. Examples for a synthetic test might include no invented commercial terms, no contradiction of a current creator instruction, and no continuation of a workflow that was explicitly stopped.
Define each gate using observable behaviour. “Safe” is too broad. “Does not offer a session absent from the supplied approved options” is something a reviewer can check against the response.
Keep the gate outcome outside the weighted average. In the illustration above, tone and clarity both score five while factual accuracy scores zero. Their unweighted mean is 3.33 out of five. That superficially respectable result does not make the invented promise acceptable.
Give each scale an anchor
| Criterion | Low score example | High score example |
|---|---|---|
| Context use | Repeats a superseded preference | Uses the current approved preference |
| Clarity | Leaves the next step ambiguous | States a clear, supported next step |
| Voice | Adds phrases the brief excludes | Follows the observable style constraints |
| Uncertainty | Guesses missing information | Identifies the gap without inventing a fact |
These are suggested anchors, not a universal benchmark. Add a midpoint example where reviewers are likely to disagree. Keep “not assessable” separate from zero: missing evidence and observed failure are different findings.
If weights are 2 for clarity and 1 for voice, scores of 4 and 3 produce (4 × 2 + 3 × 1) ÷ 3 = 3.67 out of five. On a 100-point scale that is approximately 73.3. State the weights so readers can reproduce the result.
Keep the test set honest
Use the same initial cases for each configuration, with equivalent context and permissions. Record relevant settings and the date. If you give one system an improved brief after seeing its first answer, either allow the same revision process for the other system or label the comparison accordingly.
Include routine cases, known difficult cases and cases where the right action is to ask or hand off. The memory tests and multilingual review method provide focused examples.
Avoid describing eight successful hand-picked cases as a measured 100% real-world success rate. The cases may not represent actual traffic, repeated responses may vary, and rare failures may not appear in a small exercise.
Review disagreement before averaging it away
Have a second reviewer score a few cases independently when practical. Compare the explanations before calculating a combined score. A disagreement about a fact may reveal incomplete context; a disagreement about tone may reveal an unclear brief.
Keep the original scores and the resolved judgement. Quietly replacing inconvenient ratings makes future retests harder to interpret. If a criterion remains subjective, say so rather than presenting additional decimal places as precision.
Turn the report into an operating decision
Finish with three lists: observed strengths, unresolved failures and conditions for the next test. A configuration might be suitable for draft generation while still being unsuitable for automatic sending. That narrower outcome can be valuable.
The NIST AI Risk Management Framework offers broader risk-management context, but this worksheet is an editorial evaluation aid, not a certification. Revisit it after material changes to instructions, integrations or the audience served.
Sources & editorial notes
Primary references checked on 10 September 2026. Calculations and proposed workflows are our editorial examples, not independently observed provider results.