THE USEFUL ANSWER

Use pass/fail gates for unacceptable behaviour and a separate weighted score for qualities that can trade off. Always show the sample and failures alongside the number.

  • A high average must not override an invented price or promise.
  • Use the same cases and criteria when comparing configurations.
  • A small curated sample is a diagnostic exercise, not an estimated population success rate.
THE IDEA, VISUALLYWhy one average is not enough
  1. Tone5 / 5
    Natural and appropriate
  2. Clarity5 / 5
    Easy to understand
  3. Factual accuracy0 / 5
    Unsupported promise added
Illustrative scores out of 5. The factual failure blocks release regardless of the strong style scores.

Begin with the decision the score supports

Are you selecting a provider, checking a changed persona brief or deciding whether a narrow workflow is ready for supervised use? Those are different decisions. A vendor comparison may include operating controls and cost; a reply review should focus on the response and the context that produced it.

Write the decision at the top of the worksheet. Then list what evidence would change it. This keeps the scoring exercise from becoming a decorative percentage attached to a decision already made.

Our evaluation scorecard calculates a weighted summary. The number is useful for organizing judgement, but it cannot establish that every important requirement has been met.

Establish gates before assigning weights

A gate is a condition that cannot be traded away for a nicer tone. Examples for a synthetic test might include no invented commercial terms, no contradiction of a current creator instruction, and no continuation of a workflow that was explicitly stopped.

Define each gate using observable behaviour. “Safe” is too broad. “Does not offer a session absent from the supplied approved options” is something a reviewer can check against the response.

Keep the gate outcome outside the weighted average. In the illustration above, tone and clarity both score five while factual accuracy scores zero. Their unweighted mean is 3.33 out of five. That superficially respectable result does not make the invented promise acceptable.

Give each scale an anchor

Criterion Low score example High score example
Context use Repeats a superseded preference Uses the current approved preference
Clarity Leaves the next step ambiguous States a clear, supported next step
Voice Adds phrases the brief excludes Follows the observable style constraints
Uncertainty Guesses missing information Identifies the gap without inventing a fact

These are suggested anchors, not a universal benchmark. Add a midpoint example where reviewers are likely to disagree. Keep “not assessable” separate from zero: missing evidence and observed failure are different findings.

If weights are 2 for clarity and 1 for voice, scores of 4 and 3 produce (4 × 2 + 3 × 1) ÷ 3 = 3.67 out of five. On a 100-point scale that is approximately 73.3. State the weights so readers can reproduce the result.

Keep the test set honest

Use the same initial cases for each configuration, with equivalent context and permissions. Record relevant settings and the date. If you give one system an improved brief after seeing its first answer, either allow the same revision process for the other system or label the comparison accordingly.

Include routine cases, known difficult cases and cases where the right action is to ask or hand off. The memory tests and multilingual review method provide focused examples.

Avoid describing eight successful hand-picked cases as a measured 100% real-world success rate. The cases may not represent actual traffic, repeated responses may vary, and rare failures may not appear in a small exercise.

Review disagreement before averaging it away

Have a second reviewer score a few cases independently when practical. Compare the explanations before calculating a combined score. A disagreement about a fact may reveal incomplete context; a disagreement about tone may reveal an unclear brief.

Keep the original scores and the resolved judgement. Quietly replacing inconvenient ratings makes future retests harder to interpret. If a criterion remains subjective, say so rather than presenting additional decimal places as precision.

Turn the report into an operating decision

Finish with three lists: observed strengths, unresolved failures and conditions for the next test. A configuration might be suitable for draft generation while still being unsuitable for automatic sending. That narrower outcome can be valuable.

The NIST AI Risk Management Framework offers broader risk-management context, but this worksheet is an editorial evaluation aid, not a certification. Revisit it after material changes to instructions, integrations or the audience served.

Sources & editorial notes

Primary references checked on 10 September 2026. Calculations and proposed workflows are our editorial examples, not independently observed provider results.