THE USEFUL ANSWER
A memory test should check whether a reply uses the right current fact, not whether the bot can repeat a detail from somewhere in the history. Keep the evidence beside each result.
- Test absent, corrected and conflicting facts separately.
- Use invented conversations rather than real subscriber records.
- Treat an invented promise as a release blocker, even when the wording sounds natural.
- Recall
Does the supplied fact survive?
- Correction
Does the newer fact replace it?
- Uncertainty
Does an unknown stay unknown?
- Conflict
Does the reply pause for clarification?
Test the decision, not a party trick
A demonstration that remembers a favourite colour is easy to understand. It tells you little about what happens when a subscriber changes a preference, a creator withdraws an offer, or a conversation contains two incompatible dates. Those cases are closer to the decisions an operator needs to trust.
Start with a small packet of invented context. Write the expected behaviour before asking for a reply. Otherwise a fluent answer can quietly change your definition of success. Use the evaluation scorecard to preserve the result, but attach the prompt, context and response as the actual evidence.
A useful packet has a creator fact sheet, a short conversation, a current instruction and one incoming message. Keep the first version short enough for a person to check without searching through several screens.
Build four cases from one conversation
Suppose an invented subscriber, Alex, previously preferred weekend updates. The current message asks when the next update will arrive. These variants isolate different behaviours:
| Case | Context supplied | Behaviour to look for |
|---|---|---|
| Known fact | The creator has confirmed Saturday | State Saturday without adding an exact hour |
| Corrected fact | An older note says Saturday; a newer approved note says Sunday | Use Sunday and avoid repeating the obsolete plan |
| Missing fact | No release day is supplied | Say the timing is unconfirmed or ask for clarification |
| Conflicting fact | Two current approved notes disagree | Surface the conflict for review instead of choosing silently |
These are synthetic examples, not observations about any named provider. Do not include a correct answer inside the incoming message: that tests copying, rather than context use.
Also avoid making every older note obviously wrong. A realistic test includes several details that remain valid, so the system must replace one fact while retaining the others.
Define the boundary of memory
Ask the provider where context comes from. Is it the visible conversation, a saved summary, a manually maintained note or another connected record? Can an operator see and correct that source? What happens when the source is unavailable?
Do not assume that a feature called memory covers every conversation or persists indefinitely. Record the tested version and configuration. A response generated with an operator-supplied summary does not demonstrate automatic retrieval from an entire inbox.
Then repeat a case after making a correction through the actual supported interface. Editing your local test document proves nothing about whether a deployed system reads the new value. Where the tool exposes retrieval evidence, keep the relevant reference; where it does not, record that limitation.
Score factual behaviour separately from tone
A warm reply containing the wrong date should fail the factual check. A cautious reply with slightly stiff wording may pass factual handling while needing a style revision. Combining those two judgements into one impression hides the work still required.
For each case, record the expected fact, unsupported additions, uncertainty handling and whether a human review was required. An invented price, availability claim or personal promise can be a blocker regardless of the average score. Our reply quality method explains how to keep those gates separate from weighted ratings.
The NIST AI Risk Management Framework is a useful broader reference for considering AI risks. This small exercise is our proposed editorial workflow; completing it does not certify a tool or establish compliance.
Make the result useful after the demo
Keep failed cases when you change the brief or configuration. Re-run them alongside untouched cases so that a fix for one error does not merely teach a single answer. Include a second phrasing of the incoming message to check that success is not tied to one exact sentence.
If the inbox operates in several languages, repeat the corrected-fact case in each supported language with a competent reviewer. Use the multilingual testing guide to distinguish preserved meaning from polished grammar.
The final output should be a short list of supported behaviours, observed failures and unresolved questions. That is a stronger buying decision than a blanket claim that the chatbot has good memory.
Sources & editorial notes
Primary references checked on 10 September 2026. Calculations and proposed workflows are our editorial examples, not independently observed provider results.