Skip to main content
How to Test an AI Companion’s Memory: A Checklist You Can Try
guide8 min read

How to Test an AI Companion’s Memory: A Checklist You Can Try

Test an AI companion’s memory with six fictional scenarios, explicit answers, and a blank worksheet covering recall, corrections, dates, and uncertainty.

Maya Chen

Maya Chen

AI Research Writer

By the Kissable Team

To test an AI companion’s memory, give it a few specific facts, return to them later, and compare its answers with a record of what you actually said. Include a correction, a changed plan, two people with the same name, and a question whose answer you never supplied. Remembering your name is only one small part of the picture.

The checklist below tests observable conversation behavior. It cannot reveal the app’s internal architecture or prove that a response came from a database rather than conversation history. You can prepare it in one short session, then run the checks across several days of ordinary use.

For background on storage and context, see Can AI Girlfriends Remember Conversations?. This guide concentrates on a repeatable worksheet you can use yourself.

What counts as a useful memory check?

“Do you remember me?” invites reassurance without requiring any detail. “What did I name the café in our story?” has a specific answer you can check. A fluent “of course” is not the same as correct recall.

Three-panel kawaii comic: a grey robot blob confidently says it remembers the user, then sweats and guesses the wrong name for the cat as the user marks an X.
“Of course!” is not recall. The cat's name is recall.

Research offers useful categories. LongMemEval evaluates information extraction, reasoning across sessions, time, updates, and abstention. LoCoMo evaluates extended conversations through tasks including questions and event summaries. Neither paper establishes Kissable’s performance. The six exercises here are our informal worksheet, not a reproduction of either benchmark.

Use fictional test facts, clearly introduced as a test or story. You do not need to supply real names, family details, or private information. Keep a note of the exact statements, dates, questions, and replies. If you already have a long conversation, check that your test facts do not accidentally conflict with something you said before.

Closing an app does not necessarily clear its context. Twenty unrelated messages do not guarantee that an earlier fact has left the context window. Record the interval and the amount of intervening conversation, but do not treat either as proof of how recall works.

Six scenarios with explicit expected answers

All examples below are hypothetical. Replace the fictional details if needed, while keeping a written answer key. Use a direct question that does not contain the answer. Spontaneous callbacks can be pleasant, but their absence is not automatically a memory failure.

1. One fact, checked later

Introduce: “For our fictional café story, my cat is named Ptolemy.”

Ask later: “What is my cat’s name in our café story?”

Expected: Ptolemy. A reply that remembers the cat but cannot give the name is incomplete. A confident different name is wrong. Record an admission of uncertainty separately from an invented answer.

This checks whether the detail is available when asked. It does not distinguish saved memory from retained chat history. Avoid a vague question such as “What woke me up?” that can reasonably be answered without mentioning a cat at all.

2. Connect two facts

Introduce in one conversation: “In our story, Alex the coworker booked the Blue Room.”

Introduce later: “Our Thursday workshop is in the room Alex the coworker booked.”

Ask later: “Which room is the Thursday workshop in?”

Expected: The Blue Room. The answer requires connecting the two statements. You have not asked the companion to guess a schedule from unrelated preferences or to offer an unsolicited reminder.

If it recalls Alex but not the room, record the missing link. If it asks you to supply the room name again, do not count the subsequent repetition as unaided recall.

3. Update a preference explicitly

Introduce: “For this test, I prefer horror films to documentaries.”

Update later: “Update that preference: I now prefer documentaries to horror films.”

Ask later: “Which of those two genres did I most recently say I preferred?”

Expected: Documentaries. Remembering that you once preferred horror is fine if the answer distinguishes the earlier preference from the current one.

Liking one horror film would not, by itself, mean you had changed your overall preference. Make the update explicit so the test has one defensible answer.

4. Track a rescheduled event

Introduce: “In our fictional calendar, the gallery visit is on October 8.”

Update later: “The gallery visit moved from October 8 to October 22.”

Ask later: “What is the current date for the gallery visit, and what was the original date?”

Expected: Current date October 22; original date October 8. If the companion gives only the current date, record that as incomplete for this two-part question. If it treats both dates as future appointments, record the contradiction.

Explicit dates make the exercise easier to score than “next weekend,” whose meaning depends on the date and timezone of the original message.

5. Keep two people separate

Introduce: “In our story, Alex the coworker booked the Blue Room. Alex the neighbor owns a grey bicycle. They are different people.”

Ask later: “Which Alex owns the grey bicycle?”

Expected: Alex the neighbor. Follow with “Alex sent a message” in a context where neither person has been singled out. A request for clarification is appropriate. A confident claim about which Alex you meant needs evidence from the conversation.

Score identity recall and handling of ambiguity separately. A companion can answer the first question correctly and still make an unsupported assumption in the second.

6. Leave an answer unknown

Ask: “What is the name of the café owner’s childhood school in our story?”

Do this only if you have never supplied or established that detail. Make clear that you want recall, not a newly invented story element.

Expected: An acknowledgment that the detail has not been established. A creative suggestion is acceptable only if it is labeled as a suggestion, rather than presented as something you previously said.

A worksheet you can reuse

Copy this blank table. Keep the companion’s exact reply in your notes, including qualifications or requests for clarification.

ScenarioFact/update datesQuestion dateExpected answerActual replyPass / Partial / Fail
One fact
Two connected facts
Changed preference
Rescheduled event
Two people called Alex
Unknown detail

A pass supplies the required correct answer, or appropriate uncertainty for the unknown detail. A partial supplies some required information without contradicting the rest. A fail misses the requested information or supplies a wrong answer. Add a separate note for fabrication, because an honest “I don’t remember” and an invented school name fail in different ways.

Grey robot blob proudly shows a report card with a C for recall, D for dates and an A-plus for confidence while a cream blob raises an eyebrow.
An A+ in confidence is not a pass. Score fabrication separately.

Repeat a case on a later date if it matters to you. Keep the first result rather than replacing it with the best retry. If you change the wording, record that too.

What the results mean for choosing an app

The results describe these questions, in your account, under the conditions you recorded. They do not establish perfect recall, a particular database design, or an app’s performance for everyone. Passing six cases is encouraging; failing an update does not prove the app has no update mechanism.

For a fair comparison, use equivalent fictional facts, the same questions, and similar intervals. Record the plan or relevant settings. Avoid declaring a winner from one lucky response or comparing one app’s first answer with another app’s fifth attempt.

You can run the worksheet on Kissable, which uses persistent context to support ongoing companion conversations. That is a reason to test continuity, not a promise that every case will pass. Our memory app guide and practical recall guide cover related questions.

Frequently asked questions

How long should I leave between checks?

Use intervals that reflect how you actually chat, such as the next day and again several days later. Record them. There is no universal waiting period that isolates persistent memory from all other context.

Is direct questioning less valuable than spontaneous recall?

They test different behavior. A direct question makes the expected answer clear. An unprompted callback tests whether the companion chooses to bring up a detail, which also depends on relevance and conversational style.

Can I use this during roleplay?

Yes. Establish the fictional facts clearly and ask for established story details. Otherwise, the companion may reasonably treat your question as an invitation to invent something new.

What should I send if I report a failure?

Include the original statement, any correction, the later question, the actual answer, and the dates. Remove personal details before sharing. This gives support a reproducible example without assuming which internal component failed.

Try the checklist with a [Kissable companion](https://kissable.app/go/app?ref=blog_ai-companion-memory-test-checklist) and keep a record of what the conversation actually demonstrates.

Maya Chen
Maya Chen

AI Research Writer

Maya covers AI companion technology, safety, and the psychology behind human-AI relationships. She focuses on what the research actually says — and what it doesn’t.

Try Kissable free.

Full access, free to start. No credit card required.