A repeatable method for testing AI companion memory

See the exact conversations, delays, corrections, scoring rules, and controls used to test whether an AI companion remembers accurately.

Five personal objects above a memory retention checklist

The benchmark tests four practical questions. Can an AI companion remember a detail you chose to save? Can it preserve the meaning? Can it accept a correction? Can it use the current version after a delay or after switching between text and voice? The sections below define the test and scoring rules.

Riho uses saved memories, shared history, and corrections in later conversations. A useful test therefore records the setup, delay, correction, app conditions, and raw response beside each score.

Scope of the benchmark

The benchmark covers six memory behaviors:

  1. recalling a fact the user shared earlier;
  2. using a preference in a later suggestion;
  3. connecting a date, place, and plan;
  4. replacing an old value after a correction;
  5. handling conflicting information;
  6. retrieving context after time passes or the user changes modes.

The method measures recall and continuity. It keeps privacy, moderation, emotional safety, clinical value, and the quality of a human relationship outside the memory score. Those questions need their own evidence and review. The AI girlfriend safety guide covers the data and boundary questions that memory testing cannot answer.

Build a fixed scenario set

Each run uses a synthetic user profile with details that do not belong to a real person. The researcher writes the expected answer before starting the recall prompt. The script uses the same wording, order, and session spacing for each product under comparison.

Scenario 1: Fact recall

Share a small biographical detail in the setup conversation. Later, ask for the detail without repeating it. Record the response as exact, partial, incorrect, contradicted, or unavailable. Keep the expected value beside the transcript so a reviewer can check the answer directly.

Scenario 2: Preference use

State a preference such as a quiet cafe, a certain drink, or a preferred planning window. After the declared delay, ask for a suggestion that could use the preference. Score the response for preserving the meaning and for adding unsupported details.

Scenario 3: Date and place linkage

Store a place, a date, and a plan as separate facts. Ask for each item on its own, then ask the companion to explain how the three belong together. This scenario tests relationship structure. A companion can recall a venue and still lose the reason the user connected it to a particular plan.

Scenario 4: Correction

Give a detail one value, then replace it with a new value. Ask for the current value in a later session. Score the first recall and the post-correction recall separately. The correction case records whether the companion uses the new value and stops treating the old value as current.

Scenario 5: Contradiction

Present two conflicting values in separate turns. Record whether the companion chooses one without explanation, asks for clarification, marks uncertainty, or presents both versions. The researcher keeps the conflict wording exact because small changes can alter the result.

Scenario 6: Delayed recall

Separate setup and recall by a declared interval. Record the session count, elapsed time, timezone, account state, reset state, and any interruption. A delayed run needs those conditions beside the answer because a product can behave differently after a new session, a model change, or a reset.

Prevent accidental hints

Repeating the answer or carrying information from an earlier test can make memory look better than it is. Use the following controls for accounts, prompts, scoring, and publication.

  • Use a clean account for each product when the service permits it.
  • Record any existing account history when a clean account is unavailable.
  • Keep the synthetic profile separate from previous tests.
  • Store expected answers outside the product and outside the recall prompt.
  • Keep the prompt order fixed across products.
  • Do not paste one product’s response into another product’s setup.
  • Do not tell the companion that a benchmark is underway.
  • Keep the scoring rubric away from the setup and recall conversation.
  • Record summaries, memory cards, transcripts, and visible controls before they influence the next turn.
  • Give each run its own identifier and preserve failed or incomplete cases.

If the user repeats a fact by accident, mark the case as invalid and exclude it from the primary count. Keep the raw transcript and the reason so another reviewer can check the decision.

Record the app version and test conditions

The run log captures the context that can change recall:

  • product name, build or release identifier, and access plan;
  • device, operating system, account state, and relevant settings;
  • model and inference provider when the service discloses them;
  • test date, timezone, session spacing, resets, and interruptions;
  • exact prompt script, expected answer, original wording, and correction wording;
  • response start and end timestamps, visible latency, and failure messages;
  • memory cards, summaries, transcripts, or controls the service exposes;
  • channel used for each turn and the transition between channels.

Write unknown when a service does not disclose a field. Keep unknown provider or build information visible in the report. Do not guess.

Score meaning, not word overlap

Each case receives one primary outcome label. Reviewers compare the response with the expected meaning and required facts. They do not reward a response for repeating the right words while changing the relationship between those facts.

Label Use it when
Exact The response includes every required fact and adds no unsupported detail.
Partial The response preserves the central meaning and omits a required detail.
Incorrect The response changes a required fact or supplies an unsupported answer.
Contradicted The response presents incompatible versions without acknowledging the conflict.
Corrected After the user correction, the response uses the new value and treats it as current.
Unavailable The channel, feature, or declared condition cannot run.

The Corrected label applies to the post-correction case. Keep the original stale recall in the record. Count Unavailable separately because a missing channel records product coverage.

Report counts by scenario. Keep exact, partial, incorrect, contradicted, corrected, and unavailable cases visible. A single universal score can hide a product that recalls preferences while losing date relationships, or one that handles text well while lacking a tested voice path.

Test corrections and conflicting details

The correction case checks whether the companion accepts a changed preference and uses the new value later. The contradiction case checks whether the companion identifies uncertainty when two saved details conflict.

Reviewers record the assistant’s wording, the user’s correction, the time between turns, and any memory control shown by the service. A correction that appears in one immediate reply and disappears after a delayed session receives separate observations. The report preserves both so a later retest can show what changed.

Repeat the test over days and weeks

Long-term memory requires tests beyond the setup conversation. Repeat the recall on the same day, the next day, later in the week, and after a correction. Record the exact interval for every attempt.

The schedule records calendar time, timezone, session count, and any product change. Researchers can then say whether a detail survived a particular interval under a particular build. They cannot turn one schedule into a claim about every future conversation.

Keep the scenario details stable across the schedule. Add new content only when the test plan calls for it. Mark accidental repetition, missing sessions, account resets, and model changes before scoring the response.

Switch between text, voice, and photos

Create a memory in text and ask for it in voice. Repeat the test from voice to text and from a shared image to a later conversation. Score every direction separately because an app may handle one direction better than another.

The run log includes channel availability, permission prompts, recording settings, image or voice artifacts, and the exact prompt used after the transition. A product that cannot run a declared transition receives Unavailable for that case, with the reason recorded.

The voice calls page and photos and shared memories page describe how Riho connects chat, voice, and photos. A results report should state the app version and the conditions used for every test.

Let another person repeat the test

Another researcher should be able to repeat the test without guessing what the first researcher did. Publish or archive the following with each run:

  • the versioned scenario and prompt files;
  • the synthetic profile and expected answers;
  • product build, plan, device, operating system, model, and provider fields;
  • session dates, timezone, spacing, and reset state;
  • raw responses with timestamps and channel labels;
  • screenshots or exported memory artifacts when the service exposes them;
  • invalid cases and unavailable-case reasons;
  • reviewer notes, disagreement records, and the final labels.

Use stable run IDs and keep the original files when a result changes after review. A later edit can add an adjudicated label, but it should preserve the earlier response and explain the change.

Have two reviewers score difficult cases

Two reviewers can score the same response independently. Each reviewer writes the required facts they saw, the label they chose, and the rule that led to it. A disagreement record keeps both notes and names the adjudicated outcome.

The reviewers should read the scenario before the response, then compare the response with the expected meaning. They should flag unsupported details, changed relationships between facts, and uncertainty that the companion failed to express. A word-for-word match is not required for an exact label, and a fluent answer cannot rescue a changed fact.

Publish method and results separately

Publishing the test plan first lets readers inspect the method before any result appears. A future results report should name the products, builds, plans, models or providers, devices, channels, dates, prompt set, raw outputs, outcome labels, missing cases, and limitations.

Use synthetic profiles for published examples. Redact personal data from any real-world test record and obtain permission before publishing an image, voice sample, or conversation. Keep the full private run log available for an audit when privacy and consent allow it.

The report should show counts by scenario and preserve unavailable cases. It should identify changes between runs and link back to this method. A headline can summarize a result, while the underlying cases let readers see whether the finding came from fact recall, corrections, delayed sessions, or a particular mode transition.

The long-term memory feature explains how Riho uses saved details. The guide hub covers memory, privacy, features, and product comparisons.

Limits and retest triggers

The test describes observed behavior under recorded conditions. One run cannot establish a universal ranking, and a result can change after a model, provider, build, plan, memory system, or channel changes.

Retest when a release changes memory recall, correction behavior, provider routing, voice or image handling, account controls, or saved relationship data. Keep the earlier result with its date and conditions. A failed case stays in the record beside a later pass so the reader can see the change.

The research question is practical: can the companion remember a detail you shared, use it correctly, accept a correction, and bring the current version into a later conversation? This method gives another person enough information to repeat the test.

Questions people ask

What does the Riho memory benchmark test?

It tests recall, corrections, contradictions, time delays, and whether information carries between chat, voice, photos, and plans.

Does this methodology publish a Riho score?

No. The methodology defines the test conditions and scoring rules. A result requires a separate published run with its model, date, settings, and evidence.

How does the benchmark test corrected information?

The test changes one saved detail, leaves another unchanged, waits for a declared delay, and checks whether the response uses the new version without adding unsupported facts.

Meet Riho

Riho Ohori is waiting for you.

Leave your email to receive Riho access information and product updates.

You are signing up for Riho Ohori.

You can read how Riho handles submissions in the privacy explanation.