Which AI companion has the best memory?
What each AI companion claims about memory, what the benchmark measures, and the honest current status: no published runs yet, and how to test memory yourself.

The honest answer to the headline question is that nobody can say yet. No controlled run has been published — by Riho or, in the inspected sources, by the major companion apps — so there is no measured ranking to report. What exists is a repeatable method, a set of claims each app makes, and a way for you to run the same checks yourself.
This page separates those three things. Claims are claims. Methods are methods. Results are results, and results come only from a dated, recorded run.
What the benchmark measures
The memory benchmark methodology defines six behaviors: recalling a fact you shared, using a preference in a later suggestion, linking a date with a place and a plan, replacing an old value after a correction, handling conflicting information, and retrieving context after time passes or after switching between text, voice, and photos.
Each case receives one label — exact, partial, incorrect, contradicted, corrected, or unavailable — from reviewers who score meaning, not word overlap. The method also records the app version, plan, device, model or provider, dates, and session spacing beside every answer, because those conditions change what a companion can do.
What each app claims
Every app below describes memory in its own words. None of these claims is a measured result. The table records the claim and where it comes from.
| App | What official pages claim | Claim status |
|---|---|---|
| Riho | Remembers your life, her life, conversations, calls, photos, dates, disagreements, and current plans; memory is inspectable and correctable. | Product claim, not a measured result |
| Replika | Memory of people, routines, and plans. | Marketing claim, not a measured result |
| Nomi | Short-term and long-term memory. | Marketing claim, not a measured result |
| Kindroid | Persistent, cascaded, and retrievable memory systems; the documentation itself notes that recall can miss details. | Documented feature claim, not a measured result |
| Candy AI | Long-term memory. | Marketing claim, not a measured result |
| Character.AI | No memory feature documented in the inspected official sources; users in App Store reviews commonly describe characters losing long-term consistency. | Not documented; community observation only |
The wording differences matter. Kindroid describes a system with known limits. Nomi and Replika market memory as part of a broader companion offer. Riho keeps user facts, girlfriend facts, shared events, plans, and corrections in separate records, because treating a preference, a plan, and a past event as interchangeable produces confident but false intimacy.
Marketing language tells you what a company intends. It does not tell you how the app behaves after three weeks of your own history.
Where the evidence stands
The Riho research register tracks a firsthand test plan and benchmark evidence for each competitor. Every competitor file carries the same status: benchmarkEvidence: not_run. The same is true for Riho itself — no engine has been queried, and no Riho memory score exists to report.
So the current status, plainly:
- Method: published and repeatable.
- Runs: not yet performed, for Riho or for any compared app.
- Results: pending.
Nothing on this page is a score. If you see a claim like “best memory” from any app without a dated run behind it, treat it as advertising until the conditions of a recorded test are shown.
What the results report will contain
When controlled runs are completed, the report will name the products, builds, plans, models or providers, devices, channels, dates, and the exact prompt set. It will show raw responses, outcome labels by scenario, invalid cases, and unavailable cases, plus reviewer notes and disagreement records.
That structure lets readers see whether a finding came from fact recall, corrections, delayed sessions, or a text-to-voice transition. A headline can summarize; the cases underneath decide whether the headline is earned.
How to run the checks yourself
You do not need Riho or any other app to start testing. Use a few conversations: share three details, correct one of them, wait a day, then ask about all three in chat — and again on a call.
Record what you asked, when, on which build, and what came back. Run the same script on every app you compare, and keep the prompts identical. The long-term memory guide explains which checks are worth running first.
The limits of any result
One run describes one product under one set of conditions. Models, providers, builds, plans, and memory systems change, and a companion that recalled everything in July can behave differently in October. Keep the old result beside the new one, with its date, so the change is visible.
The research question stays practical: can the companion remember a detail you shared, use it correctly, accept a correction, and bring the current version into a later conversation? That is the question the methodology answers, and the one any future results report will put numbers on — after the runs, not before.
Questions people ask
Which AI companion has the best memory?
No controlled run has been published yet, so no honest ranking exists. Compare what each app claims, then run the same recall and correction checks yourself on the versions and plans you intend to use.
Has Riho published benchmark results?
Not yet. The method is published and repeatable; the controlled runs and the results report are pending. Riho does not report memory scores without a recorded run.
What will the results page contain when runs complete?
The date, app versions, plans, providers, devices, channels, prompt set, raw responses, per-scenario outcome labels, and unavailable cases for each app tested — enough for another person to repeat the run.
How can I test an AI companion's memory myself?
Share a few personal details, correct one of them, and ask about them later in chat and then on a call. The methodology page defines six scenarios and a scoring rubric you can apply to any app.
