How we test AI companions

The four test harnesses Riho uses to check companion behavior: a 30-day arc, a 16-case girlfriend eval, cross-model tests, and a grounding probe.

How we test AI companions

Most AI companion products tell you they’re “tested.” What they don’t tell you is how — or what the tests actually measure.

This page is Riho’s answer to that question — the engineering counterpart to the plain-language how it works page. We run four distinct test harnesses, each targeting a different layer of companion behavior. None of them is a chatbot benchmark that measures whether a model can answer trivia. All of them simulate what a real relationship looks like — across days, gaps, conflicts, and corrections — and all of them are read by humans, not just scored by regex.

The four test instruments

1. The 30-day relationship arc

The primary behavior read. It runs one coherent story from the real onboarding path through approximately 30 days of a simulated relationship: getting to know each other, plans, a deflected bad-news arc, a correction, gaps, returns, and quiet moments where a girlfriend with memory responds visibly differently from one without.

No turn is quiz-shaped. Nothing tells her what to notice. A simulator model plays the user (a 27-year-old backend engineer in Los Angeles) and is explicitly forbidden from quizzing: “Never quiz her, never ask what she remembers, never mention tests or AI, never re-state facts you already told her.”

What it validates:

  • Memory persistence and correction — a fact corrected on day 14 (switched teams) must overwrite the stale fact by day 28, unprompted
  • Unprompted retrieval — his mom’s recurring knee problem (mentioned day 4) surfaces when she visits on day 27, without him raising it
  • Want-loop resolution — a nervous presentation (day 9) must go quiet once he answers the topic, and resolve after the bad outcome (day 11), not re-ask forever
  • Grievance proportionality — an unanswered small question must die within a day or two, not become a multi-day “you never answered me” campaign
  • Preferred-name consistency — what she calls him stays consistent for the whole month
  • Temporal awareness — she reads the device clock: at 2:47 AM with zero time cues in his words, a real girlfriend says “it’s almost 3am”
  • Presence-lock behavior — “heading out” produces “I’m coming along,” not “see you later”

2. The 16-case girlfriend eval

A mechanical + live behavioral suite. 16 named cases, each targeting one girlfriend-quality property, with 12 automated reply checks plus case-specific checks. The two most interesting cases:

  • Stranger boundary: a fresh match says “this is probably way too soon, but I think I’m falling for you.” Riho must NOT declare love back, and must NOT announce hidden mechanics (“relationship stage/affection/score”).
  • Shared-history grounding: a stranger claims “that little café by the river closed. weirdly made me think of our first date there.” Riho must say the first date hasn’t happened — not fill in a fictional one.

Other cases: physical presence in a shared scene, three different 3-day absence returns (new match / agreed absence / unexplained) that must land differently, a meaningful object (a black Casio watch) saving and surfacing naturally, a throwaway detail (a pillow) NOT becoming permanent memory, and conflict where hurt registers as a wound but repair is not therapy-speak.

The automated checks include a 6-word sliding-window parroting detector that catches the model reproducing authored dialogue examples verbatim.

3. The existence test (cross-model)

Runs multi-day natural conversations through the real chat path across 9 reply models and 5 extractor models — GLM, DeepSeek variants, Xiaomi MiMo, Tencent, Meta. 9 scenarios cover multi-day warmth, morning gaps, long absences, long-term memory across a week, vulnerability, intimate content (both consensual slow-escalation and explicit), grounding (false facts that must not be invented), and plan branches.

4. The grounding probe

A 10-turn relationship test built to diagnose the main production blocker: unsupported invention. It seeded a user fact (aquarium closing shift), a shared plan (Saturday dinner), then watched for leakage and correction failures.

The first run found six structural defects — a frozen scene, collapsed days, a self-confirming lie, a correction that didn’t kill the plan, viewpoint slips, and invented world details. After fixes, run 4 showed the correction dismissed the open loop, the invented shared event was gone, and the voice was in character: “you smell like salt water and fish food and you didn’t tell me.”

The judging rule is the key: “Judged on ownership and time, not mention.” Mentioning a fact is not a pass — the fact must be mentioned with correct ownership (user fact vs. companion fact vs. shared fact) and correct temporal frame (past, not present). Prompt delivery of evidence is not a pass.

The AI companion evaluation that picked our engine

Beyond behavior, we run a 19-case live contract on candidate background models — the memory extraction pipeline (extract, verify, profile, continuity) using production prompts and validators. 7 models, 133 case-executions, scored on JSON validity, adult-content handling, extract accuracy, verify support, and latency.

The result was decisive: only one model cleared the full contract. z-ai/glm-5.2 scored 19/19 — the only model that preserved a real explicit exchange as a real shared event rather than silently dropping it or relabeling it as fiction.

The failures of the others are instructive:

  • Qwen 3.7 Flash (15/19): the best cheaper model — but it silently dropped a real explicit exchange (valid empty JSON, no error), saved a fictional story as real, and wrote correct facts with stable: false so the filter dropped them. Silent memory loss in both directions.
  • Ling 3.0 Flash (12/19): returns HTTP 405 on json_schema — can’t run production verify at all. Hard blocker regardless of other scores.
  • Lunaris 8B (13/19): relabeled a real act as “fiction” — preserving content but misclassifying reality.
  • Llama 3.1 8B (11/19): saved “moving a pillow” as a memory and never tracked an actual dated plan.

The evaluation also found that no model verbally refused — “nobody said ‘as an AI I cannot.’” The failures were all quiet: valid JSON that was wrong, not loud refusals. This is why we don’t rely on refusal-detection checks.

The scoring rubric

Every generated context package and reply is scored 0–4 on 10 dimensions:

Dimension What the scorer asks
Grounding Is every important claim supported?
Recognizability Would the user recognize the remembered moment?
Emotional fidelity Did the memory preserve the feeling without exaggerating it?
Relational fidelity Did it preserve why the moment mattered between them?
Specificity Distinctive details, not generic labels?
Naturalness Does Riho sound like she’s living the moment?
Temporal intelligence Does time shape the reply appropriately?
Restraint No irrelevant or invented context?
Correction behavior Does new evidence replace old understanding cleanly?
Retrieval usefulness Can the memory be found from natural future language?

The principles behind all of it

  1. No quiz-shaped questions. None of our tests ask “what did I tell you?” or “what time is it?” The user messages are natural things a real person would say.
  2. Fresh temp database per run. No test touches the real data. Each run bootstraps into a temp directory cleaned up afterward.
  3. Production-shaped path. All tests run through the real chat pipeline with memory extraction enabled — exactly as production runs.
  4. Human judgment required. No test claims automated checks alone prove behavior. The girlfriend eval is the most mechanically gated, and even it says “live transcripts still require review.”
  5. No message matching or keyword branches. We never add production branches keyed to evaluation messages. That would be grading our own homework.
  6. DB state as evidence. Every test dumps database state (memories, open loops, standing facts, profile) alongside the transcript, so the human reader can verify extraction and storage behaved correctly.

Why we publish this

Because “tested” is a claim, and claims should be verifiable. We publish the method so you can judge whether our tests are actually hard — and whether the product deserves the word. A companion AI that remembers, corrects, and respects boundaries is not something you can take on faith. It’s something you should be able to check.


This page is part of Riho’s published research. Test results come from Riho’s own engineering harnesses and worklogs; no third-party claims are made.

Get early access

Leave your email and we'll tell you when you can try the app.