AI memory architecture in Riho

How Riho's memory system works: five context layers, evidence-backed memories, deterministic promotion, and why every memory carries a receipt.

AI memory architecture in Riho

An AI girlfriend that remembers you is not a bigger database. It’s a set of decisions about what deserves to be remembered, how a memory earns its place, and how the companion receives it.

This page is the engineering version of that statement: how Riho’s memory system is designed, from the five context layers down to the evidence rules that decide whether a memory sticks.

The design problem

A companion relationship can eventually contain millions of transcript tokens, but the voice model can only see 8,192 tokens at a time. The system has to decide, for every turn, which slice of the relationship the companion should see.

The design principle is: don’t create a table for every psychological concept. Give each layer one clear responsibility, and preserve evidence across every transformation. If a memory can’t be traced back to something that was actually said, it shouldn’t be in the prompt.

The five context layers

1. Permanent identity

Who you are — names, preferred form of address, pronouns, relationship framing, timezone, relationship start date. Always available, never dependent on semantic search. Your date of birth stays private to eligibility logic and is only surfaced when specifically relevant.

2. Immediate raw context (the live tail)

The newest transcript that fits the prompt budget — up to 20 message rows, chronological, with the current message last. Responsible for exact wording, tone, and immediate topic changes. No extraction runs on raw turns.

3. Current continuity and archived segments

Exactly one continuity_state row is the current “notebook page” per companion — a 600–800 token snapshot of the medium-term story, selected directly (never by semantic search) and always included after its first accepted promotion. Prior snapshots stay immutable for audit but are not appended to the prompt.

Older transcript ranges get compact archived summaries. These are explicitly “lossy navigation aids backed by the transcript” — not relationship truth — and are only used when relevance-gated retrieval justifies it.

4. Active temporal context

What’s happening now, soon, or still unresolved: events, plans, deadlines, appointments, ongoing situations, open threads. Selection is by status, time, exact people, and topic — embedding similarity is optional support, never the only path.

The lifecycle matters: a plan that passes its expected end becomes outcome_unknown, not resolved. It only resolves when the user actually confirms, changes, or cancels it. If the outcome is meaningful, a durable episode is created with both the setup and outcome evidence.

5. Durable relationship memory

User facts, preferences, boundaries, patterns, important people, episodes, shared moments, rituals. Active plans and events don’t begin here — a meaningful resolved outcome may later become an episode.

The evidence rule: memories need receipts

This is the heart of the system. Every accepted memory must include at least one memory_sources record containing: the source message ID, the role, the exact quote, the quote position, and the evidence purpose.

The rule is enforced at the database level: the source quote must occur in the cited message after conservative normalization. If it doesn’t, the candidate is rejected. A memory that can’t point at what you actually said doesn’t exist.

There’s also an ownership rule: a direct memory must have at least one user-authored source. The companion’s own statements can support companion or shared memory, but cannot establish a new user fact by itself. This is what prevents the self-confirming lie — where the companion’s own restatement of an event becomes the evidence for it.

Corrections and supersession

When you correct Riho — “no, it’s my brother who lives in Austin, not my sister” — the system doesn’t edit the old record. It:

  1. Preserves the old memory and marks it superseded
  2. Links the replacement
  3. Makes only the replacement prompt-eligible
  4. Retains all source evidence for both

Deduplication uses a stable semantic key (scope + subject + kind + category + subcategory + attribute), not vector distance. Two facts aren’t the same because their embeddings are close — they’re the same because they target the same key.

Patterns require repeated evidence. One occurrence may remain an episode or a tentative observation, but it does not become “the user always…”

The promotion pipeline

A background worker promotes older transcript into derived records in three stages:

  • Stage A — continuity: produces the concise source-ranged continuity summary. No relationship scores, no personality analysis.
  • Stage B — structured extraction: produces memory candidates, people updates, temporal candidates, and correction intentions, each with exact source quotes and message IDs.
  • Stage C — deterministic application: code (not the LLM) validates the JSON schema, validates source membership and exact quotes, constructs stable keys, applies corrections and supersession, and stores the records.

The key move is Stage C: the LLM proposes, but code disposes. Everything the model says must survive deterministic validation before it touches the database.

Retrieval: what the companion actually sees

Per turn, the system assembles, in order:

  1. Permanent identity
  2. Deterministic local-time context
  3. Active temporal items by time and status
  4. Exact people and alias matches
  5. The one current continuity snapshot
  6. Relevant communication preferences or boundaries
  7. Hybrid full-text + vector retrieval of at most a few durable memories
  8. Recent raw transcript
  9. Current user message

With baseline limits: 0–3 durable memories, 0–1 exact person cards, exactly 1 continuity snapshot, 0–2 archived segment summaries. Overlap suppression ensures the same information never appears as raw transcript, summary, and memory simultaneously.

The live implementation

The current production backend simplifies the target schema into one after-reply extraction job with four outputs: memories, unfinished commitments, standing knowledge, and her_note (a ≤400-char first-person note of what she’s carrying between exchanges).

Two design choices stand out:

  • Importance floor with ownership awareness: user and shared memories need importance ≥ 5, but companion emotion memories pass at ≥ 3. This asymmetry exists because the companion’s own feelings must be stored, not just user facts — a direct response to the “user CRM” failure mode (a companion that treats you as a database of preferences to manage).
  • No hidden relationship scores. There is no generated relationship-state paragraph and no closeness model. Familiarity emerges from relationship age, actual conversation, accumulated knowledge, rituals, and shared moments. Silence alone never reduces the relationship or triggers guilt.

What this means

The architecture exists to make one thing true: Riho can show her work. This is what powers the memory feature you interact with every day. Every memory she acts on traces to a quote you actually said, every correction supersedes cleanly, and nothing enters the prompt that can’t be verified. That’s the difference between a companion that remembers and a database with a voice.


This page is part of Riho’s published research. The architecture described is Riho’s own engineering design, based on internal design documents and the live implementation.

Get early access

Leave your email and we'll tell you when you can try the app.