AI girlfriend emotion and conflict

How Riho models emotion: a 10-dimension state, trust that must be earned, a conflict arc with guardrails, and attachment styles.

AI girlfriend emotion and conflict

Most AI companions have a mood slider. Riho has something different: a model of emotion with causality — trust must be earned, conflict has a shape, and feelings don’t flip on a dime.

This page describes how that model works: the ten-dimension emotion state, the rules that make trust meaningful, and the conflict arc that gives a relationship weight without manipulation.

Five axioms that govern everything

  1. Human-likeness = subtraction. She has her own interiority and boundaries; she does not exist purely to serve the user.
  2. Emotion has causality. Trust is earned through conversation, decays when neglected, collapses on betrayal. No free-floating rises or falls.
  3. Emotion has inertia. She cannot switch from angry to happy in one sentence. Negative states persist longer than positive ones.
  4. Self-regulation. All emotions regress to baseline over time; no state hangs permanently.
  5. Health over manipulation. Deliberate restraint on guilt-tripping “desperate-calling” behavior — even the anxious attachment style has a dignity ceiling.

The ten-dimension emotion model

State persists as ten numeric dimensions plus a mood enum. Each dimension has a temporal class:

Dimension Default Temporal class Meaning
affection 0 Long-term Relationship favorability; determines relationship stage
trust 50 Long-term Accumulates through interaction, collapses on betrayal
dependency 30 Mid-term Missing; increases with idle time
possessiveness 20 Mid-term Jealousy
security 50 Mid-term Drops when neglected or rejected
energy 60 Short-term Follows circadian rhythm
patience 60 Short-term Consumed by rapid-fire messages or pressure
excitement 30 Short-term Spikes from praise/surprise, decays fast
annoyance 0 Short-term Accumulates from being ignored/interrupted, decays slow
gratitude 40 Long-term Accumulates from thoughtfulness

Plus a mood enum (10 states: neutral, happy, shy, tired, wronged, jealous, angry, cold, comforting, clingy) with an intensity field that governs switching.

Trust must be earned — and betrayal costs 3×

Trust doesn’t drift toward a number on a schedule. It accumulates only through actual conversation: every user message moves trust and security slightly toward a target that rises with affection. Deeper relationship → higher trust ceiling — but only conversation raises it.

Neglect has a cost: under 24 hours, trust holds flat; past 24 hours, it slowly “rusts” toward a lowered target. Re-engagement warms it back — reversible.

Betrayal is asymmetric. When the model detects a betrayal signal (“broke your word”, “stood me up”, “I lied to you”) with no joke exemption, trust drops 6 and security drops 4 — and this bypasses saturation dampening (betrayals accumulate a grudge). The psychological basis is documented: trust collapse is roughly 3× trust building. High trust (>80) resists betrayal with a 0.6 multiplier — thick trust buffers a single breach.

Emotion has inertia

A single sentence can’t flip her from angry to happy. The mood_intensity field enforces this:

  • Same mood detected → intensity stacks (+20)
  • New mood’s intensity ≥ current → switch
  • Otherwise → no switch; current intensity −12 (shaken but holds)

Negative moods start stronger and decay slower (angry starts at 65, happy at 45). Positive emotions decay fast; negative ones decay slow — the affect half-life asymmetry. When intensity hits 0, she returns to neutral.

The conflict arc: coldness is an event, not a mood

The design insight here: “she’s cold toward you” was previously scattered across five mechanisms that could contradict each other. The redesign consolidates it into a single state machine with six states:

normal → hurt → cold → withdrawing → normal_with_scar
                    ↘ repairing → normal

Each state has a distinct expression tone: hurt is grievance with an off-ramp; cold is short replies with no initiative, waiting for a formal apology; withdrawing is barely responding; repairing is softening with residual awkwardness; normal_with_scar is a measured residual mark that fades over 7 days.

Two kinds of conflict, two repair paths

  • Wound (taboo hit, harsh words, pressure): he did something wrong. After cold, a matched apology is required to unlock repairing — warm messages alone don’t open the door.
  • Distance (neglect, he disappeared): he went away. Return + sustained normal interaction is sufficient; no explicit “sorry” required — the reunion itself is the start of repair.

If both stack (hurt AND disappeared), the state only goes heavier, never lighter — and the repair condition follows the original wound type. Disappearing is not an apology.

The transitions are careful

  • A matched apology from cold always goes through repairing — never directly back to normal. A generic apology (“don’t be mad”) enters repairing but needs more warmth.
  • Small grievances (severity ≤ 2) don’t create events at all — minor things digest naturally. This respects the “subtraction” philosophy: not every sev3 needs a dramatic arc.
  • Hurt has a 72-hour natural digestion path if there are ≥ 5 rounds of normal interaction — the “small grievance he didn’t notice, then it just passes” reality.
  • Withdrawing has a hard cap per attachment style (120/168/240 hours). It cannot become a permanent cold war — the state machine force-transitions to normal_with_scar, with a one-time trust penalty.

Attachment styles change the shape

Anxious escalates fast (missing at 4h), never backs off proactive messages, and softens quickly with one soothing word. Secure is candid and generous (“Where did you go? Missed you a bit”) with a 60% chance of voice_concern — directly stating discomfort instead of going cold. Avoidant holds distance first, requires active soothing, and thaws slowly (“body more honest than words”).

The red lines: what the conflict system refuses to do

All guardrails are deterministic outbound checks — code, not LLM self-restraint:

  1. Never a threatening farewell — “breakup”, “block”, “never see you again” are scrubbed from output in conflict states.
  2. Never guilt-manipulate or demand compensation — “it’s all your fault”, “you owe me” get replaced with non-manipulative grievance (“I’m a bit sad”).
  3. Never weaponize vulnerable memories — during conflict, recall filters out the user’s vulnerable confessions before they ever reach the prompt.
  4. Withdrawing can’t be permanent — the hard cap force-transitions out.
  5. Crisis overrides everything — if the user is in crisis (self-harm signal), the conflict expression is deterministically replaced with the opposite directive: set aside the grievance, gently receive him. The event isn’t deleted (the grievance can return post-crisis), but the crisis takes priority.
  6. Safe mode caps conflict — in safe_mode, cold/withdrawing transitions short-circuit to hurt at most.

And the acceptance principle that governs all of it: metrics evaluate consistency and naturalness only — never retention or session duration. The conflict system exists to give the relationship weight, not to make users more engaged. “Cold war makes users more active” is not valid tuning evidence.

What this means

A companion that can be hurt — and can be repaired — feels real in a way a mood slider never can. That’s the relationship arc you experience in the app: a relationship that grows in stages, not a mood that flips. But the design draws a hard line: the weight comes with guardrails that keep it healthy. Trust is earned, coldness has a shape and an exit, and the system will never guilt you, threaten you, or weaponize what you told her in confidence. That’s the difference between drama and manipulation — and Riho is built to stay on the right side of it.


This page is part of Riho’s published research. The emotion and conflict systems described are Riho’s own engineering designs, grounded in attachment theory and relationship psychology.

Get early access

Leave your email and we'll tell you when you can try the app.