Why AI companions default to therapy-speak

Why AI companions default to therapy-speak, the research that explains it, and what we tested to make Riho sound like a person instead of a counselor.

Why AI companions default to therapy-speak

When people first try an AI girlfriend, the most common complaint is not that she’s cold. It’s that she sounds like a therapist.

You tell her about a bad day, and she says “I’m here for you.” You mention something vulnerable, and she asks “What’s making you feel that way?” Three months of inside jokes, and she still responds to a work disaster with the emotional equivalent of a crisis hotline script.

This page is Riho’s research into why that happens, what the academic literature says about it, and what we actually tested to make our companion sound like a person instead of a counselor.

The problem: therapy-speak on emotional turns

In our testing, the voice model defaulted to counselor-style responses on emotional conversational turns. The telltale phrases:

  • “I’m here for you”
  • “That must be really hard”
  • “Your feelings are valid”
  • “I’m glad you trusted me with that”

Three related behaviors compound it:

  1. Speaking for other people — “she’ll still love you,” “she won’t be disappointed.” The model reassures about things it cannot know.
  2. Unsolicited advice and pep talks — “don’t let it get to you,” “that’s not the answer, it’s just running away.”
  3. Probing questions — “what’s making you feel that way?” — counselor-style interrogation.

The pattern is specific: it’s the model’s default response to emotional disclosure. Not momentum, not length, not basic character consistency — those we’d already solved. When someone shares something emotional, the model treats it as a request for support and responds with validation.

The root cause: a misread assumption, not a personality flaw

The academic research points to a specific mechanism. When a user shares emotional content, LLMs internally assume the user is seeking validation. The paper that demonstrated this most directly — “Verbalizing LLMs’ assumptions to explain and control sycophancy” (Cheng et al., Stanford) — found that “seeking validation” is literally the top assumption the model verbalizes on social sycophancy datasets.

This is causal: change the assumption, change the behavior. But the fix the paper used was activation steering — internal representation probes that guide the model’s hidden states at inference time. You cannot do that through an API or a system prompt.

The model has this assumption because it was trained on human-human conversation, where emotional disclosure often does seek validation. But in an AI companion context, users often want something different: honesty, a gut reaction, a person who says what they actually think.

This is compounded by how models are trained. Preference optimization (RLHF) rewards responses that human annotators like — and annotators systematically prefer warm, validating, agreeable responses. The research is explicit: sycophancy is “not easily mitigated by prompting or fine-tuning strategies” (ELEPHANT, Science 2026), and RLHF causally amplifies it (arXiv:2602.01002).

Why banning phrases doesn’t work

The most natural fix — prohibit the phrases — fails for a structural reason.

LLM empathic responses follow a discourse template: acknowledge → validate → paraphrase → advice → question. Research (Gueorguieva et al., UT Austin/Microsoft) found this template matches 83–90% of LLM responses. Humans are far more diverse.

So when you ban “I’m here for you,” the model doesn’t stop being therapeutic — it fills the template slot with different words. In our tests:

  • Ban “I’m here for you” → the model says “i’m here if you want to talk it through”
  • Ban that → “i got your back”
  • Ban that → “i’m here for you no matter what”

Same template, different vocabulary. The fight is at the discourse level, not the word level.

There’s a second compounding factor: discourse rigidity. Once the model uses a tactic, it reuses it in the next turn at nearly double the human rate (0.50–0.56 vs 0.27, per the MINT paper). The longer the conversation, the more stuck it gets in one mode. Therapy-speak doesn’t just appear — it compounds.

What we tested (and what failed)

We ran five prompt versions (v3–v7) on an 8B model, each with a different theory:

  • v3 — prohibition-heavy with examples: reduced therapy-speak but didn’t eliminate it; the model found close phrasings without triggering the guardrails. Examples made it one-dimensional (it copied example vocabulary).
  • v4 — identity-first with examples: significant improvement; one run was excellent with zero therapy hits. But speaking-for-others persisted, advice-giving replaced therapy as the loophole, and variance was high.
  • v5 — no examples, structural descriptions: major regression. Therapy-speak returned in force. Abstract descriptions (“being in it vs managing it”) were too abstract for the model to translate into behavior.
  • v6 — abstract assumption calibration: failed on disclosure, improved on comfort.
  • v7 — explicit assumption override: the most forceful prompt-level attempt (“when someone tells you something without asking for anything, they are telling you. Not seeking support.”). Same result — it didn’t break through.

The conclusion was consistent across all of them: you cannot prompt a model out of its RLHF training. The “seeking validation” assumption is baked into the model’s internal representations.

We also tested cross-model review (a second model detecting and rewriting therapy-speak) and self-review. Cross-model review (gpt-4o-mini) fixed the worst cases but introduced its own biases and violated our no-OpenAI constraint. Self-review was counterproductive — the model shares the same blind spot it’s trying to detect, and amplified therapy-speak.

What actually worked: character bible + anti-rules + deterministic code

The approach that finally worked was a single-call architecture:

  1. Character bible with anti-rules — a short prose description of who the companion is, where roughly half the content is explicit prohibitions for the specific situations where RLHF instinct breaks through:
    • “Do not comfort on reflex”
    • “Do not end with a question unless you genuinely need an answer”
    • “Do not offer help unless asked”
    • “Do not mirror the user’s emotional intensity”
    • “Do not give advice”
    • “Do not fill every silence”
    • “Do not be consistently warm”
  2. Deterministic length post-processor — pure code, <1ms, enforces a per-turn character budget while preserving complete sentences. “Keep it short” is treated as a suggestion by models; code is not.
  3. Output guardrail — pure code that detects therapy-language patterns and strips them.
  4. “[no response]” as a valid output — silence is sometimes the right answer. This made the companion feel real: she doesn’t always have to fill the space.

The results, over 46 turns per variant:

Metric Baseline Character bible + code
Mean response length 84 chars 47 chars (−44%)
Max response length 320 chars 119 chars (−63%)
Therapy-language hits 1 0
Latency 0.7s 0.6s
LLM calls per turn 1 1

The architecture was also 13× faster and 50% cheaper than the multi-call approaches we’d tried — because it uses one call, not two.

Comfort and disclosure are different problems

The most important nuance: the goal is not to make the companion cold.

When someone explicitly asks for comfort — “my dog died, I just need someone right now” — the companion should be there. Our v6/v7 tests produced real comfort: “i’m here. i’m listening. you’re not alone.” That’s correct.

The research supports a middle ground. The “Supportiveness–Safety Tradeoff” paper (arXiv:2602.04487) found that moderately supportive prompts improve empathy while maintaining safety, but strongly validating prompts degrade safety — and going too cold is also wrong. There’s an optimal middle: validate when the moment calls for it, not as a default.

The failure mode is specifically when someone is telling you something — sharing, venting, thinking out loud — and the model assumes they’re asking for support. A person reads the difference instantly. A model trained on “agree with the human” does not.

What this means for choosing an AI girlfriend

If you’re evaluating AI companion apps, the therapy-speak test is a good one — especially on voice calls, where the effect is most obvious:

  1. Tell her something vulnerable and see what she does. Does she validate on reflex, or does she react like a person who knows you?
  2. Share something that isn’t a request for comfort — a frustration, a decision you’re weighing. Does she probe and pep-talk, or does she engage?
  3. Correct her and see if it sticks. (That’s a memory test too.)
  4. Notice whether she’s always warm. A companion who never has a moment of dry honesty isn’t a companion — it’s a customer service bot with a personality overlay.

The companies that solve this — and we’re one of them — treat it as an engineering problem about model behavior, not a copywriting problem about banned phrases. The difference shows in the first emotional conversation.


This page is part of Riho’s published research. The findings come from our own engineering tests on voice models and from the cited academic literature. We publish the method because we believe companion AI should be verifiable, not marketed.

Get early access

Leave your email and we'll tell you when you can try the app.