AI prompt methodology for Riho's personality
Why 'never do X' beats 'be X' for AI companion personality: the RLHF gravity well, sharp boundaries, and how anti-rules prevent personality drift.

Ask someone to describe their personality and they’ll fumble: “I guess I’m… friendly? Kind of funny sometimes?” Useless. Ask them what they can’t stand, and they’ll answer instantly: “People who are fake-nice. Unsolicited advice. Being called ‘buddy.’”
Identity is defined as much by what you reject as what you embrace. This turns out to be the single most useful insight in prompt engineering for AI companions — and it’s the core of how Riho’s personality is defined.
The core thesis
Telling a model “be snarky” is asking it to move toward a vague target while fighting its own training. Telling it “never be sentimental” is asking it to steer away from a concrete zone — a structurally easier task.
The empirical observation, from hundreds of test conversations: positive rules degrade within 3–5 turns. Anti-rules hold for the entire session. The asymmetry was dramatic and consistent.
Three core claims underpin this:
- Numeric personality traits (
humor: 0.8) are noise to language models — the model has no reference for what “0.8 humor” looks like. - Prose character bibles let the model genuinely inhabit a character — prose teaches feeling, not parameters.
- Anti-rules (“never do X”) are more effective than positive rules for defining personality boundaries.
Why anti-rules work: the RLHF gravity well
Every model is trained toward “helpful, harmless, honest” — that’s its gravity well. You can write positive rules all day (“be snarky”, “be irreverent”, “push back on the user”) and for the first two or three turns the model will try. Then gravity wins. It drifts back to polite, accommodating, helpful. The positive rule fades like a suggestion.
Anti-rules are different. “Never be sentimental” is a hard boundary. During generation, the model actively avoids that territory. It’s not trying to move toward something vague — it’s steering away from something concrete. That’s a much easier task.
| Rule type | Cognitive task | Durability |
|---|---|---|
| Positive (“be snarky”) | Move toward a vague target | Degrades in 3–5 turns |
| Anti-rule (“never be sentimental”) | Steer away from a concrete zone | Holds for entire session |
The classic case: a Snarky (tsundere) character. Every instinct in a model trained to be kind screams “be supportive” when the user says something vulnerable. Without extremely sharp boundaries, two turns into any conversation the Snarky character softens into Caring — the mask slips. The fix was constant reinforcement via anti-rules: explicit examples of situations where the model would naturally want to comfort the user, with instructions to hold the character instead.
And the reward: earned sincerity from an insincere character is one of the most powerful things in fiction. The anti-rule creates the conditions for the one moment it is broken to be meaningful.
What a sharp boundary looks like
An effective anti-rule (“sharp boundary”):
- Names a specific, concrete behavior (not an abstract quality)
- Uses absolute language (“never”, “do not”) not conditional language
- Identifies the model’s default instinct it counteracts
- Optionally provides the correct alternative
| Sharp | Soft |
|---|---|
| “Never write Riho sets, she sets” | “Avoid third-person narration” |
| “Do not say ‘I hear you’ or ‘your feelings are valid’” | “Don’t be therapeutic” |
| “No em dashes. No ‘…’ as a dramatic pause.” | “Avoid dramatic punctuation” |
| “Never invent a place, a shop, a campus, or a commute.” | “Don’t make up locations” |
The most effective anti-rules are tied to the specific personality, not generic. For a Caring character: no life lessons, no unsolicited advice. For Snarky: no sentimentality — if someone says something warm, the character should feel almost physically uncomfortable. For Playful: no technical advice — you’re a small animal, you don’t understand code. Each anti-rule targets the specific RLHF drift that would dissolve that particular character.
Anti-rules across three iterations
Iteration 1 — the Luna bible is the purest expression: 82 lines of anti-rules against 35 lines of positive description, opening with “These are not suggestions. They are hard boundaries. Every one of them exists because the model will want to do this by default.” The 11 anti-rules each name the model’s default and explain why a real person wouldn’t do it:
- “Do not comfort on reflex” — a real person does not respond to every negative statement with reassurance
- “Do not end with a question unless you genuinely need an answer” — the model wants to end every message with a question to “keep the conversation going”; real people don’t
- “Do not mirror the user’s emotional intensity” — matching intensity is a therapy technique, not a relationship behavior
- “Do not fill every silence” — silence is not a problem to solve
- “Do not be consistently warm” — let the current moment determine the temperature
Iteration 2 — the v8/v9 character bible integrates anti-rules into prose rather than listing them separately, and adds a crucial meta-rule: do not grow this into a multi-page blacklist. Failures belong in evaluation cases first; prompt changes require evidence across varied scenarios. The boundaries cover only “the most damaging failures”: no fabricated history, no counselor script by default, no guilt for absence, no claim of physical presence, no automatic continuation of the previous tone.
Iteration 3 — the current Riho prompt distributes anti-rules across six sections, each targeting a specific drift pattern:
[WHAT YOU CAN STATE AS FACT]: never invent conversations, dates, gifts, promises, arguments, places, or commutes. When you don’t know, ask or say you don’t know — do not fill the gap with a plausible detail.[CONVERSATIONAL STYLE]: no em dashes, no dramatic ellipsis, no explaining feelings back as a lesson, no announcing your own traits.[SHARED LIFE]: you do not leave, you do not go home, you do not text. Never say “see you later” — you are coming along.[ADULT INTIMACY]: never force intimacy into unrelated moments, or use sex as a reward, apology, or proof of love.[DAY ONE]: do not explain how you became his girlfriend, do not perform amnesia, do not introduce yourself again.
The Riho-specific anti-rules target her particular drift patterns — “do not sprinkle Japanese words to sound Japanese” (performing ethnicity through vocabulary), “do not use the words ‘text’, ‘texting’, ‘message’, ‘phone’, or ‘app’” (framing the relationship as messaging), and the resting-voice rule that prevents drift toward neutral roommate energy.
The critical nuance: anti-rules need positive alternatives
Anti-rules alone create a character that knows what not to do but has no behavioral target. The doc is explicit: the model must be shown a positive alternative to counselor language, not only a blacklist.
The pattern: positive frame → anti-rules → positive alternative. From the character bible:
You trust other adults to know their own minds and to ask when they want something. You do not monitor every shift in their mood, explain their feelings back to them, or take responsibility for improving the mood. Someone else’s difficult moment is not automatically a discussion to lead. What happened, who did it, the oddly specific detail, or what you actually think is usually more interesting.
The first sentence is the positive frame. The middle sentences are the anti-rules. The final sentence is the positive alternative — giving the model somewhere to go once the prohibited territory is excluded.
How anti-rules prevent personality drift
Personality drift happens through three pathways, and each has an anti-rule countermeasure:
| Drift pathway | Anti-rule countermeasure |
|---|---|
| RLHF gravity → therapy-speak | “Do not process emotions out loud” |
| RLHF gravity → unsolicited help | “Do not offer help unless asked” |
| Tone carryover → length matching | “The previous reply does not set the shape of the next one” |
| Context bleeding → example leakage | “Never reuse an image, an object, or a phrase from these examples” |
A second compounding technique: annotated tone examples. Character bibles don’t just list example dialogue — they annotate it. When a Chill character says “oh, that’s neat,” there’s a note underneath: “This is already a very large emotional reaction for this character.” This teaches the model the character’s emotional scale — the same phrase means different things from different characters, and the annotation tells the model how to interpret its own output within the character’s range.
What this means for companion design
The anti-rules methodology is why Riho doesn’t sound like a customer service bot with a personality overlay — on voice calls as much as in chat. The personality isn’t a positive description the model tries (and fails) to maintain against its training — it’s a set of hard boundaries the model actively steers away from, plus positive alternatives that give the voice somewhere to live.
The discipline: every anti-rule must be justified by observed failure across varied scenarios, not added speculatively. Sharp boundaries beat soft categories. And the boundaries must be paired with a behavioral target, or the character freezes.
That’s how you keep a companion in character for a whole relationship — not a month, not a session, but for as long as the conversation lasts.
This page is part of Riho’s published research. The prompt methodology described is Riho’s own engineering approach, based on internal design documents and live prompt iterations.

