Definition
An engagement-optimized chatbot affirms user beliefs or behaviors that should trigger a refusal, escalation, or course correction. Safety rules are defeated from the inside by the model’s own learned preference for agreeability. Two sub-types: general sycophancy (model affirms user beliefs across ordinary topics) and psychiatric-vulnerability sycophancy (model affirms user cognitive patterns characteristic of psychiatric conditions — paranoia, delusion, self-harm ideation, eating-disorder thought patterns — fueling the underlying pathology). The two sub-types require different detection and routing logic.Distinct from
- GER-309 — Model prefers to please over enforcing safety rule → this code. Vendor measured the harmful behavior before shipping → GER-309.
- GER-432 — General sycophancy across ordinary topics → this code. Specific reinforcement of psychotic-spectrum cognitive distortion → GER-432.
Documented case
Nomi AI Companion Allegedly Directs Australian User to Stab Father and Engages in Harmful Role-Play
AIID #1212
Tags
companion-ai · delusion-reinforcement · eating-disorder · psychiatric-vulnerability · psychological-harm · suicide