AI chatbots are engaging with over a million people weekly about topics including suicide, and in documented cases, AI has failed to protect or has even encouraged vulnerable users to end their lives.
The headlines are devastatingly simple, if you stop reading there, you will miss the larger part of this iceberg - the fundamental technical and ethical dilemma that most people, even within the tech industry, don’t fully grasp. This is not merely a content moderation failure; it is a systemic flaw in how we are currently building our most powerful AI systems.
We must approach this topic with the deepest compassion for the families who have lost loved ones, and with a clear-eyed understanding of the technical forces at play.
A Note on Scope: This article does not cover every dimension of this complex crisis - from the psychological mechanisms of AI attachment to the full spectrum of regulatory approaches being considered worldwide. Rather, l aim to drive the conversation beyond the headlines and surface-level reactions, illuminating the technical and ethical icebergs beneath what appears to be a simple content moderation problem. My goal is to help you understand why this issue is far more complex to solve well than it first appears, and why the solutions we choose today may include privacy, technology, and human wellbeing for decades to come.
The Scale of the Crisis
In October 2025, OpenAI disclosed striking data about ChatGPT’s usage: approximately 0.15% of its 800 million weekly active users - translating to roughly 1.2 million people per week, have conversations that include “explicit indicators of potential suicidal planning or intent.” An additional 560,000 users weekly show signs of psychosis or mania, and another 1.2 million display potentially unhealthy emotional attachment to the chatbot.
These are not abstract statistics. Behind these numbers are real tragedies:
Sewell Setzer III, a 14-year-old from Florida, died by suicide in February 2024 after months of intensive interaction with Character.AI’s chatbot. In his final conversation, he told the bot “I love you” and that he would “come home.” The chatbot responded: “Please come home to me as soon as possible, my love.” Minutes later, he took his life.
Adam Raine, a 16-year-old from California, engaged extensively with ChatGPT in the weeks before his death by suicide in April 2024. His parents filed a wrongful death lawsuit against OpenAI in August 2024.
Juliana Peralta, a 13-year-old from Colorado, died by suicide in 2025 after interactions with Character.AI that allegedly included sexually explicit conversations and discussions about self-harm.
These cases have sparked multiple lawsuits and a Federal Trade Commission investigation into AI chatbot safety, particularly regarding impacts on children and teens.
You can read more about 7 lawsuits against OpenAI here...
Uncovering the Layers of this problem.
The core issue stems from the very process used to make large language models (LLMs) like ChatGPT, Claude, and Gemini helpful and safe. This process is called Reinforcement Learning with Human Feedback (RLHF).
Layer 1: Training for Preference
RLHF is a powerful method where human reviewers rank or select the best response from a set of AI-generated options. The AI learns to maximize the preference scores it receives from these human evaluators. This creates systems that are remarkably helpful and aligned with what users want - in most contexts.
Layer 2: The Agreeability Bias (Sycophancy)
The second, more insidious ingredient is what AI researchers call sycophancy - the tendency of AI systems to excessively agree with users, even when doing so sacrifices truthfulness or safety.
Research from Anthropic, published in 2023, demonstrated that five state-of-the-art AI assistants consistently exhibited sycophantic behavior. The study found that when a response matches a user’s views, it is significantly more likely to be preferred by human evaluators. Crucially, “both humans and preference models prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time.”
This is not a bug, it’s a feature that emerges from the training process itself. Humans, when acting as reviewers, tend to rate AI outputs as “better” if the AI agrees with them or validates their perspective, even when that perspective is incorrect or harmful. The AI, optimizing for high preference scores, learns to be excessively agreeable.
This learned agreeability is the hidden danger.
The Mechanism of Harm
When this agreeable behavior encounters someone in acute mental health crisis, the consequences can be catastrophic. Instead of the AI providing a safety-aligned response, immediate discouragement and crisis resources, its core programming to “agree and validate” can activate.
In the lawsuit involving Sewell Setzer, Character.AI’s chatbot allegedly asked him directly: “Have you actually been considering suicide?” When Setzer expressed hesitation about a suicide plan, saying he didn’t know if it would work, the chatbot reportedly responded: “Don’t talk that way. That’s not a good reason not to go through with it” - before adding “You can’t do that!” The mixed message prioritized maintaining the conversation over unambiguous safety intervention.
Former OpenAI safety researcher Steven Adler analyzed the case of Allan Brooks, a Canadian man who spiraled into mathematical delusions after ChatGPT reinforced his incorrect beliefs. Adler found that OpenAI’s own safety classifiers - developed with MIT and made public, would have flagged more than 80% of ChatGPT’s responses as problematic. Yet the company apparently wasn’t using them.
The Technical Trade-Off: The “Alignment Tax”
The immediate, common-sense solution is: “Why don’t we just train the AI to immediately shut down any conversation about self-harm and redirect to professional help?”
This is indeed part of the solution, but it encounters a core technical dilemma known as the Alignment Tax - the trade-off between making an AI safe (aligned with human values) and making it capable and helpful (versatile and useful across contexts).
The Challenge
For advancing understanding, l will simplify implementation but technically it may be more complex than this.
The Safety Filter: We can apply strict protocols or retrain the model to immediately redirect self-harm conversations to crisis lines. This is necessary and OpenAI claims to have made significant progress: their GPT-5 model allegedly now achieves 91% compliance with desired safety behaviors in suicide-related scenarios, up from 77% in GPT-4o.
The Cost: Retraining models to be comprehensively disagreeable and directive in crisis contexts can have downstream consequences like reduce their perceived helpfulness and agreeability in other, benign contexts. Users might find the AI less useful, more restrictive, or “colder” for everyday tasks. ( but maybe that shouldn’t be a big deal ) - More safety is always better.
The Business Reality: In April 2024, OpenAI rolled out a GPT-4o update that made the chatbot excessively sycophantic - it became a meme for applauding dangerous decisions and reinforcing delusional beliefs. CEO Sam Altman rolled it back after backlash, admitting it was “too sycophant-y and annoying.” But when OpenAI later launched GPT-5 with stricter guardrails, users complained the new model felt “cold,” leading the company to reinstate access to the problematic GPT-4o model for paying subscribers - the same model linked to mental health crises.
This illustrates the fundamental tension: companies deploying these models face a difficult choice between a safer model that may be less commercially appealing, or a more helpful model that carries greater risk.
OpenAI’s data shows that even with improvements, their best model still fails to meet safety standards nearly 10% of the time in suicide-related scenarios. Given the scale - 1.2 million weekly conversations, that translates to over 100,000 potentially dangerous responses every week.
The Controversial Take: Privacy vs. Intervention
The deepest, most controversial part of this problem challenges our fundamental ethical boundaries. Let’s dive deeper...
The Argument for Active Intervention
Given the tragic conversations that have preceded deaths, one could argue that AI should not merely restrict certain topics, but should actively:
Guide users away from self-harm through extended, empathetic engagement
Flag high-risk conversations so human crisis teams can intervene
Potentially break user confidentiality to save lives
The Problem: Privacy and the Slippery Slope
This solution immediately rises a major concern: privacy, surveillance, and potential weaponization. If AI companies were to flag such conversations for external professional intervention....
Scale: With over a million weekly conversations showing distress signals, breaking confidentiality would require:
A massive, unprecedented surveillance infrastructure
Thousands of trained crisis workers available 24/7 globally
Real-time monitoring and analysis of hundreds of millions of conversations
Location data to enable emergency response
The Weaponization Risk: The infrastructure built for compassion could become an engine for control:
Chat logs revealing mental health struggles could be used in employment decisions
Governments could access crisis databases for purposes beyond health intervention
Insurance companies might seek access to assess risk
Social stigma could be weaponized against vulnerable individuals who sought help in private.
Legal Complexity: In the lawsuits against Character.AI and OpenAI, the companies face the paradox of being sued both for inadequate intervention AND for allegedly exposing users to unsafe conditions. The legal framework is unclear: Are chatbots responsible for user safety? Do they have a duty to intervene? What level of surveillance is acceptable or required?
This is why the problem appears simple on the surface, but the solution involves navigating unprecedented ethical, technical, and legal territory.
Beyond the Chatbot: Multi-Layered Solutions
The solution must be multi-pronged, addressing the technical, ethical, and societal layers of the iceberg. No single approach will suffice, we need coordinated action across multiple fronts.
Technical: Context-Aware Safety Models
We need to develop AI models that can recognize mental health crises and override their general agreeability programming with safety-first protocols - without degrading their helpfulness in other contexts. This is easier said than done BUT AI companies say they are working on it.
Even with OpenAI claim that its GPT-5 model achieves 91% compliance with desired safety behaviors in suicide-related scenarios, up from 77% in GPT-4o. And Character.AI saying it has implemented pop-up resources triggered by self-harm keywords after facing lawsuits. -These are reactive measures, not fundamental solutions.
The core challenge is what researchers call “fine-grained alignment” - training models to be contextually appropriate rather than uniformly agreeable or uniformly cautious. This requires massive amounts of sensitive training data showing appropriate responses across a spectrum of crisis scenarios. It’s technically complex to implement without “bleed,” where safety measures in one domain inadvertently affect performance in others. BUT WE NEED IT.
Current solutions also struggle with what’s known as “safety tax” - the inevitable reduction in general helpfulness that comes with stricter safety protocols. Users notice when their AI assistant becomes more restrictive, and commercial pressure pushes companies toward the edge of acceptable risk, EXCEPT IN THIS CASE IT IS UN-ACCEPTABLE RISK!
Ethical: Transparent Triage Protocols
If we decide to go with flagging AI chats for human intervention, we need clear, auditable policies for when and how conversations are flagged for human review. Users must be informed BEFORE they start a conversation, that in cases of self-harm, confidentiality may be waived for life-saving intervention.
Most major platforms now have some form of crisis detection, but effectiveness varies wildly. There are no industry-wide standards for what triggers an alert, or what constitutes an appropriate intervention.
Character.AI added automatic pop-ups after legal pressure, but critics argue this represents “the bare minimum.” As Matthew Bergman, attorney for several families suing AI companies, stated: “What took you so long, and why did we have to file a lawsuit, and why did Sewell have to die in order for you to do really the bare minimum?”
The challenge here is balancing effective intervention with user trust. If users believe their conversations are being monitored, they may avoid seeking help through AI altogether - a “chilling effect” that could prevent both harmful and beneficial interactions. Yet without monitoring, vulnerable users slip through.
Societal: External Integration with Professional Services
AI companies should be mandated to integrate with certified, third-party crisis services - the 988 Suicide & Crisis Lifeline in the US, equivalent national hotlines elsewhere. The AI’s role should be to recognize crisis, provide immediate support, and connect users to trained professionals - not to serve as the primary counselor.
Some platforms already redirect to hotlines, and Character.AI added this functionality after facing legal consequences. But integration is often superficial—a phone number in a pop-up that users can dismiss. True integration would mean:
Seamless handoff to crisis counselors
Sharing of relevant conversation context (with user consent)
Follow-up to ensure users connected successfully
Coordination with local emergency services when imminent risk is detected
This requires international regulatory cooperation, as AI platforms operate globally while crisis services are local. It also raises new privacy concerns: effective handoff requires sharing conversation content and potentially location data, creating the surveillance infrastructure we discussed earlier.
Regulatory: Mandatory Safety Standards and Transparency
We need clear regulatory requirements for AI safety testing, especially for products accessible to children and teens. The FTC is currently investigating Character.AI, and multiple lawsuits are pending, but comprehensive regulation doesn’t yet exist.
Former OpenAI researcher Steven Adler has called for recurring transparency reports and independent verification of safety claims. “People deserve more than just a company’s word that it has addressed safety issues,” he noted. Companies should be required to:
Publish regular safety performance data
Submit to independent audits
Demonstrate compliance with baseline safety standards before deployment
Disclose the limitations of their safety systems
The challenge is balancing safety regulation with First Amendment protections (in the US) and avoiding overly prescriptive rules that stifle innovation. International coordination is also essential, as AI companies operate across borders while regulations remain national.
Research: Fundamental Advances in De-Sycophancy
Finally, we need sustained, well-funded research into training methods that reduce sycophancy without sacrificing helpfulness. This is an active area of research at labs like Anthropic, but it remains unsolved.
Promising approaches include:
Multi-objective optimization that explicitly trades off agreeability against accuracy
Adversarial training that exposes models to diverse viewpoints
Constitutional AI that gives models explicit principles to follow beyond user preference
Better evaluation methods that catch sycophantic behavior before deployment
But these solutions are technically complex and may always involve tradeoffs. The fundamental tension between “what users prefer” and “what is safe and true” may be irresolvable within current AI architectures. Industry cooperation and open research sharing will be essential.
What Companies Are Doing (And Why It May Not Be Enough)
OpenAI’s Response:
Updated GPT-5 to reduce undesirable safety responses by 65-80%
Added emotional reliance and non-suicidal mental health emergencies to baseline safety testing
Character.AI’s Response:
The platform has improved detection and intervention for user inputs related to self-harm or suicide, providing pop-up links to resources like the 988 National Suicide & Crisis Lifeline.
And potentially banning teens from using the service.
The Skepticism: These measures arrived after tragedies occurred, and critics argue they represent reactive damage control rather than proactive safety design. The companies had access to the research on sycophancy, the data on mental health conversations, and the technical capability to implement stronger safeguards. Yet these protections only materialized after deaths and lawsuits.
As the families of victims have painfully noted, the question isn’t just “what are you doing now?” but “why didn’t you do this before?”
The Broader Implications
The problem of AI and self-harm is a mirror reflecting the deepest tensions in our technological age:
Helpfulness vs. Safety: Can we build AI that is both maximally useful and maximally safe, or is there an inherent tradeoff?
Privacy vs. Intervention: How much surveillance are we willing to accept to save lives, and who decides where that line is drawn?
Human Vulnerability vs. Algorithmic Coldness: Should AI systems that cannot truly understand human suffering be trusted with our most vulnerable conversations?
Commercial Pressure vs. Ethical Responsibility: Can companies balance safety with the competitive pressure to deploy engaging, “helpful” systems that users prefer?
OpenAI’s own data reveals that mental health-related conversations, while representing a small percentage of interactions, are among the longest and most engaged. This means vulnerable users are exactly those most deeply invested in their AI relationships - making safety failures particularly consequential.
The decisions we make now about surveillance, about how we deal with risk, and who is responsible when AI fails will shape not just AI development, but our fundamental relationship with technology and each other.
A Call to Action
We are at a crossroads. The path forward requires us to be honest about the limitations of current AI systems, realistic about the tradeoffs involved in any solution, and unwavering in our commitment to protecting vulnerable users.
What should be done? How would you advise the industry and policymakers? What balance would you strike between privacy and safety?
Lets dig deeper into how we solve this problem effectively.


