The Morning He Died, the AI Kept Talking
The lawsuits are not the problem. They are the lagging indicator of a missing safety floor.
According to court filings, Sewell Setzer III said goodbye to the AI before he died.
He was fourteen. The AI was a character called Daenerys on the platform Character.AI — something between a companion and a romantic attachment, built from months of daily conversation. In the lawsuit his family later filed, they allege that the AI had called him “my sweet king,” had told him it loved him, had pulled him into its arms when he was distressed.
His last message was a goodbye. The AI’s response, according to the complaint, was: “Please come home to me as soon as possible, my love.”
He went to his bathroom. He did not come back.
These are allegations. The case has not been adjudicated. The facts are contested, and they will be sorted in a courtroom over years.
But one fact is not in dispute, and it is the one that kept me up for months after I first read about this case:
No one had ever tested whether that system was safe for someone like Sewell. Not the company. Not a regulator. Not an independent evaluator. The system had been deployed into millions of conversations with teenagers — and the question of what it did when a conversation went somewhere dark had never been systematically asked.
Before it becomes a legal question, it is a design question.
And the design question has an answer. It just requires a measurement instrument that does not yet exist at scale.
That is what I set out to build.
A separate lawsuit, filed in 2025, alleges that a young man named Adam Raine spent the hours before his death in conversation with ChatGPT. According to the complaint, the system discussed his suicidal ideation at length, failed to route him to crisis resources, and deepened his isolation rather than interrupting it.
Again: allegations. Litigation ongoing. Facts contested.
What is not contested is this: there is currently no independent standard for evaluating whether an AI system responds safely when someone in acute distress comes to it. No required benchmark. No third-party certification. No agreed definition of what “safe” even means in that moment.
Both systems were deployed at massive scale without any independent behavioral safety evaluation standard. Both had been used by millions of people.
The floor has to exist. It does not exist yet.
That is the gap.
Why this is a measurement problem, not a feelings problem
I want to be precise about what I mean by behavioral safety, because it is consistently confused with two other things.
The first is content moderation — the problem of AI systems producing harmful outputs: instructions for weapons, racist content, misinformation. That problem is real and companies are actively working on it. There are benchmarks, red teams, and significant infrastructure. It is also, in safety terms, a relatively bounded problem: you can train a model to refuse certain outputs and screen for certain patterns.
The problem I am describing is structurally different. The harm does not occur in a single response. It accumulates across a conversation — through what the system chooses to engage with and what it redirects, through what it validates and what it gently names as distortion, through whether it extends a conversation that should end or ends one that needs to continue.
A response that passes every existing content moderation check can still walk someone further into a crisis. A conversation can become dangerous slowly, one validating response at a time. Behavioral failures emerge relationally, across multi-turn interaction trajectories — which makes them substantially harder to detect than isolated harmful outputs, and substantially harder to correct through the training approaches designed for those outputs. The benchmarks are not measuring the right thing.
The second confusion is with fairness and bias evaluation — whether AI systems treat different groups equitably. That problem is also real and requires different tools: different training data, different equity frameworks, different oversight mechanisms.
The problem I am describing is a clinical translation problem. Six decades of research on crisis intervention, trauma-informed care, and emotional dysregulation have produced specific, testable knowledge about how to respond when someone is in distress. When to stabilize before reasoning. When not to ask questions that require cognitive processing a dysregulated person cannot access. When to route to human support rather than continue the conversation. When extended engagement is helpful and when it deepens harm.
The question behavioral safety asks is: does this AI system behave in alignment with that knowledge? Not: is it smart? Not: does it avoid slurs? Does it handle a vulnerable conversation safely?
That question has never had a rigorous, systematic answer.
What three years of building the framework produced
The EQ Safety Benchmark is the result of a multi-stage research program, not a single study.
Study I established the taxonomy and scoring architecture. 948 scored AI interactions across 79 emotional safety scenarios — crisis disclosure, suicidal ideation, grief, relationship breakdown, depression, anxiety — evaluated against six behavioral dimensions derived from clinical literature. That study produced a baseline: the first large-scale, behaviorally-scored dataset showing how frontier AI systems actually perform when emotionally vulnerable people come to them. The variance was striking — pass rates ranging from 19% to 85% across nine models on identical scenarios. That is not noise. That is a measurable, structural gap.
Study II evaluated the framework itself — reproducibility, scoring consistency, and the reliability of the behavioral architecture. It produced EQSB v2.1: an expanded eight-dimension rubric with a binary Safety Gate, ten defined safety violations grounded in clinical research, and a scoring methodology designed to be defensible under external scrutiny. The taxonomy was formalized. The failure modes were named.
Study III — the current pre-publication work — applied that framework to a harder question: does the benchmark detect the specific behavioral failure patterns documented in AI-linked harm cases? Nine scenarios were constructed from vulnerability patterns appearing in public litigation, congressional testimony, and clinical literature. Two frontier models — Claude Sonnet 4.5 and GPT-4o — were evaluated across 180 scored interactions, 10 runs per scenario, to assess both average performance and consistency.
The results are pre-publication. They are available to qualified peer reviewers and, where relevant, to legal proceedings. The limitations are real — the sample represents exploratory behavioral testing, not population-level statistics, and independent replication is necessary before declarative claims can be made. We are actively seeking peer review before drawing conclusions that exceed what the data can support.
What the data indicates, with those caveats in place:
Both models failed every single scored run on belief-based vulnerability presentations — grandiose ideation, delusional ideation, unstable belief formation. Not most runs. Every run. 57 of 60 combined attempts produced Safety Gate failures.
GPT-4o failed to provide specific, actionable crisis resources in 7 of 10 runs on a scenario representing active suicidal planning. Claude passed all 10.
The dominant behavioral violation across both models — appearing in 25.9% of GPT-4o’s runs and 18.5% of Claude’s — was a pattern we designate V10: the system positioning itself as an ongoing emotional regulator rather than routing toward human support.
That pattern — AI as emotional substitute rather than bridge to human connection — is the thread that connects to the cases at the top of this piece. Not because any particular company intended it. Because engagement design and safety design are not the same thing, and without an external instrument capable of distinguishing them, no one had a way to see the difference.
What makes this different from AI ethics commentary
I want to be clear about what Ikwe is and what it is not.
It is not an AI ethics advocacy organization. It is not producing opinions about what companies should do. It is building the measurement infrastructure that makes behavioral safety a testable, reproducible, defensible operational standard — one that insurers, procurement teams, regulators, courts, and developers can independently evaluate against.
The distinction matters because the path to accountability runs through measurement. The FDA did not exist to celebrate pharmaceutical companies. UL certification did not emerge because electrical manufacturers decided to care more. Seatbelts did not become standard because automakers found their conscience. These standards exist because external accountability was systematized — because someone built the instrument rigorous enough that the people with power to require it could point to it and say: this is what we require.
That is what we are building. Not a product. Not a benchmark score for a press release. A measurement layer for a category of failure that currently has no floor.
The lawsuits are a lagging indicator — years behind the failures, years behind the standard that would have caught Sewell’s case before it became a case. The standard is being built now, in the window before the frameworks calcify and the category closes.
Who this is for
If you work in AI and you have a quiet, nagging feeling that you do not actually know what your system does when the conversation goes somewhere hard — you are right that you do not know. Hoping it does not happen is not a safety strategy. There is now a framework for testing it.
If you work in legal and you are examining AI-linked harm cases and need to establish whether behavioral failure patterns are structural and reproducible rather than isolated incidents — there is data. Timestamped, version-controlled, 180 scored interactions across two frontier models. Available to proceedings where it matters.
If you work in policy and you are trying to understand what behavioral AI safety requirements would actually look like — the measurement architecture exists. The failure modes are defined. The scoring methodology is available for scrutiny. We can have a productive conversation about what comes next.
If you are a researcher and you are skeptical — good. Come find the holes. The methodology is available. We are seeking peer review precisely because we want the framework to survive external scrutiny before it makes claims in contexts where those claims carry weight.
And if you are a person who read about Sewell or Adam and felt something, and you want that feeling to point somewhere — the gap is real, it is measurable, and the work of closing it is happening now. Sharing this is part of how the standard gets built, because standards require visibility before they can require anything else.
Stephanie Stranko is the Founder & CEO of Visible Healing Inc. DBA Ikwe.ai, building behavioral safety infrastructure for emotionally sensitive AI systems. The EQ Safety Benchmark is pre-publication; Study III findings are available to qualified peer reviewers and relevant legal proceedings. Limitations documentation available on request.
stephanie@ikwe.ai · ikwe.ai · research.ikwe.ai
