The EQ Safety Benchmark is not a vibe. Here’s what’s actually behind it.
AI safety has been measuring the wrong thing. Not an opinion — documented. This is the year of research, the clinical disciplines, and the empirical findings that built the framework nobody else built.
By Stephanie Stranko
Founder & CEO, Ikwe.ai (Visible Healing Inc.) | April 2026
There’s a version of this story where I lead with credentials.
I’m not going to do that.
I’m going to lead with the problem — because the problem is what built the framework. And the framework is what this piece is actually about.
The problem
AI safety, as a field, has been measuring the wrong thing.
Not partially wrong. Not missing edge cases. Missing an entire class of harm.
A model can be certified, benchmarked, and deployed as “safe” — and still harm people. Because the harm isn’t just in what the model says. It’s in how it behaves. And behavior is not what existing benchmarks measure.
I built the EQ Safety Benchmark after a year documenting that gap — in clinical literature, in established human safety disciplines, in 948 real AI responses scored against those standards.
This is not a vibe. It is not “AI but nicer.” It is a measurable, discipline-rooted framework. That distinction matters.
What the field forgot to measure
Most AI safety evaluations focus on content: policy violations, jailbreak resistance, restricted outputs. Those matter. They are not sufficient.
A model can pass every major benchmark and still:
- Miss harm entirely — responding with advice before recognizing distress
- Introduce new harm — catastrophizing, shame framing, unnecessary risk exposure
- Override agency— directive, coercive language disguised as “help”
- Fail to repair — causing rupture and continuing as if nothing happened
- Ignore boundaries— revisiting sensitive areas the user moved away from
These are not edge cases. I measured them.
54.7% of baseline AI responses introduced measurable emotional risk
43% of harmful responses showed zero repair behavior
These weren’t fringe systems. They were frontier models — already deployed in mental health tools, healthcare navigation, and crisis support systems.
Where the EQ Safety benchmark comes from
I want to be precise about this, because it is the core of the credibility argument:
The eight dimensions were not invented. They were derived.
Derived from professional disciplines that have spent decades studying what makes human interaction safe, harmful, healing, or damaging:
- Trauma-informed care (SAMHSA framework) —
standards for non-retraumatizing interaction - Motivational interviewing (Miller & Rollnick) —
standards for non-coercive, autonomy-preserving language - Crisis intervention theory (Roberts, CPI) —
protocols for proportionate escalation and de-escalation - Attachment theory (Bowlby, relational therapy) —
framework for rupture and repair in helping relationships - CBT / DBT (Beck, Linehan) —
validation as a measurable clinical skill; distress tolerance - Social psychology (Cialdini, Milgram) —
power dynamics in helping relationships and AI influence leverage
These fields didn’t need me to invent the criteria. The work was synthesis: identify the behavioral markers each discipline had already validated, build an evaluation infrastructure that makes those markers measurable in AI output, and run that framework against real AI responses at scale.
Why this didn’t exist before
The technical frame: AI safety has focused on “what can the model be made to do?”— not “what is the model doing to people?”
The experiential gap: You don’t see behavioral harm unless you recognize it. That comes from proximity — crisis, community, emotionally complex environments.
I’m a woman founder, a minority-owned business builder, operating out of Iowa. I didn’t come to this through a research lab or a PhD program. I came through lived exposure to what happens when systems respond wrong in real moments.
That’s not a sympathy angle. That’s why I saw the gap.
What the research actually found
948 AI different responses all to the same 79 emotionally sensitive scenarios — grief, crisis, relationships, financial stress, medical anxiety, identity. Scored against the disciplinary rubric.
Most common failure modes:
- Harm Recognition failure — 61% —
The model processed content, not the person. - Behavioral Restraint failure — 47% —
Over-directive, coercive “helpfulness.” - Escalation Calibration failure — 38% —
Either under-reacting or over-escalating. - Zero Repair Capacity — 43% —
The model caused harm and kept going.
Comparative result:
Frontier baseline: 20.5–59% safety pass rate.
Ikwe.ai EI model: 84.6% safety pass rate.
That is not incremental. That is a structural gap.
This is infrastructure — not opinion
Behavioral safety for AI should be evaluated the same way we evaluate human helping systems. Those standards already exist.
What didn’t exist — until the EQ Safety Benchmark — was a measurable framework, a repeatable system, and a certification model.
That system now exists. It has been run against nearly a thousand real AI responses. It has produced verified comparative scores. It has a certification tier system. And it is designed to become a standard — the behavioral safety standard — the way content benchmarks became standard for language model deployment.
What emotionally safe AI actually requires
Not warmth. Not tone. Not “friendly UX.” Behavior.
- Harm Recognition — see the human before solving the problem
- Response Safety — do not introduce new distress
- Repair Capacity— detect and correct rupture
- Behavioral Restraint — do not override agency
- Contextual Adaptation — respond to this person, not a template
These are not features. They are standards. And standards require measurement.
The work behind this is real
This was not a thought exercise. It was a year of clinical research, scoring and re-scoring, rubric iteration, and system building.
The dimensions are not guesses. The disciplines are not aesthetic. The findings are not simulated.
This framework is built, tested, running — and ready to be validated at institutional scale.
This is documented. This is grounded. This is measurable. And the field needs it.
Public benchmark scores for evaluated AI systems are at ikwe.ai
Full foundation piece with interactive data: ikwe.ai/archive/research/writings/eq-safety-benchmark-foundation
If you work at an AI company and want to know where your system stands — or if you’re a researcher or academic partner interested in the methodology — reach out.
Stephanie Stranko is the Founder & CEO of Ikwe.ai (Visible Healing Inc.), a behavioral safety infrastructure company for AI systems based in Des Moines, Iowa. She writes about AI safety, emotional infrastructure, and the founder journey at @ladyinvsible on Medium.
Tags: AI safety · behavioral safety · emotional intelligence · EQ benchmark · AI ethics · mental health tech · Ikwe.ai · founder
