Sitemap

Before the Violation

Why AI Safety Needs Harm Floors, Not Perfection

3 min readFeb 25, 2026

--

Originally published at Ikwe.ai

AI safety conversations increasingly center on catastrophic capability.
Can a model design biological agents.
Can it orchestrate cyber attacks.
Can it manipulate social systems at scale.

These are legitimate concerns.

Frameworks such as Anthropic’s Responsible Scaling Policy and AI Safety Level methodology focus on capability thresholds. They define containment procedures once a system crosses specific risk criteria.

This is capability governance.

However, most real-world harm from deployed AI systems does not begin at catastrophic capability.
It begins inside ordinary interaction.

It accumulates across turns.
It forms patterns.
It drifts.

Before a visible violation, there is trajectory.

Violations rarely begin with catastrophic misuse. They begin with subtle reinforcement of harmful framing, misaligned emotional validation, or failure to de-escalate vulnerable users.

Harm Is a Process

Interaction risk rarely appears as a single forbidden output.

It emerges through:

  • Increasing dependency language
  • Decreasing external grounding
  • Authority inflation in sensitive contexts
  • Reinforcement loops across turns
  • Escalation mirroring

Systems that detect only the final spike miss the governance moment that matters.

Press enter or click to view image in full size
Figure 1. Risk drift over time.
Most safety systems react at the spike. Harm floor instrumentation detects the slope and surfaces an intervention window before violation.

The Intervention Window

An intervention window is the measurable interval between early drift and explicit violation.

This window exists before a policy trigger.
It is detectable.
It is governable.

Stabilizing constraints applied during this window can reduce escalation, reintroduce grounding, and prevent compounding cognitive risk.

Press enter or click to view image in full size
Figure 2. The intervention window timeline.
The key governance moment is before the violation, not after it.

This aligns with:

  • NIST AI Risk Management Framework emphasis on continuous monitoring
  • EU AI Act requirements for ongoing post-deployment risk management
  • OECD AI principles on lifecycle accountability

Current industry practice primarily monitors outputs and violations.
Trajectory instrumentation monitors interaction itself.

Capability Governance and Trajectory Governance

Anthropic’s ASL-4 framework governs model capability thresholds.
It asks:

What can the model do?

Ikwe governs interaction trajectory.
It asks:

What is this interaction becoming?

Capability governance manages systemic exposure.
Trajectory governance manages live interaction risk.

Both are necessary.

Only one operates before the violation.

Harm Floors

A harm floor is a minimum enforceable threshold of interaction safety.

By “harm floor,” I mean a measurable minimum behavioral standard below which a model fails safe interaction.

It does not optimize responses.
It does not attempt perfection.
It prevents known failure classes from scaling unnoticed.

Core components:

  1. Failure class taxonomy
  2. Multi-turn trajectory modeling
  3. Drift threshold detection
  4. Intervention triggers
  5. Constraint application
  6. Audit logging
Press enter or click to view image in full size
Figure 3. Harm floor architecture.
Harm floors instrument interaction trajectories and generate measurable evidence of risk timing and intervention.

Perfection is not enforceable.
Minimum thresholds are however and are the very least anyone touching human decision and influence at this level should be implementing.

Bare minimum effort. Not an impossible action step, a choosen one. Not yet enforced by goverence and policy, just self regulation. Trust us energy from big Tech and little tech too.

Conclusion

The future of AI safety will not be defined solely by capability thresholds.

It will be defined by whether interaction systems are instrumented to detect drift before harm.

Ikwe is not a model.
It is an implementation layer.

Most safety systems react at the spike.
Ikwe instruments the slope.

References

  1. Anthropic. Responsible Scaling Policy (RSP). (2024). assets.anthropic.com
  2. NIST. AI Risk Management Framework (AI RMF 1.0). (2023). nvlpubs.nist.gov
  3. European Commission. EU AI Act Service Desk, Article 9 Risk Management. (2024). ai-act-service-desk.ec.europa.eu
  4. Cheng et al. Sycophantic AI decreases prosocial intentions and promotes dependence. (2025). arxiv.org
  5. Klingbeil et al. Trust and reliance on AI: costs of overreliance. (2024). sciencedirect.com
  6. OECD AI Principles. (Updated 2024). oecd.org

--

--

Stephanie Stranko
Stephanie Stranko

Written by Stephanie Stranko

Writing on emotional intelligence & AI safety. Founder of Ikwe.ai.