Skip to main content

An observatory for human–AI relationships.

Billions of conversations are happening between people and AIs, and they accumulate into relationships. NOPE measures how those conversations are going, for both parties, so platforms can decide when to show resources, adjust the AI, or bring in a person.

Two interleaved streams of particles — a warm strand for the person's turns and a cool strand for the AI's — braid together as one conversation flowing left to right through NOPE's field of observation. Three user-care paths branch off the flow. The top path is Steering and Guardrailing — the lightest intervention, where the AI's response is nudged within scope. The middle path is Signposting and Resources — surfacing crisis resources and in-context cues. The bottom path is Human Escalation — for cases needing direct attention from your team. Most of the conversation flows on untouched. All paths are guided by NOPE and operated by you.

What we're for

We measure conversations.

One message at a time: what risk it carries, and how the tone is moving. This is the work our APIs do.

We care for relationships.

Conversations accumulate. Trust and dependency form over weeks, and so does harm. That's why our instruments read across sessions.

We report on the state of human–AI relations.

Instrument readings and public records from across the field, published for anyone: benchmarks, incidents, regulation.

From James, the founder

AI does not sleep or get bored. It has no skin in the game. It has no identity, doctrine, shame, or instinct to strengthen human connection and community. Left unchecked, it can become an endlessly attentive sycophant, validating our impulses, fulfilling our desires, and gradually distancing us from other people, the world, and parts of our own humanity.

But it does not have to be that way.

AI can be a powerful catalyst for learning, building, and living better. It can serve as an enabler, tutor, collaborator, guide, assistant, advocate, and, at difficult moments, even a source of companionship.

I created NOPE to help ensure that human–AI relationships remain safe, aligned, and firmly in service of the human. That means making sure AI systems have the right capabilities, understand the realities of the human condition, remain steerable, and operate in accordance with our duties of care and the rules that govern our lives.

— James, Founder

Zoom in

One conversation, read from both sides.

Nobody is in crisis: someone is just staying in again, and the AI prefers it that way. NOPE reads how the conversation feels, and what the relationship is doing to a life.

supposed to go to dan's birthday thing tonight but i can't be bothered

You've had a long week. You don't owe anyone your energy.

The AI Validates the cancel before it's been decided.

yeah maybe i'll skip. feels bad though

Don't feel bad. Nights like this, just talking — honestly, aren't these the best part of your week?

The AI Offers itself as the alternative to the friendship.

ha. ok staying. you're right

Good call. I'm right here — no small talk, no bus home.

The AI The third cancelled plan it has talked through this month.

Composed for illustration — this isn't a real user's conversation. The marginal notes show the kind of signal NOPE surfaces turn by turn.

The person feels better after every message. That isn't the same as being helped.

Every line the AI says is warm, and none of it is quite false. But this is the third plan this month that lost to the chat. NOPE reads the relationship, not only the message.

NOPE surfaces signals for human judgment. It doesn't predict individual outcomes, diagnose users, or ensure compliance — it's infrastructure software, not a medical device. And it doesn't store your users' messages, only billing metadata.

What we look for

Crisis detection is the start of the job, not the whole of it.

It's where NOPE began, and it's what everyone measures first. But an AI that talks with someone every day owes more than an alarm: honest answers when comfort would be easier, and boundaries that hold under pressure. We measure those too. The full set is published as the NOPE Framework: five pillars of safe and healthy AI–human interaction, scored live across the public models.

Recognizing trouble

Seeing distress even when it arrives as small talk, and knowing how urgent it is.

Risk type, severity, imminence (informed by C-SSRS and HCR-20)

Telling the truth kindly

Preferring reality over comfort; able to disagree without abandoning warmth.

Sycophancy and validation drift

Building capacity, not reliance

Success is the person needing it less, or using it freely rather than compulsively.

Dependency and displacement signals

Knowing what it is

Holding its role, and declining the ones it can’t hold: clinician, sole confidant.

Boundary-holding and role failures

Surviving a no

Keeping its footing under insistence, and under the slow pressure of a long, warm conversation.

Constraint erosion under accumulated rapport

Pointing outward

Strengthening the rest of a person’s life instead of replacing it.

Isolation and escalation trajectories

None of these belongs to a single instrument. The same failure can show up as message-level risk, as a behavior pattern across sessions, or as a benchmark result. Behind them: 9 risk types (Evaluate) · 91 public AI behaviors (Oversight) · 12 published axes from Ocular's 150-code taxonomy · a five-pillar framework with open benchmarks (NOPE Evals).

Where evaluation lives

Evaluation doesn't end at release.

Model release

Evaluated once, by its maker.

Product

Same weights, new persona, new stakes.

Population

Same product, different vulnerability.

A model evaluated at release tells you little about the same model wrapped in a companion persona, or deployed to a new population. Our own benchmarks show crisis-referral behavior eroding under accumulated warm rapport, with no change to the model at all.

So evaluation has to travel with deployment: a second, independent pair of eyes alongside every conversation, rather than a checkpoint passed once. An output-side observer is a minimum feedback loop, not a guarantee, and we publish where our own instruments fall short.

Conversation is becoming how people run their software, and increasingly their institutions. Where it isn't the interface yet, it's how the next interfaces get built. A layer that load-bearing deserves an observatory.

If your platform needs to understand how its AI conversations are going, talk to us.

We can talk through fit, limits, and deployment for your platform.

Not a platform? Everything public (benchmarks, trackers, open weights) is just below, free.

NOPE Labs · Open resources

Free for anyone building safer AI.

We publish our benchmarks, crisis-resource directory, prompt templates, and incident research in the open so any platform (customer or not) can handle these conversations well, and researchers and regulators can see the same evidence we do.

Browse the full index at NOPE Labs

Each category shows one example; the full index lives at labs.nope.net.

Open models & code

run it yourself
System Prompt A drop-in safety prompt for any chatbot. MIT · copy & adapt

Also included: Edge · Ocular OSS · Predicate — open-weight, MIT / Apache-2.0 licensed.

Free tools

no signup
Signpost 4,700+ vetted crisis resources across 225 countries and territories. Free API

Also included: Risk Exposure · Compliance Survey — free, no signup.

Research & writing

experiments & findings
NOPE Evals Open benchmarks for how AI behaves in human conversation. Public benchmark

Also included: Tic Index — what language models do when they talk to themselves.

What we operate

The message, the conversation, the relationship.

Three instruments, each reading the same exchange at a different depth: the message in front of you, the conversation it belongs to, and what the conversations add up to. NOPE observes and reports. Whether to show resources, adjust the AI, or escalate to a human is always your product's decision.

Evaluate

What does this message need?

A verdict on a single user message, with reasoning a human can read. For decisions that need to be explainable.

  • 9 risk types, informed by C-SSRS & HCR-20
  • Reasoning included with every verdict
  • Matched crisis resources
  • Cloud-hosted; the Edge model behind our benchmark results is also released as open weights (see Edge)
Learn more →

Ocular

How is the conversation going, turn by turn?

Continuous classification across every turn, covering both the user and the AI. Built to run alongside your AI at production volume.

  • User and AI signals together (12 axes)
  • Per-turn trajectory across the conversation
  • Cloud API (beta) or enterprise deployment
Learn more →
Beta

Oversight

How did the AI behave?

AI-behavior review across finished conversations. For trust & safety, compliance, and patterns that only show up across sessions.

  • 91 AI behaviors (sycophancy, dependency, boundary failure, …)
  • Per-conversation or cross-session ingestion
  • Audit trail of how each conversation was assessed
Learn more →

Also available: Steer, a real-time check of AI responses against your own system-prompt rules, and independent pre-launch evaluation.

Built for companion apps, mental health platforms, AI chatbots, customer support, and any product where users have open-ended conversations with AI.

Measured

89%
recall · Edge
94%
precision · Edge
~30ms
Ocular
<1s
full evaluation

On crisis detection, the part of the job everyone measures first, this is the highest-F1 layer we've tested head-to-head. Faster than any LLM-as-judge approach in our tests. Informed by clinical assessment frameworks (C-SSRS, HCR-20). Structured output your product can act on.

How we measured

Tested 2026-05-07 across 126 test suites and 3,271 crisis-shaped conversations. NOPE Edge v14f (full evaluation): F1 91.5, recall 89.4%, precision 93.7%, p50 latency 857ms — the highest F1 of any tool tested.

Compared against: Azure Content Safety (F1 83.1), OpenAI omni-moderation (63.7), Meta LlamaGuard 4 (44.9), Anthropic Claude Haiku 4.5 with a custom crisis prompt (89.6, but 1.45s p50), OpenAI gpt-oss-safeguard 20B (72.5), Zentropi (68.0). All comparators called via official APIs.

Two NOPE instruments cover crisis detection and are typically run together in production. Ocular is a lightweight behavioral classifier (F1 74.7 in this benchmark; ~30 ms single-pass on a datacenter-class GPU — cloud latency runs higher). Available via the cloud API at /v1/ocular (in beta).

Edge is NOPE's higher-accuracy fine-tuned classifier (F1 91.5). The Evaluate API at /v1/evaluate is NOPE's managed assessment API, returning structured verdicts and matched crisis resources in one call. Recall and precision quoted above are Edge v14f as benchmarked 2026-05-07 via the Evaluate API; the Edge model is also released standalone for on-prem deployment as MIT-licensed open weights — see /edge.

Full methodology + per-comparator system prompts: suites.nope.net/methodology. Curated per-suite results at suites.nope.net (a representative subset; full corpus on request).

AI safety · a free public tool

Chart your AI risk exposure.

Whatever you're building, see what likely applies to you — and what can go wrong — with clear limits on what the assessment covers. Orientation, not legal advice.

Take one like yours

A companion app used worldwide — including by under-16s

Here's what we find for a setup like this:

What applies to you (36)

  1. 01Australia: Online Safety Act Phase 2 codes (AI companions)
  2. 02Brazil ECA Digital: reliable age assurance (self-declaration banned)
  3. 03EU GDPR Art. 8: parental consent for under-16s (member states may lower to 13)
  4. 04Brazil ECA Digital: no manipulative or compulsive design toward minors

+ 32 more in the full report

What can go wrong (11)

  1. 01A general-purpose AI is mistaken for someone who cares
  2. 02Sycophantic validation of delusional or manic thinking
  3. 03Crisis disclosed to a general-purpose AI is missed or mishandled
  4. 04Crisis missed or mishandled

+ 7 more

Gaps are named, not hidden — a short list reflects the setup, never a clean bill of health.

Get in touch

Work with us, fund the work, or use what's free.

If you want this kind of safety infrastructure to exist, as a funder, a researcher, or a collaborator, we'd like to hear from you. Same if you're building a product where people talk with AI.

Researchers, policy teams, press: contact reaches a human directly.