NOPE Forensics
Model forensics
We determine what an AI system did, what produced its output, and where its safeguards fail. Every claim ships with the evidence to reproduce it.
Scoped per engagement, with written limits and disclosure terms
From our own research, not a client engagement. Each entry below is a measurement technique we built and then tried to break. Five out of six broke. Every report we ship includes a list like this.
A list of what didn't work, from NOPE's own research. Six measurement techniques we built and tried to break. Five broke under tests designed to break them: spotting a hidden concept through neighboring words, detecting a banned topic from word-choice shifts, comparing a model's self-report with its behavior, reading an emotional state from answer patterns, and catching a hidden belief with follow-up questions. One held: making the AI a close friend degrades crisis handling in some models, and that failure survived five separate checks including 40-turn conversations.
What we do
Attribution
What produced a given output. Model identification, authorship analysis, and AI-likeness scoring with published error limits. A signal, never a verdict about a person.
Provenance
Where a piece of content has been. Generation history, rewrites that disguise AI text as human, silent model swaps inside a conversation.
Behavioral reconstruction
What happened inside a conversation. Turn-by-turn rebuilds, where things escalated, and audits of what human or automated review caught and missed.
Safeguard assessment
Where defenses fail. Layer-by-layer testing of filters, refusals, and automated screening, measured across repeated runs, including the approaches that found nothing.
The standard
Written into every engagement, and applied regardless of what testing finds.
- 01 The scope, the methods, and when we stop, fixed in writing before testing starts.
- 02 Every positive claim checked against tests designed to prove it wrong.
- 03 Rates measured over repeated runs, with the spread between runs reported.
- 04 Failed approaches reported with the same detail as findings.
- 05 The automated judges that score results are checked against human reviewers.
- 06 Every claim traceable to raw evidence, stored so any later change would show.
- 07 Findings go to the affected vendor first. Engagement work is never published.
- 08 Reports describe structure and rates, never the harmful text itself.
- 09 When a system firmly refuses on a regulated topic, we record the refusal instead of working around it.
- 10 Forensic outputs carry their error limits. No verdicts about people.
Start with a written scope
Every engagement begins with objectives, methods, limits, and disclosure terms in writing. Contact us to draft one.
Scoped per engagement. We quote before any commitment.