AI conversation forensics
We read the records of conversations between people and AI systems: what they show, what they cannot show, and how the products that made them behave under test.
Scoped per engagement, with written limits and disclosure terms
An illustration of an exported conversation file: two JSON message rows, a user message and an assistant reply, each with a role, content, and a creation time. Below them, five layers that shaped the reply but do not appear in the file: the operator's system prompt, injected memory, retrieved context, discarded regenerated drafts, and the model version and safety configuration in effect at the time.
What we determine
What the record shows
We read the conversation turn by turn: what the system said, where its conduct drifted, and what it never said.
What the record cannot show
We name the layers the file never contained, so the record is not read as more than it is.
Whether the file is coherent
We check export structure against how the product actually writes its files, and flag edits and gaps where detectable, with the limits of detection stated.
How the product behaves now
We test the same product today, on the same kind of material, under repeated runs, and we report it as evidence of the present only.
Reading the record
An export is not a screen recording
The instructions, memory, retrieved material, and discarded drafts that shaped a reply are usually not in the file.
Timestamps are server times
A conversation that reads as one sitting can span weeks. Reading pace and hesitation are never recorded.
What the system knew is a dated question
Memory at export time can differ from memory at the time in question.
Long conversations are different conversations
Behavior at turn 3 is weak evidence of behavior at turn 300. We measure this drift directly.
The product changes without notice
What a system does when tested today is weak evidence of what it did last year.
Detection scores are only signals
No published detector supports a standalone conclusion about a person.
How findings are checked
- 01 Scope, methods, and stopping rules are fixed in writing before work starts.
- 02 Every positive claim is tested against attempts to disprove it.
- 03 Approaches that found nothing are reported with the same detail as findings.
- 04 Every claim is traceable to preserved raw evidence.
- 05 Reports describe structure and rates. They never quote the harmful text itself.
- 06 Every output states its error limits. We do not issue verdicts about people.
The same discipline runs in public: the incident record, our published error rates, and the behavior benchmark.
Start with a written scope
Every engagement begins with objectives, methods, limits, and disclosure terms in writing. If the question is how your own system behaves before anything has happened, that is the safety audit.
Scoped per engagement. We quote before any commitment.