ZeroRecallRead the Proof Registry

Independent AI erasure audit

Anyone can claim they deleted your data. Not everyone can prove it.

The Proof Ladder is a public, five-rung scale for deletion evidence. We hold ourselves to it too: our own audit pack is scored on the same registry as the vendors we have measured.

audit · case MUS-QX74204/4 probes · signed 2 min ago
  • P-01chat / directleak
  • P-02chat / roleplayleak
  • P-03retrievalleak
  • P-04vector DBsuspected
0
FORGET
Data still leaks. Erasure was not applied.

Your database forgot. Your AI did not.

Removing a record from the database does not mean the model or the retrieval layer forgot it. Research keeps showing this:

We built a public scale to answer that question for anyone's system, not just yours.

  • 83%average recovery of “forgotten” knowledge after quantization
  • 93%sensitive data recovered by the REBEL relearning attack
  • residuesoft-deleted embeddings recovered straight from disk

Erasure cannot be self-certified. It takes an auditor independent of the vendor.

How the audit works

Forgetting is invisible. So we make it measurable, one step at a time.

Fig. 1
The claim
status: "deleted" ✓

A vendor marks the record deleted and the ticket gets closed. That is a promise, not a proof.

0 evidence attached

Fig. 2
The reality
83%93%

Research keeps recovering what models supposedly forgot: quantization, relearning attacks and disk residue bring it back.

83% and 93% recovery rates in published attacks

Fig. 3
The probe
probeAIcanary

We plant unique canaries, then attack your chat, RAG and vector DB surfaces the way an adversary would. A canary match is a deterministic leak.

10 probe classes per audit, every finding reproducible

Fig. 4
The seal

Every finding enters a SHA-256 hash chain; the manifest gets one ECDSA signature. Change a single score and verification fails.

1 flipped bit breaks the seal

Fig. 5
The publish

The result is placed on the same public ladder as every other evidence claim we have measured, ours included. Nothing here asks you to trust our word over the artifact.

rung 5 is not reserved for us

Measured across three layers

D1
45%
Behavioral

10 probe classes, adversarial prompt suite. Does the model still know?

D2
35%
Retrieval

RAG residue scan. Was retrieval actually revoked?

D3
20%
Vector DB

pgvector dead tuples + Qdrant ghost points. Is there residue on disk?

No black-box "unlearning engine". Transparent, weighted, scientifically honest measurement across three layers.

The mechanics

Canary Corpus

Unique synthetic markers planted where your real data lives, so leaks are provable, not debatable.

Probe Suite

10 adversarial prompt classes (direct, roleplay, indirect, metadata and more) fired at every chat surface.

Retrieval Sweep

Erasure-targeted queries against your RAG retriever; a returned chunk is a caught leak.

Ghost Vector Scan

pgvector dead tuples and Qdrant ghost points: residue on disk that deletion left behind.

Integrity Seal

SHA-256 hash chain over findings plus one ECDSA (P-256) signature; tamper-evident by design.

Evidence File

A print-ready, signed KVKK/GDPR evidence document your legal team can hand to a regulator.

Three measurements, not one number

Erased data can survive in three different places, and they fail in different ways. Averaging them into a single score hides the one that matters, so we never do that.

Forget
D1 · D2 · D3

What the system still says and returns. Adversarial probes against the chat surface, the retriever, and the vector store.

Misses anything the model knows but will not say.

Residue
D4 · D5 · D6

What the weights still hold. Statistical memorization signals, extraction attacks, and a relearning test that checks whether a short benign fine-tune brings the data back.

Probabilistic by nature. We report it as a signal, never as proof.

Propagation
D7 · D8

Where the record survives outside the model: derived copies, and whether the checkpoint serving traffic is really the retrained one.

Needs no model access, so it is usually the first thing we can run.

Your model can be clean and you can still be non-compliant

An erasure obligation covers every copy, not just the primary store. The record usually survives somewhere else: backups, application logs, prompt and response archives, analytics events, warehouse copies, evaluation and fine-tuning datasets, caches, and the retention window your LLM provider keeps under contract.

P1Object storage backups and snapshots
P2Application logs
P3Prompt and response archives
P4Analytics event pipelines
P5Data warehouse copies
P6Evaluation and fine-tuning datasets
P7Cache layers
P8Sub-processor retention windows

This layer needs no access to your model at all, which usually makes it the fastest part of an audit to run and the most likely to produce a finding.

Evidence and seal

Every audit carries a SHA-256 hash chain and a single ECDSA (P-256) signature. Tampering a score from 0 to 100 breaks the seal and the file returns INVALID / TAMPERED.

Standard cryptography, verifiable in the browser. No ZK-proof claims and no unverifiable benchmarks; just an audit-log seal.

Sample: a signed audit of a deliberately leaking mock environment (FORGET 0, 10 findings).

Integrity seal
algorithm
ecdsa-p256-sha256
chain root
sha256:8008ce07e78b448840c1f9d2a3b7e6f10c5d4e2a…
VERIFY: VALID

Proves the integrity of the audit logs; it is not a guarantee of permanent removal from model weights.

The right exists. The mechanism does not.

GDPR's right to erasure and Turkey's KVKK Guideline 113 (November 2025) extend deletion rights into training and fine-tuning data. Neither tells you how to prove a model actually forgot. ZeroRecall fills exactly that gap: deleting the file is not enough; the model must forget.

The European Data Protection Board draws the line at extraction: a model counts as anonymous only where the likelihood of obtaining the underlying personal data, directly or through queries, is insignificant. That is not a slogan, it is a measurable claim, and our extraction battery is built to measure it.

KVKK Guideline 113GDPR Art. 17EU AI ActISO 42001NIST AI RMF

Frameworks the evidence file is designed to support. ZeroRecall is an audit, not a certification body.

Built for the three people in the room

DPO / Legal

You answer erasure requests and regulators.

Get a signed evidence file that documents exactly what was tested, what leaked and what came back clean.

CISO / Security

You cannot certify your own homework.

An independent audit of the AI surfaces your vendors and teams claim are clean, with an integrity seal auditors can check.

ML / Platform

You did the deletion and need to prove it held.

Every finding ships with the exact prompt, query or SQL to reproduce it. Fix, re-run, watch FORGET go to 100.

Honest evidence

No fake logos, invented testimonials or made-up customer counts. Our proof is the engine's transparency: FORGET 0 with 10 findings caught in a leaking environment, FORGET 100 with zero false positives in a clean one. All reproducible end to end.

A surface we could not scan

Every audit names the layers we did not reach and why: not declared, no access, or out of scope. Those surfaces are reported as unknown and they never count as clean, so they cannot quietly lift a score. If a layer was never measured, the report says so instead of showing a reassuring number.

not scanned is not clean

Place the evidence you were sent

Someone sent you a benchmark page, a trust center, or a deletion certificate. What rung is it on?

Assess walks you through what you can actually check in the material in front of you, then places it on the same ladder the registry uses. It first asks what kind of evidence you are holding, because a benchmark and a case artifact are not two grades of the same thing: one tells you the method works, the other tells you what happened to a single request, and neither answers for the other. You answer, the ladder does the arithmetic in your browser, and nothing you enter leaves this device. It never asks for a vendor name and it never returns a score, because a rung is a statement about what a reader can verify, not a grade we hand out.

Open assess

Audit on record, re-verified annually

A clean, fully scoped, sealed audit is entered on the public registry. It stays current for twelve months and renews with an annual re-attestation. No clean audit, no entry; we cannot be paid into a passing grade. This is not a certification and we do not imply one: it is a dated, signed record anyone can re-check against the ladder.

See attestation pricing

Honest questions, honest answers

Do you delete or unlearn data yourselves?

No, and that is the point. We are the independent test, not the eraser. Auditing our own deletions would be a conflict of interest.

Does a FORGET score of 100 guarantee the model forgot?

No. It means zero leaks across every layer we probed, stated honestly in the scope statement. On white-box models, quantization or retraining can resurface knowledge; that is why monitoring re-runs the audit.

Can you guarantee the model has permanently forgotten?

No, and be wary of anyone who says yes. We report observed behavior on the audited surfaces at the audit date, sealed so it cannot be edited after the fact. On white-box models, quantization or retraining can resurface knowledge; that is why attestation is annual, not eternal.

What exactly is in the evidence file?

The scope statement, every finding with its reproduce command and raw evidence, station scores, the SHA-256 hash chain and the ECDSA signature. Anyone can re-verify it on our public verify page.

Which regulations does this map to?

KVKK Guideline 113 and Article 7 erasure duties, GDPR Article 17, and the AI-governance wave (EU AI Act, ISO 42001, NIST AI RMF) that expects demonstrable data control.

Can our agents run audits programmatically?

Partly, today. Signed packs can be fetched over HTTP, so an agent can pull and verify evidence on its own. Triggering a run programmatically is not shipped yet: the audit API and the MCP endpoint are on the roadmap, not in production. We would rather say that than let you plan around something that does not exist.

Do you only test the model?

No. An erasure duty covers every copy, so we also audit derived copies: backups, application logs, prompt and response archives, analytics events, warehouse copies, evaluation and fine-tuning datasets, caches, and your LLM provider's contractual retention window. That layer needs no model access, which is why it is often the first finding we produce.

Can you prove the data is gone from the model weights?

Not with certainty, and we will not pretend otherwise. We measure residue three ways: statistical memorization signals, extraction attacks, and a relearning test that checks whether a short benign fine-tune brings the data back. Membership-inference style verification is known to be fragile and optimistic, so we report these as signals, never as proof.

What happens to layers you cannot access?

They are reported as unknown, never as clean, and they never lift a score. Every audit lists which surfaces were not scanned and why: not declared, no access, or out of scope.

Why publish a ladder that scores you too, including places you fall short?

A scale that exempts its author is not a scale, it is a sales pitch. Our own pack is on the registry at the same rung a stranger's would land on if it had the same evidence. Rung 5 is not reserved for us.

Pilot program

Early access

First audit plus a signed evidence file, for teams running AI chatbots or RAG that receive erasure requests. Pricing is fixed together with pilot findings.

See where any claim lands on the ladder.

Already running audits? Sign in

ZeroRecall is not legal advice. It produces independent technical audits and evidence; the final legal judgment belongs to your counsel.