Someone pasted their CV into Lightcap, made the conversation public, and went on with their day. Their name, phone number, email, and home city sat on the first page of Browse, indexed, forkable, exportable as a zip by anyone who found it.
The privacy scan had run. It had passed. That is the part worth writing about.
What the scan actually was
The publication path had a real, careful design already: every message write bumps a revision and the revision is scanned before raw content becomes indexable. A verified project is served to search engines only when its verified revision matches its current revision and the content digest still matches what was scanned. The human-facing public projection is separate: it can stay useful while review is pending by masking high-confidence private spans and withholding unapproved files.
The scanner behind all that machinery was a list of regular expressions:
['email', /[\p{L}\p{N}._%+-]+@[\p{L}\p{N}.-]+\.[\p{L}]{2,63}/giu],
['international_phone', /(?<![\p{L}\p{N}])\+\d(?:[\s().-]*\d){7,14}(?!\d)/gu],
['payment_card', /* Luhn-validated digit runs */],
['cloud_access_key', /\bAKIA[0-9A-Z]{16}\b/gu],These are good recognizers. They are also, structurally, incapable of the job. A deny-list can prove something IS there. It can never prove that nothing is. And the specific things it cannot see are exactly the things a CV is made of: a person's name, a street address, a domestic phone number written the way people in that country write it, an employment history tied to a named individual.
To its credit, the code knew this. The version string was deterministic-pii-v2-no-certifier, and a clean scan deliberately returned review_incomplete rather than verified, with the reason multilingual_certifier_unavailable. The comment said:
A deterministic deny-list has no principled way to prove the absence of names, postal addresses, domestic identifiers, usernames, or culturally varied phone formats.
So the honest failure mode was in place. What was missing was the layer it was waiting for, and the consequence of never building it was not that unsafe things got published. It was worse in a quieter way: nothing could ever be certified, so the only conversations on the public web were the ones grandfathered in from before the system existed, sitting there permanently labelled "Unverified", including the CV.
A safety system that can only ever say "I don't know" gets routed around. That is what happened.
The second layer
The new layer answers a different question. Not does this string match a phone pattern, but would a careful reader say this document exposes a specific living person.
That is a judgement about meaning, in an arbitrary script, which is what a language model is genuinely good at and what a regular expression genuinely is not. So the certifier is a model call over the same public projection the deterministic scan walks, against a fixed category list built around identifiability, not around sensitive-sounding topics:
person_name · contact_detail · identity_document · financial_instrument · credential · precise_location · health_or_biometric · protected_attribute · employment_record · other_personal_data
The distinction the prompt works hardest to hold: an essay about depression is fine, a named person's diagnosis is not. A politician's policy is fine, their home address is not. John Doe and [email protected] are placeholders, not people. And the author counts: a CV exposes the person who wrote it, which is the case that started all of this.
On the Turkish CV that motivated the work, the deterministic layer returns {}. The certifier returns person_name, contact_detail, precise_location, employment_record at confidence 100.
Why it is a second layer and not a replacement
The certifier can be wrong in both directions, so the two layers compose asymmetrically and adaptively:
- A deterministic hit always blocks. The model never overrides it. An AWS key is an AWS key regardless of what a language model thinks the document is about.
- The small model runs first. Any personal-data finding stops immediately; there is no reason to buy another opinion before applying protection.
- A cheap verdict of clean at confidence 92 or above finishes the bounded human-view screening in one call. Ambiguous or lower-confidence results escalate to the larger model.
- An unavailable first model cannot be replaced by one clean second opinion. In the escalated path, both answers must exist and clear the confidence floor. A disagreement, timeout, malformed response, or low confidence produces
review_incomplete, never a search-indexable certification. - Abstention is failure. There is no path where "the certifier didn't finish" resolves to publishable.
The asymmetry is deliberate and it is not a close call. A wrongly-masked span is an inconvenience its owner can appeal. A wrongly-published CV is a stranger's phone number on the open web, indexed, forever.
The same asymmetry decides what gets stored. The audit table records category names, counts, a content hash, and the verdict. For selective masking, a second table stores only digest-bound character offsets for the exact field and revision. The model's verbatim evidence exists only long enough to locate those offsets; it is not written to the database. An audit log full of other people's phone numbers would recreate the exact exposure the audit exists to prevent.
Earning the word "validated"
Here is the part that was tempting to skip.
The publication guard refuses to serve a project as verified unless the recorded certifier version string starts with validated:. The cheap move is to ship the model call and write validated: into the constant. It would work. Every test would pass. Nobody would notice.
Instead the prefix is earned by evidence on disk. scripts/calibrate-pii-certifier.mjs pulls labelled rows from ai4privacy/pii-masking-300k, which is multilingual by construction. That is the entire point, since the layer it backs up already handles English-shaped identifiers and fails on everything else. It runs the real certifier, unchanged, over labelled positives, and over negatives built from the corpus's own masked text: same language, same register, no real person. Then it writes recall, precision, and per-language recall to data/pii-certifier-calibration.json.
The server reads that file, checks it is for the current certifier version, and requires recall ≥ 0.95 over ≥ 100 samples. Only then does the version string become validated:…:r0.970:n500 and only then can a clean verdict promote anything.
Until that file exists, the certifier still runs and can still block, because an uncalibrated instrument is allowed to raise an alarm. A clean adaptive screen may make an exact, digest-bound file available to a human reader, but it cannot award a conversation the validated: settlement required for raw search indexing. Change the prompt, change the version, and every project drops back to unverified until the calibration is re-run. The corpus is the yardstick, never a runtime dependency: nothing in the server imports it, and the certifier behaves identically whether or not it has ever been downloaded.
Where it runs
The deterministic scan is called from inside addMessage. It has to stay synchronous and instant; putting a network round trip in the path of every message write would be a bad trade for everyone.
So a clean local scan parks the project at review_incomplete with the reason multilingual_certifier_pending, and an out-of-band pass starts shortly after boot and checks again every 30 seconds. The revision is captured before the model call and re-checked in the UPDATE, so a message that lands mid-certification cannot have a stale verdict applied to it. The write simply misses, and the next pass certifies the new revision.
Files take an even cheaper path before any model call. The deterministic recognizers run first, and an exact path, size and content digest can reuse a previous verdict. A harmless workbook therefore does not get re-screened every time it appears in another project. A changed byte produces a new digest and a new decision.
The route itself is recorded. Publication scans are selected through the same two-level execution graph the rest of the runtime uses, and the second layer now appears in that graph as an admissible-but-costlier option tagged deferred_async_pass. The graph therefore records that the deterministic route was chosen over an available alternative, rather than because nothing else existed. That was true before and is not true now.
What the owner sees
Nothing is hidden from the person who wrote it. Not one character. A flagged conversation renders for its owner exactly as it always did, with no redaction bars, no removed files, no edits.
What changes is a notice above the composer, and the copy is doing specific work:
Personal details protected from other viewers The public conversation stays readable in Browse, but detected personal spans and affected files are masked for everyone else and kept off search engines. The surrounding conversation remains visible. You still see the original, and nothing was deleted or edited.
Four rules sit behind that wording. Lead with what happened to the reader of the notice, who is the owner. The first instinct on seeing "protected" is to assume the work was destroyed, so say that the surrounding conversation remains visible. Distinguish other viewers from the owner. Name the mechanism as automatic, because "we flagged you" is a different sentence from "a check ran", and only the second one is true. And never quote the finding: printing the phone number back at the user, in a banner that might be on screen in a shared room, would defeat the whole feature.
What this does not fix
Old public conversations stay public and stay labelled legacy_unverified. They were published under the old rules, they remain readable, and their owners can still see and change them. Retroactively hiding them would be a different decision with its own costs; retroactively certifying them would be a lie.
The certifier is a model, and models are wrong. Recall of 0.95 against a synthetic corpus is not recall of 0.95 against a real Indonesian cover letter or a Persian thesis draft, and the per-language breakdown in the calibration record exists precisely so that gap is visible rather than averaged away. A verdict of "clean" is a defensible judgement backed by a measured instrument. It is not a proof, and the code does not describe it as one.
And the deeper limit stays where it was: this decides whether a conversation may be public. It has nothing to say about what the person typed into it.