Pick a clip and a configuration, then compare it with the original. Every output was produced offline by the
custodian_video prototype on our own GPU server; this page only plays the precomputed files.
In processed outputs faces are blurred or boxed, and any replacement names, numbers or voices are synthetic.
The PII planted in the inputs (names, phone numbers, IDs) is invented.
Accuracy at a glance
Measured on the six demo clips against references that are independent of the pipeline's own detections. Hover a metric (or see the notes) for how it was measured.
| Metric | Blur default · safest |
Blur + fake text/speech faces blurred · local fakes |
Blur + Custodian Guardian Layer faces blurred · Guardian Layer values |
|---|
1 · Choose a clip
2 · Choose how to protect the video
Each button is a different masking method. Blur (default) and Solid box mask every detected face and licence plate and black out the personal words on screen; spoken personal data is bleeped. Blur + fake text/speech also blurs faces, but replaces on-screen and spoken personal data with realistic fake values of the same type (the same fake value on screen and in speech). Blur + Custodian Guardian Layer does the same, but the replacement values come from the Custodian Guardian Layer, Custodian's own transform; a local fake value fills in only where Guardian Layer's value is missing or incomplete. Turn on Show what changed to see every masked region.
blurred / boxed face licence plate on-screen PII text spoken PII (banner)
Audit log summary
Counts from the run's JSON audit log. The log itself never stores raw PII (only hashes and masked previews); this page shows counts only.
Accuracy of this output
This clip and configuration only, measured against independent references (see "Accuracy at a glance").
Planted PII caught
Fake values planted in the input. Screen: not OCR-readable in any sampled frame of the output. Speech: not in a Whisper re-transcript.
Known misses and limitations
Generated from the same measurements (video/accuracy.json); nothing here is hand-picked.
How it works
Two passes over the file: analyse everything first, then render. Every model runs locally from weights on disk; only the Guardian Layer configuration calls the Custodian Guardian Layer API.
Faces & plates
- EgoBlur (Apache-2.0) detects faces every frame and plates every 2nd frame.
- Tracking links boxes into tracks, fills gaps and pads each track so faces don't flicker back.
- Faces and plates are always masked with a blur or a solid box. Faces under 80 px get extra mask padding, and faces the YuNet recall net finds are masked too. (Replacing faces with synthetic ones is a research direction we evaluate internally; it is not shown here because its visual quality is not yet good enough.)
- Person filter (fewer false masks): a face track is left unmasked only if a COCO person detector (torchvision Faster R-CNN, run on the full frame and again zoomed in) finds no person's head there on any detected frame, its score is below 0.5, it is away from the frame edge, the YuNet recall net never saw it, and it was seen on at least 2 frames (or sits on an animal). Anything uncertain stays masked; every unmasked track is listed in the audit log.
On-screen text
- OCR (RapidOCR / PP-OCRv5) every 0.5 s.
- PII detection with a local LLM (qwen2.5:7b) + regex.
- Scene-context gate: a local VLM (Qwen3-VL-8B) decides if the text is a document/screen/badge (mask) or public signage (keep).
- Only the PII words are boxed, or re-typeset with a fake value of the same type.
- Blur + Custodian Guardian Layer: OCR lines also go to the Custodian Guardian Layer, which rewrites personal data into realistic values. Its flags add a value only if it contains a digit or "@" or overlaps one the local detector found, so field labels and ordinary words are never changed. Its value for a span is used only if it replaces the whole value with one of the same type and format; otherwise the local fake value is used. Each value's source is logged per span.
Speech
- Whisper (faster-whisper small.en) transcribes with word timestamps.
- The same PII detector finds spans.
- Redact: 1 kHz bleep. Transform: the fake value is spoken by local TTS in the same slot, matching what is shown on screen. In the Guardian Layer configuration the transcript also goes to Guardian Layer, and its value is spoken whenever it passed the same check.
Every run writes a JSON audit log (track ids, frame ranges, boxes, scores, actions, surrogate used; PII only as hashes and masked previews). Nothing leaves the machine, except in Blur + Custodian Guardian Layer: there the OCR lines and the speech transcript are sent to the Custodian Guardian Layer API. The clips are public Creative Commons videos and the planted PII is invented; Guardian Layer can also be deployed on-prem.
481-clip evaluation · privacy at scale
The large-scale benchmark behind the demo (custodian_video/eval, docs/video_masking_survey.md §9). Face channel only.
Masking coverage, leaks and re-identification at scale
Same 481 clips (Kinetics-400 400, MOT17 21, VoxCeleb2 60), face channel only; mean of per-clip rates with 95% clip-bootstrap CIs. Reference faces: YuNet on the original (≥ 24 px). Lower is better except coverage. These numbers were measured before the person filter was added; the filter only removes masks from low-score tracks with no person's head nearby, so it does not raise coverage and is not reflected here.
| Arm | Mask coverage | Leak | Same-place re-ID | Impostor floor | Rank-1 of 100 | Cross-video ID |
|---|
Privacy holds in every masking arm. On VoxCeleb2 (60 probes, 118-identity gallery), cross-video identification falls from 100% on the original to 0–1.7% rank-1 after Blur. Person detection on MOT17 is essentially unaffected (AP50 within ±1.2 points). These 481-clip numbers were measured before the person filter was added.