Protect the truth.Detect the fake.
Every face in a video, scored frame by frame — with the attention map that produced the verdict, and an honest account of when the model could not decide.
Synthetic video is now good enough to fabricate a statement, impersonate an official, or manufacture evidence.
The tools that detect it are foreign and closed. They cannot be audited, cannot be trusted with sensitive material, and are not maintained against locally relevant threats. A detector that can be inspected, retrained and run entirely offline is a different kind of asset.
Full fine-tuning teaches a model the specific artefacts of its training set, and it then collapses on forgeries made by any other tool. DeepShield updates 103,169 of 303 million weights — LayerNorm only. Preserving the pretrained representation is what holds cross-dataset accuracy at 0.942 where CNN baselines fall to around 0.65.
Sample
Up to 32 frames are drawn evenly across the clip, so the whole timeline is covered rather than just the opening seconds. If the first eight already agree strongly, the rest are skipped.
32 frames · even stride
Isolate
Every visible face is detected, rotated until the eyes sit level, and cropped with a 1.3× margin — the region where face-swap and reenactment artefacts concentrate.
SCRFD · 1.3× margin
Classify
A CLIP ViT-L/14 transformer scores each crop independently. Only its LayerNorm weights were tuned, preserving the pretrained representation that lets it generalise to unseen forgery tools.
103,169 of 303M weights
Explain
Attention rollout reconstructs which patches the classification token actually drew from, so the verdict arrives with visual evidence instead of a bare number.
Abnar & Zuidema rollout
The number that matters is the one it never trained on.
Bars span 0.5 (chance) to 1.0, since AUC cannot fall below chance.
Face-swap and reenactment only
The model was trained on swapped and reenacted faces. Fully synthetic video from newer generative systems is a different problem, and may pass as real.
Compression degrades it
Accuracy holds near 0.937 down to JPEG quality 60, falls to 0.878 at quality 30, and reaches 0.670 at quality 10. The app warns you when an upload lands in that range.
Tracking can merge people
Faces are matched between frames by identity embedding. People who cross paths on screen can still be confused for one another.
Attention is not proof
Rollout maps show where the classification token drew from. That is a noisy proxy for attribution, not an explanation of the decision.
Upload · analyse · inspect
See what the model sees.
Analysis runs on your own machine. The upload is deleted from disk the moment the result is returned.
Analyse a video