Deepfake detection · CLIP ViT-L/14

Protect the truth.Detect the fake.

Every face in a video, scored frame by frame — with the attention map that produced the verdict, and an honest account of when the model could not decide.

0.942
Cross-dataset AUC
0.988
In-dataset AUC
103K
Tuned parameters
32
Frames per scan
CLIP ViT-L/14FaceForensics++ c23Celeb-DF v2LayerNorm-only tuningSCRFD face detectionArcFace identity trackingAttention rollout0.942 cross-dataset AUC103,169 trainable parametersRuns offlineCLIP ViT-L/14FaceForensics++ c23Celeb-DF v2LayerNorm-only tuningSCRFD face detectionArcFace identity trackingAttention rollout0.942 cross-dataset AUC103,169 trainable parametersRuns offline
The problem

Synthetic video is now good enough to fabricate a statement, impersonate an official, or manufacture evidence.

The tools that detect it are foreign and closed. They cannot be audited, cannot be trusted with sensitive material, and are not maintained against locally relevant threats. A detector that can be inspected, retrained and run entirely offline is a different kind of asset.

Full fine-tuning teaches a model the specific artefacts of its training set, and it then collapses on forgeries made by any other tool. DeepShield updates 103,169 of 303 million weights — LayerNorm only. Preserving the pretrained representation is what holds cross-dataset accuracy at 0.942 where CNN baselines fall to around 0.65.

The method
01

Sample

Up to 32 frames are drawn evenly across the clip, so the whole timeline is covered rather than just the opening seconds. If the first eight already agree strongly, the rest are skipped.

32 frames · even stride

Isolate

Every visible face is detected, rotated until the eyes sit level, and cropped with a 1.3× margin — the region where face-swap and reenactment artefacts concentrate.

SCRFD · 1.3× margin

Classify

A CLIP ViT-L/14 transformer scores each crop independently. Only its LayerNorm weights were tuned, preserving the pretrained representation that lets it generalise to unseen forgery tools.

103,169 of 303M weights

Explain

Attention rollout reconstructs which patches the classification token actually drew from, so the verdict arrives with visual evidence instead of a bare number.

Abnar & Zuidema rollout

Measured results

The number that matters is the one it never trained on.

FaceForensics++ test split
In-dataset
0.988
Celeb-DF v2 · official 518-video list
Cross-dataset
0.942
Re-encoded at JPEG quality 60
Robustness
0.937
Re-encoded at JPEG quality 30
Robustness
0.878
Re-encoded at JPEG quality 10
Robustness
0.670

Bars span 0.5 (chance) to 1.0, since AUC cannot fall below chance.

What it cannot do

Face-swap and reenactment only

The model was trained on swapped and reenacted faces. Fully synthetic video from newer generative systems is a different problem, and may pass as real.

Compression degrades it

Accuracy holds near 0.937 down to JPEG quality 60, falls to 0.878 at quality 30, and reaches 0.670 at quality 10. The app warns you when an upload lands in that range.

Tracking can merge people

Faces are matched between frames by identity embedding. People who cross paths on screen can still be confused for one another.

Attention is not proof

Rollout maps show where the classification token drew from. That is a noisy proxy for attribution, not an explanation of the decision.

Upload · analyse · inspect

See what the model sees.

Analysis runs on your own machine. The upload is deleted from disk the moment the result is returned.

Analyse a video