A generated video has to be right thousands of times in a row. A generated image only has to be right once. That is the whole reason video is still easier to catch than a still — and it is the principle behind every technique below.
There are two different problems people call deepfakes, and they fail in different ways. A fully generated clip was painted from a text prompt by a model such as Sora, Kling, Veo or Runway; nothing in it was ever filmed. A face swap is a real recording of a real person with someone else's face composited on. Knowing which one you are looking at changes what you should check.
Signs of a fully generated clip
Temporal drift in the fine detail
Step through the clip a frame at a time and watch the small things rather than the subject: earrings, shirt buttons, teeth, the pattern of a tie, text on a background sign. In a real recording those stay identical between frames. In generated video they subtly reshape — a button migrates, a stud earring becomes a hoop, the sign's text changes letters. Detail that re-decides itself every frame is the single most reliable tell.
Physics that almost works
Watch anything that should have weight and momentum. Liquid pours slightly too slowly, or the level in the glass does not match what has been poured. A dropped object decelerates before it lands. Fabric moves as if it weighs nothing. Cars in the background travel at a speed that does not match their apparent distance. Feet slide microscopically instead of planting.
Crowds and hands, still
Hands are fixed on the main subject, but background people still lose fingers, merge shoulders and walk through each other. Count people crossing the frame, then count them again on the other side.
The camera behaves unnaturally
Real footage has handheld micro-jitter, breathing, focus hunting and rolling-shutter skew on fast pans. Generated footage often glides on a perfect virtual dolly, or jitters in a way that is too regular. Watch the edges of the frame during a pan: real parallax means near objects move much faster than far ones.
Signs of a face swap
- The boundary. Look at the jawline, hairline and ears. A swap has to blend at an edge, and that edge flickers, blurs or shifts in colour when the head turns.
- Head-turn breakdown. Swaps are trained mostly on frontal faces. Wait for a sharp profile turn or a hand passing in front of the face — that is when the mask slips.
- Lip-sync drift. Watch the mouth on plosives (p, b) and at the end of sentences. A few frames of lag is normal in re-encoded video; a mouth that closes on an open vowel is not.
- Lighting mismatch. The face is lit differently from the neck and the room. A hard light from camera-left should put a shadow under the nose on the same side as under the chin.
- Blink and gaze. Swaps still under-blink and hold gaze slightly too steadily.
What the file itself tells you
Everything above is what you can see. The file has a second, quieter story:
- Container structure. A phone camera writes a predictable arrangement of atoms with a device-specific encoder tag and a creation date. Generated and re-wrapped clips write a very different, tool-specific set.
- Encoder mismatch. A clip claimed to be straight off an iPhone but encoded by a desktop tool has been through at least one edit.
- Duration and frame rate. Many generators output fixed lengths at fixed frame rates. A run of clips that are all exactly the same length is a pattern worth noticing.
- Audio that does not match the room. Generated or dubbed audio has a noise floor that does not match the visible space. See our guide on detecting AI voice clones.
A practical workflow
- Get the original file, not a forwarded or re-downloaded copy. Every re-encode destroys evidence.
- Hash it immediately, before you do anything else, so you can prove later that the clip you analysed is the clip you received.
- Watch it once at normal speed for the overall impression, then again at quarter speed watching only the background.
- Step through any moment where the subject's head turns or a hand crosses the face.
- Run an automated check that inspects container metadata and asks a multimodal model for a second opinion.
- Write down what you found and when. If the clip ever becomes evidence, the record matters as much as the verdict.
What a detector can and cannot promise
Be sceptical of anyone quoting a single accuracy figure for video. Results depend enormously on length, resolution, compression and how many times the clip has been shared. A 4-second clip re-uploaded through three platforms has lost most of its forensic signal, and any tool that still returns 99% confidence on it is guessing with conviction.
What a good check gives you instead is evidence: the metadata trail, the specific observations, the engine whose signature traits best match, and a hash that lets someone else confirm they are looking at the same file. Run one on your clip below, or open the deepfake video detector.