Guide

How to Spot a Deepfake โ€” Deepfake Detector Guide (2026)

Consumer-grade tools now produce convincing face-swap and lip-sync video in minutes. Here is what to look for by eye, where manual review runs out of road, and how an automated detector combines per-frame and temporal signals.

A deepfake is any synthetic or manipulated video in which the face, the voice, or both have been replaced or generated by AI. The most common forms in 2026 are face-swap, lip-sync, and fully synthetic footage of a person who never recorded a frame. Open-source toolkits have collapsed production time from weeks to minutes.

That shift is why spotting a deepfake matters well outside newsrooms and security teams. The same models behind a harmless creative effect today power a scam call tomorrow, and the person receiving it is usually an employee or a parent with about thirty seconds to make a judgement.

The short version

Fake CEO video calls, synthetic testimonial ads and AI-generated news clips โ€” different tells in each case, none reliably caught by visual inspection alone.

Three patterns that already cost money

1. Fake CEO video calls

An employee receives a video or voice call from someone presenting as the CFO or a board member. The voice matches publicly available earnings calls; the face, when shown, matches conference B-roll. The ask is always urgent โ€” a wire transfer, a change of vendor bank details, an expedited credential issue. In a 2024 Hong Kong case that ended in a $25M loss, every other participant on a multi-person video call was synthetic; the victim was the only real person in the meeting.

2. Synthetic testimonial ads

Advertisers clone the faces and voices of creators and experts to attach false endorsements to crypto schemes, supplement brands, and apps promising fast returns. The clips run fifteen to forty-five seconds, the cadence of a paid social ad, and are cheap enough to produce in volume. Takedowns arrive after the campaign has rotated to a new link.

3. AI-generated news clips

Synthetic anchors and on-the-scene reporters now air on low-quality aggregator channels and circulate as breaking news on social platforms. The damage concentrates in the first ninety minutes, before a major outlet either confirms or debunks the footage.

How to spot a deepfake by eye and ear

Manual review still catches a meaningful share of low- and mid-quality deepfakes. Watch the clip full-screen at 1:1, not at thumbnail size, then work through this checklist:

  • Face-edge artifacts. Face-swap leaves a subtle halo or seam at the hairline, along the jaw, or around the ears.
  • Lighting and colour. A swapped face often carries the source video's lighting into the destination frame, producing a face lit differently from the scene around it.
  • Lip-sync. Watch the corners of the mouth and the teeth on "m", "b" and "p" phonemes. Mismatched teeth and blur during fast speech are common tells.
  • Micro-expressions. Synthetic faces flatten involuntary blinks, brow twitches and cheek raises; technically perfect but emotionally inert is suspect.
  • Audio context. Studio clarity on a claimed phone call, or zero ambient noise on a claimed street interview, means the audio was rendered out of context.
Visual inspection catches sloppy deepfakes. The convincing ones require a detector that scores per-frame artifacts, temporal consistency and audio-visual alignment together.

How a deepfake detector works

A video detector scores several independent signals rather than returning one unexplained verdict. Four do most of the work:

  • Per-frame visual artifact scoring flags texture, colour, edge and frequency-domain inconsistencies. A high score sustained across hundreds of frames is evidence; one bad frame is usually compression.
  • Temporal consistency compares frames across the clip, catching head-pose drift, flicker and lighting changes invisible in any single still.
  • Audio-visual cross-reference checks the audio track against the visible mouth movements โ€” the strongest signal available against lip-sync attacks on video calls.
  • Provenance metadata, where present, records the capture device and edit history. A valid credential alongside a clean signal report is the closest thing to positive evidence of authenticity that exists.

How those signals are combined

We do not flag a video because one detector crossed a threshold. Any detector returning within ยฑ8 points of 50 is treated as abstaining and excluded from the calculation entirely โ€” a model saying "I don't know" should not drag a confident result toward the middle. Remaining votes are weighted by confidence multiplied by a per-detector reliability factor, so a provenance credential counts for more than a pixel-level inference. Where confident detectors genuinely disagree we surface the split rather than averaging it away, and every detector's raw score is always shown.

How to use this guide

If you have a video you are not sure about, run it through the scanner in Video mode and read the per-signal report rather than the headline verdict. Then verify through a channel you already trusted yesterday โ€” a detector is evidence, not authorisation.

Check a suspicious video in about a minute

Five detection modes โ€” text, images, video, speech and documents. Three free checks a day in total across all modes, no signup required.

Try it free