SardineCon SF/2026

Learn More

What is Deepfake?

SUBSCRIBE

A deepfake is AI-made fake audio, image, or video used to impersonate a real person. Fraudsters use it to beat selfie or liveness checks, pass video ID reviews, or approve a payment over a call, so any control that trusts a face or a voice as proof can now be fooled.

What is a deepfake, plainly?

A deepfake is synthetic media, audio, an image, or video, generated by AI to look and sound like a specific real person. The name comes from the deep learning models that create it. Early deepfakes were crude and easy to spot, but the technology has moved fast, and today a convincing fake can be produced from a handful of photos or a few seconds of recorded speech.

In fraud, a deepfake is a tool for impersonation. Wherever a system or a person accepts a face or a voice as evidence of identity, a deepfake attacks that assumption directly. It can present a fabricated selfie to a liveness check, drive a fake video stream through an ID review, or clone a voice to authorize a transfer over the phone.

The key shift for operators is that seeing and hearing are no longer believing. Controls that were built on the idea that a live face or a familiar voice proves who someone is have a new, well-resourced adversary. That is why deepfakes cut across onboarding, authentication, and payment approval rather than living in one narrow place.

Where deepfakes attack

The same underlying tech shows up against several different controls:

  1. Onboarding — Beat the selfie check. A generated face or face swap is presented to liveness and photo-match during account opening.
  2. Verification — Pass a video ID review. A manipulated video stream defeats a human or automated review that assumes a real camera.
  3. Authentication — Clone the voice. A voice clone gets past voiceprint login or convinces a call-center agent it is the real customer.
  4. Approval — Fake the authorizer. An executive's face or voice is faked on a call to push through a high-value wire or a policy change.

What it looks like in practice

In practice

A finance clerk at a mid-size company joins a video call that appears to be the CFO and two colleagues. The faces move, the voices sound right, and the CFO calmly explains an urgent, confidential acquisition that needs a wire sent before end of day. Everything on screen looks normal, so the clerk sends the funds.

Only afterward does the picture come apart: the real CFO was traveling and never joined a call. The faces and voices were deepfakes, built from public recordings and conference footage, stitched into a live meeting. The one control that would have caught it was not on the screen at all; it was an out-of-band callback to a known number, which the process did not require.

Why it breaks old controls

The problem is that deepfakes attack the evidence itself, not the process around it. A liveness check that trusts a moving face, a call-center script that trusts a matching voice, an approval flow that trusts a familiar person on video, all of them assumed the media was genuine. Once the media can be faked, those checks quietly turn into weak points.

The defense is to stop treating any single piece of media as proof and instead layer signals the attacker cannot easily fake together. That means liveness detection, injection-attack detection that spots media fed straight into the camera feed rather than captured by a real lens, device integrity checks, and out-of-band confirmation for risky actions. Crucially, do not lean on the human eye or ear; good fakes now pass both, so the decision has to rest on evidence a person cannot judge by looking.

What to watch in the data

  • Visual tells, with caution. Odd lighting, warped edges around the face, unnatural blinking, or lip-sync that drifts, though quality is climbing and clean fakes exist.
  • Audio without a room. Cloned voices often lack natural background noise, breathing, or echo, sounding oddly flat or studio-clean.
  • Injected feeds. Signs that video is being piped into the camera stream rather than captured live, a hallmark of injection attacks.
  • Device and telemetry mismatch. Emulators, virtual cameras, or sensor data inconsistent with a real phone held by a real person.
  • Urgency plus a face. High-value requests pushed through on the strength of a video or voice alone, especially with pressure to skip normal steps.

Quick questions

Can I trust my own eyes to spot one?

No, not reliably. The old visual giveaways are disappearing fast, and good fakes now pass human review. Treat the eye and ear as weak signals and rely on liveness, injection, and device checks instead.

How is a deepfake different from a face swap?

A face swap is one technique for making a deepfake, replacing one face with another. Deepfake is the broader term that also covers fully generated faces, voice clones, and manipulated video.

What is injection-attack detection?

It looks for media being fed directly into the camera or video feed rather than captured by a genuine camera in real time. Because many deepfake attacks inject a prepared stream, catching the injection catches the fake.

Does passive liveness stop deepfakes?

Passive liveness that just studies a still selfie can be spoofed. Pairing it with active challenges and feed-integrity checks is far more robust than any single passive check on its own.

What stops a deepfake CEO call?

Out-of-band confirmation. A callback to a known, trusted number or a separate approval channel does not care how convincing the face or voice was, because it verifies through a path the attacker does not control.

Are deepfakes only a video problem?

No. Audio-only voice clones are among the most common and effective, driving vishing and call-center fraud. Any control that trusts a voice as proof is exposed just as much as a video check.

Go deeper

What to know alongside Deepfake