Can AI Fake a Wake-Up Verification Video?

A technical look at whether current AI video tools can beat a short accountability check-in, covering how liveness detection works, what Sora-era generators can do cheaply, and why the real defense against a faked wake-up video isn't better forgery detection.

In this article8 sections

Yes, an AI can generate a video that looks like someone waking up. Whether it can generate one convincing enough to fool the three or four people actually watching it, at the price and speed a fraud would need, is a much narrower question, and in practice the answer is usually no, for reasons that have more to do with who’s watching than with what the model can render.

The question “can AI fake a wake-up verification video” splits into two different questions. One is pure capability: can current text-to-video and face-swap tools produce a clip of a person getting out of bed convincing enough that a stranger, or a machine, would accept it as real? Mostly yes, and that’s been true for a couple of years. The other is the question that matters for an app like DontSnooze, where missing a video check sends an embarrassing photo to your friends instead: can someone generate a clip good enough to fool the small group of people who already know what they look and sound like at 6:40 a.m., on a morning that wasn’t scheduled in advance, in the ten to thirty seconds before they react to it? The second question is far harder to answer than the first, and far easier to defend against without touching a line of detection code.

What Liveness Detection Checks For

Liveness detection is the term the biometrics industry uses for confirming that a face in front of a camera belongs to a live person present at that moment, rather than a photo, a screen replay, a mask, or a synthetic feed. It’s what runs when a bank’s onboarding app asks a new customer to blink, turn their head, or read a number aloud before opening an account. Underneath, it’s checking a short list of physical properties that are expensive to fake convincingly: micro-movements in skin and eyes, depth information from a front camera, blood-flow patterns visible as faint color shifts in the face (a technique called remote photoplethysmography), and tight timing between a requested action and the response.

Two things about that list matter here. First, none of it is about identity. Liveness detection asks “is this a live human right now,” not “is this the right human.” A separate face-match step handles identity, usually against a photo ID. Second, it’s built for a single structured interaction: one face, one camera angle, one short challenge, judged by an algorithm with no prior relationship to the person in frame. That works well for opening a bank account once, a single structured moment with no history behind it. A group chat reviewing the same person’s face every morning for months is a different kind of test entirely.

Why a 10-30 Second Unscripted Clip Is a Different Problem

A driver’s-license selfie check has to work for millions of strangers it has never seen before, using nothing but an ID photo as the reference, which is the easier version of the problem for an attacker: the verifier has almost no personal context to compare against, no sense of what the person’s voice sounds like groggy, no idea what their real bedroom looks like or whether the same pillow is always on the left.

An accountability app’s check-in flips nearly every one of those conditions. The viewers aren’t strangers; they’re friends who’ve seen dozens of these clips already and have built up a working sense of what “just woke up” looks like on this one person. The moment isn’t scheduled or scripted; the alarm time is whatever got set the night before, and there’s no retake once the clip sends. And it has to hold up to someone who might text a follow-up two minutes later. Beating the kind of deepfake liveness detection alarm apps rely on, in other words, usually has little to do with beating an algorithm, since there typically isn’t one in the loop; the real obstacle is a person’s built-up, detailed memory of another person, which is a much larger and messier dataset than a single ID photo, and one an attacker can’t obtain by asking someone to smile for a license.

What Current AI Video Tools Can Actually Do Cheaply, in 2026

The tools improved fast. OpenAI’s Sora 2, released in 2025, generates convincing short clips from a text prompt or a reference photo for roughly ten cents a second through its API, about a dollar for a ten-second clip at standard resolution. It can place a real person into a generated scene through a feature called Cameo, but only once that person uploads their own likeness and agrees to it; Sora won’t build a Cameo of someone else from scraped photos, and it lets people revoke any video that uses their face. That guardrail is a big part of why Sora isn’t the tool someone would reach for to fake being a different, unconsenting person waking up.

The tools used for identity fraud today sit outside that guardrail. iProov, a company that builds face-verification technology for banks and governments, tracks attack tooling in its annual threat report and found generative face-swap attempts against remote identity checks rose 704% from the first half to the second half of 2023, driven mostly by three free or cheap apps: SwapFace, DeepFaceLive, and Swapstream. These aren’t cinematic tools. They’re built to map a source face onto a live or recorded feed in close to real time, well enough to survive the kind of casual glance a liveness check or a distracted reviewer gives it. It’s a meaningfully different threat model than a Hollywood-grade fake of someone’s bedroom. It’s cheap, fast, and good enough to pass a glance. It isn’t yet good enough to survive a follow-up question from someone who already knows a face at 6:40 a.m., and that gap matters more than it sounds like it should.

An Original Way to Weigh This: the Fraud Ledger

Most discussion of AI-faked verification treats it as an arms race: fakes improve, so detection has to improve, forever. That framing skips the variable that decides whether faking happens in practice. A more useful way to size up any short verification clip, whether it’s a bank’s onboarding selfie or a wake-up video sent to four friends, is to weigh three numbers against each other rather than fixate on one.

Generation cost is what it takes, in money, time, and skill, to produce a fake that could plausibly pass. Verification cost is what the reviewer has to spend, in attention or expertise, to catch it. Exposure cost is what happens to the person faking it if they get caught. Line those three up for any verification setup and you get a fraud ledger: a bank’s onboarding selfie has low generation cost (a swap app and a stolen photo), low verification cost if the check runs on autopilot, and low exposure cost, since an account gets frozen rather than an actual person finding out. This is a bad ledger for a defender, and it’s exactly why banks keep spending heavily on detection.

A wake-up video sent to a small group runs those numbers almost in reverse. Generation cost is higher than it looks, because a convincing fake needs material drawn from that particular person at that hour rather than just a recent photo, and it has to survive someone who’ll notice if a voice, a room, or an expression is even slightly off. Verification cost is low: the reviewer isn’t running a scan; they’re glancing at a friend’s face for two seconds, which is a surprisingly efficient way to catch small wrongness a machine wouldn’t flag. Exposure cost, if caught, is high and immediate, not a form getting rejected somewhere, but a person you have breakfast with finding out you faked getting up. Better detection only earns its keep when the first two numbers sit close together and the third is small. Here, the gap between generation cost and exposure cost is already wide before anyone builds a single detector, which is a large part of why this kind of fraud stays rare on its own.

So How Do Accountability Apps Detect Fake Videos?

Mostly, no software is doing the detecting. DontSnooze’s own product FAQ is unusually blunt about this: there’s no sensor confirming a body left a bed, no algorithm scoring the footage, nothing technical stopping someone from filming thirty seconds of nothing from under the covers. What exists instead is a person on the other end who already carries a working model of what a friend looks like right after getting up, built from every prior morning they’ve watched. That solves a harder version of the same problem than liveness detection does, and it happens to be the version an AI-generated fake struggles with most, since a diffusion model has no access to the particular texture of somebody’s ordinary mornings the way a close friend does.

Whether a given viewer looks closely enough to use what they know is a separate question, and not always a flattering one for the design. A 47-morning personal log of who actually reacts to these videos and who doesn’t found that more than half landed with no visible response at all. Human review only works as a check if the human is paying attention, and on plenty of mornings, they aren’t.

The Same Fight Is Already Happening in Banking, With One Real Difference

Financial identity verification has been fighting nearly this exact battle for years, at far higher stakes. Know-your-customer video checks, the kind used to open a brokerage account or clear a large transfer, exist specifically because criminals had already gotten good at faking selfies and static ID photos. In January 2024, a finance employee at the engineering firm Arup wired $25.6 million after a video call in which the company’s CFO and several colleagues turned out to be deepfakes, built from footage scraped off past webinars and company videos, according to Hong Kong police, who disclosed the case the following month. That employee wasn’t fooled by a low-effort clip; he was fooled by a live, responsive conference call with several familiar faces at once, assembled from real public source material.

The Arup case is a genuine preview of where this kind of fraud is headed, and it’s tempting to read it as proof that any video check, wake-up videos included, is already a lost cause. The comparison breaks down in one important place: the Arup employee had limited independent context on his CFO’s daily habits and no one else in the room to compare notes with in real time, while a small accountability group has both, every day, at no extra cost. A finance worker sees his CFO in curated calls a handful of times a year. A close friend sees an unfiltered, half-asleep face on most mornings, which is a dataset no attacker can buy or scrape.

Where the Product’s Own Check Is Weak, and Why the Fix Isn’t More Detection

DontSnooze’s video check-in is not deepfake-proof, and pretending otherwise wouldn’t hold up. There’s no liveness algorithm running, no challenge-response step asking someone to say a random word or hold up a number of fingers, no depth sensor confirming a real face is in frame. Someone with enough footage of themselves, a face-swap tool, and a group of friends who don’t look closely could, in principle, get a fake past a check-in more than once.

The instinct, staring at that gap, is to reach for more forgery-detection technology: a liveness SDK, a randomized on-screen prompt, a depth check using the phone’s own camera. Almost none of it would help as much as it sounds like it should, for the same reason banks keep losing money to deepfakes despite running exactly those checks: technical detection has to work for millions of anonymous strangers, so it can never use the one signal that matters most here, which is one person’s accumulated familiarity with another person’s face. What the fraud ledger points to instead isn’t a better detector. It’s keeping the review group small and the relationship real enough that exposure cost stays high, which is precisely the setup DontSnooze already runs on: a handful of people who know someone, not a crowd or an algorithm, on the other end of the video.

What Stops Someone From Faking It

What stops someone from faking it is a friend who’d notice, and who the person faking it would rather not have to explain themselves to. This is a different kind of security model than the one biometrics companies sell, and it happens to be the model that scales worst for a bank serving ten million anonymous customers and best for a check-in among four people who already know each other. AI-generated video will keep getting cheaper and more convincing for generic use. The systems staking their defense on detecting the fake will keep needing new detectors every year. The ones staking their defense on making the cost of getting caught by someone who matters higher than the effort of just getting up won’t need to.

Keep reading