A woman climbs a church wall on all fours while a pastor, on the stage below, commands her. The clip spread widely, and the fact-checker Factly found that no camera recorded it. It was AI-generated, and that is how AI ghost videos are made. A text-to-video model draws every frame of a short clip at once, starting from noise, steered by a sentence of text. You tell one from a recording by looking where the model is weakest: not at the ghost, but at the room around it.

The rest of this page is the long version. What these models do, where their weakness shows on screen, which labels to look for, and the order in which to check a clip someone sends you.

How a text-to-video model draws a ghost

The tools have names you have probably seen: OpenAI’s Sora, Google’s Veo, Kuaishou’s Kling, Runway’s Gen models. They work on the same broad principle. The model begins with a block of noise the size of the whole clip, every frame of it, and removes that noise over many steps until what is left matches the prompt.

The prompt can ask for a figure on a staircase. It can also ask for the look of the camera that supposedly caught it: a ceiling-mounted CCTV view with a timestamp in the corner, a phone in night mode with grain in the shadows, a camcorder with its washed-out color. The fixed, timestamped view is the one Paranormal Activity (2007) built a whole film around. The models imitate those looks well, because they have seen a great deal of footage that looks like that.

What they have not done is build the room. A camera records light bouncing off a real staircase, so the staircase holds still between frames because it is there. A video model has no staircase. It has a statistical sense of what staircases look like, applied frame by frame and kept roughly consistent across the clip. Most of the time, roughly is enough to fool a phone screen. The tells live in the places where it is not.

Where the room gives it away

Every tell below comes from the same gap: the model tracks appearance, not objects. Look for these, and look for them away from the figure, because the figure is where the model spent its attention.

  • Counts that change. A hand has five fingers in one frame and six in the next. A doorframe gains a second edge. A window has four panes, then three. Limbs and objects merge into each other and separate again.
  • Shadows that disagree with the light. Find the lamp or the window. The shadows in the frame should fall away from it, and they should all agree. In generated clips they often point in different directions, or a figure casts none.
  • Text that will not hold still. Timestamps and camera overlays are where CCTV imitations tend to slip. Digits garble, change shape, or count in an order no clock uses. Signs on the wall turn into letter-shaped marks.
  • Textures that swim. Carpet, wallpaper, brick and foliage should be fixed to their surfaces. In a generated clip they can crawl or shimmer while the camera is supposed to be still.
  • Physics that drifts. A door swings, but the hinge side never moves. An object falls at the wrong speed, or stops where nothing stops it.
  • Framing that is too good. A real security camera is bolted where the wiring allowed, usually at an awkward angle. A generated one tends to be centered, level and symmetrical, with the subject exactly where a director would put it.
  • Blur that belongs to one object. A figure’s head smears into motion blur while the rest of the frame, including things that are also moving, stays sharp. A real camera applies the same shutter to everything in the frame.

None of these is proof on its own. Real footage has compression smears and odd angles too, and the page on glitches that look like ghosts goes through those one by one. What marks a generated clip is several tells at once, clustered away from the subject.

Labels, credentials and watermarks

Some of the answer may already be attached to the file. Some platforms label uploads “Made with AI,” either because the uploader declared it or because the platform detected it. Some models embed content credentials under the C2PA standard, a signed record of how the file was made that travels in its metadata. Google embeds an invisible watermark called SynthID in video generated with Veo.

Each of these is useful when present and weak when absent.

SignalWhere it livesWhy its absence proves nothing
“Made with AI” labelOn the platform, declared or detectedLabels depend on the platform
C2PA content credentialsIn the file’s metadataLost when a clip is downloaded, cropped, re-encoded or screen-recorded, which is how clips usually travel
SynthIDAn invisible watermark in video generated with VeoInvisible by design; stepping through frames will not find it

Treat a label as a strong reason to stop, and the lack of one as no reason at all.

Three cases from the past year

The church wall clip is the first. According to Factly’s fact-check, it was first posted on October 20, 2025, by an Instagram account that regularly posts AI-generated clips. That history alone is a reason to look twice, before a single frame.

The second is a household scene. A cat walks past, and a woman is thrown down a staircase as if by something no one can see. Lead Stories fact-checked it in August 2026, and the clip carried a “made with AI” label. The label was there. It did not stop the clip from being shared as a haunting.

The third is a still image, and it had consequences off the screen. In June 2026, police in Palembang, Indonesia, detained a 24-year-old man who had made and spread an AI-generated image of a pocong, the shrouded ghost of Indonesian folklore; the picture had gone viral and alarmed residents of the city’s Gandus district. The case is recorded in the OECD AI incident database and was reported by Antara. Generated photographs have their own tells, set out in how an AI ghost photo gives itself away.

How to examine a clip someone sends you

Work in this order. The early steps are cheap and often settle it.

  1. Ask for the original file. Not the repost, not the screen recording. The original carries whatever metadata survived, and its resolution lets you see detail that a compressed copy has smeared away. Ask who shot it, where and when.
  2. Check for labels and credentials. Look at how the platform labels the post, and look at the uploader’s other posts. An account that posts generated clips every day is a finding in itself.
  3. Watch the edges and the background. Play it once looking only at the corners of the frame, the ceiling, the floor and whatever is behind the figure. That is where the model was paying the least attention.
  4. Count. Fingers, doorframes, window panes, stair treads, chair legs. Count them in an early frame and again in a late one.
  5. Step through frame by frame. Most video players and editors can advance one frame at a time. Watch the timestamp digits, the shadows and the hinge side of any door. Motion that looks natural at full speed often falls apart when you stop it.
  6. Reverse-image-search a still. Take a clean frame and run it through a reverse image search. You may find the original post, an earlier version with a watermark, or the same scene in another clip.

If you have a clip you cannot place after all six, send it to us. We look at it the way every piece in the On Camera section looks at a clip, and we say what we can see.

What a tell cannot tell you

A tell tells you about the file, not about the house. A clip can be generated and still have been made about a real place, with a real story attached that was told long before the model existed. And a real recording can carry so much compression that it fails half the checks above.

So the question at the monitor stays narrow: did a camera record this? Sometimes the frame answers it. Sometimes the doorframe has four edges, then three, then four again, and the figure in the doorway never casts a shadow at all.