Back to blog
Video16 min read

Physics Failures in AI Video (And How to Prompt Around Them)

AI video physics fails in specific, testable ways: water, cloth, collisions, gait. What a real physics benchmark measures, and the prompting habits that reduce each one.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: AI video physics fails in specific, testable ways: water that moves like one syrupy mass, cloth that clips through bodies, objects that pass through each other, and feet that slide instead of planting. None of it is solved. The best current approach on a real physics benchmark reaches barely half the maximum score. Prompting narrows each failure. It doesn't remove any of them.

Why Does AI Video Get Physics Wrong?

Because there is no physics engine anywhere in an AI video model's pipeline. Every AI video physics failure on this page, water, cloth, collisions, gait, traces back to that one gap. A diffusion-based video model has never run a fluid solver, never resolved a collision between two rigid bodies, and never checked a footfall against a ground plane. What it has done is watch an enormous amount of real footage and learn which next frame plausibly follows a given one, then repeat that guess forward for the length of the clip.

Whether that process amounts to understanding physics, or just to convincingly faking it, is the question a team from Google DeepMind and INSAIT set out to test. Predicting the next frame of a real video is, they argue, "impossible... if the model has no understanding of how objects move (trajectories), that things fall down instead of up (gravity), and how pouring juice into a glass of water changes its color (fluid dynamics)" (Motamed et al., "Do generative video models understand physical principles?", arXiv:2501.09038, read September 4, 2026). Their benchmark, Physics-IQ, is built from 396 real, high-resolution recordings of genuine physical events, scored by how closely a model's predicted continuation matches what the real footage did next.

The headline result, from the original study: across Sora, Runway Gen 3, Pika, Lumiere, Stable Video Diffusion and VideoPoet, the paper found "a massive gap to the physical variance baseline, with the best model scoring only 29.5% out of the possible 100.0%." A separate, sharper finding matters more for anyone prompting today: the researchers "find no significant correlation between visual realism and physical understanding". A gorgeous frame is not evidence the next one will make sense.

One example from the paper makes this concrete, and it is a genuine fluid-dynamics failure, not a rendering glitch. Asked to continue a clip of a burning match lowered into a glass of water, "Runway Gen 3 generates a continuation where as soon as the flame touches the water, a candle spontaneously appears and is lit by the match." As the researchers put it: "Every single frame of the video is high quality in terms of resolution and realism, but the temporal sequence is physically impossible." That is hallucination applied to physics: not a wrong fact stated in text, but a wrong sequence of cause and effect rendered as video.

What follows is the same failure, broken into the four places the editorial brief for this piece asked about by name: water, cloth, collisions and gait. Each gets its own mechanism and its own prompting lever, and each lever is something you can test on your own clip in one rewrite.

Why Does Water Look Wrong in AI Video?

Because a fluid's behavior across time, not its look in one frame, is what a model is actually guessing at, and time is where the guess breaks down soonest. The common symptoms: a splash that appears fully formed with no build-up and no aftermath, a surface that moves as one connected, slightly viscous mass instead of scattering into independent droplets, and a liquid that fails to mix or change color the way it should when something is poured into it, exactly the "juice into a glass of water" case Physics-IQ's own researchers named as a basic test of fluid-dynamics understanding.

No video vendor publishes a fluid-behavior parameter of any kind, the same way none publishes a general motion-amount dial; our motion-strength reference covers that absence in full. So the only lever here is naming the fluid behavior itself, as a specific word, tied to the exact thing causing it, in one clause rather than two separate ideas.

Before: A wave crashing on a beach, dramatic, powerful ocean.
After:  A wave crests and breaks against the rocks, water sheeting off
        in a fan of spray as it recedes down the sand.

Before: Someone pouring red wine into a glass, elegant, close-up.
After:  Red wine streams from the bottle's neck and swirls into the
        glass, darkening the water already in it as the two mix.

Before: A boat moving fast through calm water, cinematic wake.
After:  A speedboat cuts across the water, its hull throwing a curling
        white wake that fans out and flattens behind it.

The pattern in every after line is the same one the site's no-motion diagnosis established for subjects in general, applied specifically to a fluid: name the verb the water performs, not an adjective describing water. Dramatic ocean gives the model a mood. Sheeting off in a fan of spray gives it something to render. Keep the interaction to one beat, too. A splash, a pour, a wake, one continuous event, asks less of the model than a whole sequence of fluid interactions crammed into five seconds, and our duration reference by model has the per-vendor ranges if the clip needs to be longer than the default.

Why Does Cloth Clip Through Bodies or Move Like Stiff Plastic?

Because cloth simulation, like fluid simulation, is not happening anywhere in the request. Runway's own schema is explicit that motion of any kind, on the branches that accept a text field at all, has to be written into the prompt itself, described as "motion or changes in the output video" rather than set as a parameter. That is the entire mechanism for fabric too. A generic flowing dress or billowing cape gives the model an adjective with nothing to attach it to, and the two most common failures follow directly: cloth that clips straight through a limb because nothing enforced the two surfaces staying apart, and cloth that hangs rigid because the prompt never named a force acting on it.

The fix is the same discipline the water section used, applied to material and force instead of fluid behavior: name the fabric's weight, and name the specific thing moving it, in the same clause as the action.

Before: A woman in a flowing red dress walking through a garden.
After:  A woman in a lightweight silk dress walks through a garden, the
        hem lifting and swaying with each stride as a breeze catches it.

Before: A superhero's cape billowing dramatically in the wind.
After:  A heavy wool cape snaps and trails behind a figure sprinting
        forward, catching the wind at the shoulders and lagging a beat
        behind the turn.

Before: A knight in armor with a cloak, standing still, epic.
After:  A knight turns to face the camera, the cloak's thick canvas
        swinging a half-second behind the shoulder's rotation before
        settling.

Naming lightweight silk versus heavy wool canvas is not decoration; it is the only signal a model has for how fast fabric should respond to a force, since no current vendor exposes a material-physics field to set instead. Be honest about the ceiling, too: if a source image already shows a sleeve clipped through an arm, prompting cannot un-clip a frame that has already committed to the wrong geometry. That is the same anchoring limit the morphing and warping piece covers for identity drift, and it applies to fabric geometry just as much as to a face.

Why Do Collisions and Falling Objects Look Wrong?

Because nothing in a video model checks whether two solid objects can occupy the same space, or how fast something should fall. Physics-IQ's own scenario design names this directly: the dataset tests "collisions... trajectories under the influence of forces (e.g., gravity), material properties and reactions" as core categories, not edge cases. Two solid objects overlapping mid-frame, a dropped object hitting the ground with no deceleration, an impact with no reaction from the thing it hit: all three are the model predicting a plausible-looking frame without any constraint that two things can't share a location.

The match-and-water example from earlier is really a collision failure in disguise, a physically impossible chain of cause and effect rather than a static glitch, and it is worth remembering here: high per-frame quality is not evidence the sequence between frames holds together.

The prompting lever that actually helps is the same one our sports and action motion piece already uses for contact shots, just applied with physics in mind rather than framing in mind: name the impact and its immediate physical consequence in the same clause, rather than naming only the moment of contact. Two cars collide is an event with no result attached. The car's hood crumples inward and the bumper sheers off on impact gives the model a cause and an effect it can render together, because the reaction is stated, not implied.

Before: A glass falling and shattering on the floor.
After:  A glass tips off the counter, falls, and shatters on the tile,
        fragments skidding outward from the point of impact.

Before: Two cars colliding at an intersection, dramatic crash.
After:  A car strikes another broadside at the intersection, the hood
        crumpling and the struck car's rear end swinging out from the
        force.

Before: A stack of boxes falling over.
After:  A stack of boxes tips past its balance point and collapses,
        the top box tumbling free as the rest scatter across the floor.

The honest limit: this gets worse, not better, as the number of colliding objects goes up, for the same reason a video with more subjects drifts more, covered for identity in the morphing piece. A single dropped glass is one physical state to track. A collapsing stack of boxes is several, each with its own trajectory the model has to keep consistent with every other one in the same frame, and no vendor publishes a rigid-body or collision-detection setting that would do that tracking for you.

Why Do Feet Slide Instead of Planting When a Character Walks?

Because the visual symptom borrows a name from an older problem it isn't actually the same as. Animators have called this footskate since at least Kovar, Schreiner and Gleicher's 2002 paper on cleaning it out of motion-capture data, where it describes a foot that moves when it's supposed to be planted, a real artifact from editing footage that was captured with an actual, physically constrained skeleton. AI video has no such skeleton to violate in the first place. What you're seeing is the same frame-to-frame guessing process that produces water and cloth failures, applied to legs, and legs in contact with a moving ground plane happen to be an unusually hard case for it: the foot has to appear to stop dead for an instant while everything else keeps moving, and nothing enforces that stop.

Two things help, and one sidesteps the problem entirely rather than reducing it. First, name the footfall itself as an event, the way the collision section named an impact and its consequence: her heel strikes the pavement, weight rolling forward onto the ball of the foot gives the model a contact moment to render, where walking naturally does not. Second, keep the clip to about one stride. A full gait cycle is a repeating structure, and asking for several strides in a short clip is the same overcrowded-clip problem our no-motion diagnosis already names for any subject: more beats than the duration can hold, each one finishing worse than the last.

Before: A man walking down a street, natural and smooth.
After:  A man's foot strikes the pavement, weight rolling from heel to
        toe as his other leg swings forward for the next step.

The genuine workaround, where it's available, isn't a wording trick at all: driving the walk from a real captured performance instead of asking the model to invent one. Kling's Motion Control, which is image-to-video and distinct from its Motion Brush feature covered in our no-motion piece, does exactly this: "Motion Control enables precise control of a character's movements and facial expressions based on a reference image... The motion can be extracted from an uploaded video or selected directly from the motion library" (Kling AI, Motion Control User Guide, read September 4, 2026). That doesn't ask the model to guess where a foot should land from a sentence; it copies the footfall timing from footage where a foot actually did land. It's the closest thing on this list to a real fix rather than a mitigation, and it's only available on the vendors and workflows that document motion transfer, not as a general prompting habit.

Which of These Four Actually Responds to Prompting?

Unevenly. Know which end of that range you're on before you spend an afternoon rewriting.

Four physics-failure categories, the mechanism behind each, and the prompting lever that narrows it. No vendor field addresses any of the four directly.
FeatureWaterClothCollisionsGait
Typical failureSplashes appear or vanish with no build-up; surface moves as one massFabric clips through the body, or hangs with no response to motionSolid objects overlap; an impact has no visible reactionFeet slide or float instead of planting on contact
What's actually happeningNo fluid solver; next-frame prediction onlyNo material-physics field on any current vendorNo collision check between generated objectsNo skeleton or ground-contact constraint exists at all
One lever that helpsName the fluid behavior and its cause in one clauseName the fabric's weight and the force moving itName the impact and its immediate consequence togetherName the footfall as an event; keep the clip to one stride

Read across the middle row and the pattern is the same everywhere: nothing here is a parameter you can flip, because no vendor ships one. That puts every failure on this page closer to a capability limit than to a wording bug, which is a different balance than our no-motion diagnosis found for static clips, where most causes were prompt problems a rewrite genuinely fixed. Physics splits the other way: prompting narrows how often you see the failure. It does not remove the underlying gap Physics-IQ measured.

Does Any Model Actually Solve This?

No, though the trend line is real rather than flat. The Callout above already has the dated numbers; the honest summary of them is that the best achievable score roughly doubled since the original paper, but the gain came from stacking multi-frame generation and best-of-N sampling on a base model, not from a better single prompt. Whatever you type into a consumer app is still closer to the lower end of that range than the top.

Test your own shot rather than trust a vendor's showcase reel, the same advice our morphing and warping piece gives for identity drift. A water shot that holds up on one model can fail on another with the same prompt, because the gap Physics-IQ measures is architectural, not something a wording change reaches.

Do You Need a Paid Prompt Architects Plan for Video Prompts?

Video Prompt Generation is an Advanced and Team feature. Our pricing page shows it unavailable on the Free and Pro columns and available on Advanced and Team, with Advanced at $9.99 a month at the time of writing, verified in the plan comparison table in September 2026. The Free plan includes 5 prompt enhancements per day, forever, per our FAQ page.

We build the prompt. We don't run the model, and nothing about a well-built prompt gives a diffusion model a fluid solver or a skeleton it doesn't have. What Prompt Architects does is turn make the water look better into the specific, causally-attached wording this piece argues for, saved once so the next water, cloth, collision or walking shot starts from something that already earned its physics, rather than from a blank prompt box and a hope.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account