Video models produce impressive four-second clips. Ask for twenty seconds of continuous action and objects start to melt. This is not a defect you can prompt around — it is a consequence of how frames are predicted.
Where the errors come from
- Error accumulation. Each frame is generated with reference to previous ones, so small inaccuracies compound.
- Identity drift. Faces and clothing shift gradually across a long take.
- Physics approximation. Objects pass through each other; liquids behave unconvincingly.
Designing around it
- Build sequences from multiple short shots, the way conventional editing already works.
- Cut on movement so the transition hides the join.
- Use image-to-video with a strong starting frame to lock composition.
- Keep camera motion deliberate rather than letting the model improvise.
The practical result
Treated as a shot generator rather than a film generator, current models are genuinely useful for previsualisation, transitions and b-roll. Accepting that framing produces far better output than fighting it.
Comments (0)
Log in to join the discussion
Log InNo comments yet