Adjacent shots must share facts, not just share a look

Cross-Shot Continuity for AI Video

From the state ledger, action phases, spatial axis, eyelines and editing handles through to J-cuts and L-cuts: a cross-shot continuity workflow for AI video that actually cuts together.

Continuity is not "a similar look"

Two shots can use the same character and the same grade and still refuse to cut together. Viewers simultaneously track where people stand, screen direction, eyeline targets, what is in someone's hand, how far an action has progressed, the direction of the light, the ambience and the intensity of the feeling. If any of those facts jumps without cause, the shots stop belonging to the same event.

The most robust approach is to give every shot a state_in and a state_out: which facts it starts from, and what it has changed by the end. The next shot reads only the results the previous one committed; it does not re-guess them from a prompt.

Split action into phases

Do not write "she picks up the file and walks to the door." Split the action into:

  1. Preparation: her gaze lands on the file, the shoulders begin to lean.
  2. Initiation: the hand leaves the desk, reaching toward the file.
  3. Contact: fingers press the edge of the folder.
  4. Peak: the file leaves the desk and arrives at chest height.
  5. Release: the body turns toward the door.
  6. Aftermath: others react, and ambience or music takes over.

The previous shot can cut out after contact, and the next can cut in as the file leaves the desk. If both shots perform the whole pickup, the edit repeats itself; if one stops at preparation and the other has already reached the door, the action in between disappears.

Maintain the spatial axis and the eyelines

Draw a minimal scene map first: door, window, table, seating, principal light sources and the camera zones you may use. A two-person conversation establishes a 180-degree axis, and each character records an orientation and an eyeline target. When you need to cross it, use a neutral frontal, a move across the line, an explicit establishing shot or character movement to re-establish the space.

Multi-person scenes especially need a directed "who looks at whom." A reaction shot generated on its own with no eyeline target gets a studio-style straight-to-camera look, and once cut in the character seems to be looking into another world.

Leave handles for the edit

Beyond a shot's target duration, keep a brief stable entry and exit. Do not let the picture collapse the instant a performer finishes; leave usable stillness before a camera move begins. Handles let an editor adjust rhythm, match action, insert a reaction or hide a generation artifact.

Sound handles matter just as much. Room tone, fabric, breath and the tail of a line should be able to continue across a cut. A J-cut lets the next scene's sound arrive early; an L-cut lets a line run past the picture change. Both usually preserve continuity better than a dazzling morph.

Define the narrative relationship before choosing a transition

Before you pick a transition, answer: is time continuous, does the location change, whose subjective experience dominates, and does the audience need impact or flow?

Relationship Common solution What acceptance checks
One continuous action Hard cut, match on action Action phase and screen direction
A look reveals its object Eyeline match Gaze angle and the order of revelation
A jump in time or place Sound bridge, establishing shot, graphic match Whether the audience reorients immediately
Entering a subjective memory Sound first, controlled dissolve, texture change Whether the subjectivity is unambiguous
High-speed ellipsis Montage, insert shots Whether the causality is still readable

One continuity acceptance pass

Watch muted first and check action, space, light and props. Then close your eyes and listen, checking dialogue, ambience, music and rhythm. Finally put them together and check whether attention lands on the right information at the right moment. When you find an error, label it as identity, state, space, action, light, sound or editing, and send it back to the matching workstation.

Frequently asked

Why won't two individually good shots cut together?

Because they may not share action phase, screen direction, eyelines, spatial relationships, lighting state, a sound bed, or enough editing handles.

Can an AI morph transition fix a discontinuity?

Usually not. A morph can mask a small visual difference; it cannot repair wrong narrative causality, spatial direction or character state.

How should an action that crosses shots be generated?

Split it into preparation, initiation, contact, peak, release and aftermath, then fix the previous shot's exit and the next shot's entry — rather than letting two models each perform the whole thing.