Part V — Continuity Engineering: Character, Space, Action
Chapter 26. Multi-Character Continuity#
In this chapter
26.1 Adding people multiplies errors#
A two-person shot does not merely add a second face. It adds height proportion, left/right position, occlusion, eyelines, costume ownership, hand contact and consistent lighting. Three or more adds multiple axes and competing attention. Never assume that stable single-character work implies stable ensemble work.
26.2 Lock each person first, then build ensemble assets#
Every character passes single-person identity, wardrobe and lighting tests first. A two-person composition image then determines height, positions and shared light — it never replaces the individual golden identities.
cast_layout:
id: CAST_LINYUN_LINWEI_BOARDROOM
characters:
CH_LINYUN:
height_ratio: 1.00
zone: door_left
wardrobe: WD_LINYUN_GREY_SUIT
CH_LINWEI:
height_ratio: 0.98
zone: table_right
wardrobe: WD_LINWEI_IVORY_SUIT
axis: LINYUN_TO_LINWEI
minimum_separation: 0.8m
An ensemble asset showing any face bleed must never enter the reference library.
26.3 The feature ownership table#
Multi-person generation frequently swaps hair ornaments, moles, colors and jewellery. Give every distinctive feature a single owner and a list of characters it must never transfer to.
feature_ownership:
left_eye_mole: CH_LINYUN
pearl_hair_clip: CH_LINWEI
old_pen: CH_LINYUN
dark_green_brooch: CH_ZHOULAN
Review does not only ask whether a feature is present. It asks whether it has migrated onto someone else.
26.4 Production order: master shot, then coverage#
Produce one spatially correct master first, fixing headcount, positions, eyelines and lighting. Then produce the singles, over-shoulders, hands, proof objects and reactions separately. The master carries the geography; the coverage carries performance and identity.

Figure 26-1 The plan on the left assigns four characters fixed seats and color slots. The master, over-shoulder, reaction and insert on the right all inherit that same space. The dashed off-screen slot means a character is temporarily out of frame — it does not mean their position and state have disappeared.
The color slots are a production marker and never reach the finished cut. They help review confirm that one character has not been replaced by a similar face in a different coverage shot, and help check over-shoulder foregrounds, eyelines and seat proportions. Multi-character continuity is about preserving relationships, not about getting everyone into every frame.
Do not reverse-engineer the space from two independently attractive close-ups; that easily puts two people in the same position.
26.5 When a two-shot is worth it#
Shared frames suit distance, power, contact and relational change. If the value of a scene is only the dialogue, shot/reverse is more stable. Important two-shots should hold to simple action: one person approaching while the other stays still; one handing something over while the other watches. Avoid both people moving substantially and speaking at length simultaneously.
In Backlit Takeover, the authorization reveal can use a Lin Yun and Lin Wei two-shot to show status, then cut to singles for the lines — without requiring sustained two-person lip sync.
26.6 Identity risk in over-shoulder shots#
The foreground shoulder has to belong to the correct person too: costume color, hair edge and height must match. Models can fuse fragments of a foreground face or hair into the background character.
The foreground can use a separate matte or a defocused asset while the background character still references their single golden image. Do not frame the over-shoulder so that half an unstable foreground face shows.
26.7 Height and proportion#
Record relative heights, heel heights and seated heights. Characters must not grow and shrink across shot sizes. A low angle changes the visual impression while world proportion stays fixed.
Use eyelines, shoulder lines and known furniture as scale references when reviewing two-person keyframes. Do not let a model generate the antagonist taller because of their status.
26.8 Occlusion and depth order#
Who is in front, who is occluded, and whose body a hand passes in front of must all be explicit. Prefer layers for contact shots: background, character A, character B, hands and props. Wrong occlusion sends a hand through a body or blends clothing.
When characters swap depth positions, do it through visible movement or by cutting to a new establishing shot.
26.9 Assigning lip sync in dialogue#
In a shared frame, let one person carry the primary lip sync while the others hold low-amplitude reactions. Two characters speaking at once can overlap in sound while the picture chooses one clear mouth and puts the other in profile, rear view or off screen.
Group shouting suits a sound layer and a wide. It does not require every mouth to sync accurately.
26.10 Background crowds#
Split extras into designated reactors and atmosphere. Give two or three people stable positions and actions; keep the rest rear-facing, defocused and low-motion. Crowd faces should not compete with the principals, and do not need cross-episode identity unless the story later requires it.
Build crowd plates and reaction versions so that headcount, seating and costume do not change completely in every shot.
26.11 Layered compositing and color matching#
Generating characters separately and compositing reduces face bleed, and requires handling contact shadows, depth of field, grain, edges, color temperature and eyelines. Two portraits that look like stickers usually lack shared light and spatial occlusion.
Compositing is not free repair. It is worth the investment on hero two-shots; ordinary dialogue is better served by shot/reverse.
26.12 Multi-character QC#
Check identity, feature ownership, height, wardrobe, position, eyelines, occlusion, contact, lighting and who is speaking, item by item. Review the stills first, then play the motion. Face bleed in even a few frames disqualifies a key close-up.
26.13 SOP for multiple characters#
First, lock each person separately. Second, build the cast layout and feature ownership. Third, generate the spatial master. Fourth, decide which information genuinely requires a shared frame. Fifth, complete performance through single coverage. Sixth, layer contact and occlusion. Seventh, limit simultaneous movement and lip sync. Eighth, tier the background crowd. Ninth, check proportion, ownership and shared lighting.
26.14 Fault tree#
Symptom: the two characters converge toward one face. Ensemble images replaced individual identity and the model averaged features. Carry each golden reference back into every shot.
Symptom: a mole or hair ornament migrates to the other person. There is no feature ownership table. Make transfer a blocking error.
Symptom: heights change between shot and reverse. There is no proportion reference, or the camera difference is unexplained. Verify against the master and furniture as a ruler.
Symptom: faces fuse during two-person lip sync. Both are speaking with complex motion. Move to single coverage with off-screen sound.
Symptom: extras change identity and seats every shot. The background is being regenerated each time. Use crowd plates and designated reactors.
26.15 Checklist, exercises and deliverables#
Check that characters are locked separately; that features have single owners; that the master fixes the space first; that shared frames have narrative value; that heights and eyelines are stable; that occlusion is correct; that simultaneous lip sync is limited; that extras are tiered; and that composites share light and shadow.
Exercise one: build a cast layout for a three-person meeting. Exercise two: convert a sustained two-person dialogue into a master plus coverage. Exercise three: design a layered composite for a wrist grab. Exercise four: assign three designated reactors and a background crowd for the banquet.
Deliverables for this chapter: the cast layout, feature ownership, the two-person proportion chart, the master shot, the coverage plan, occlusion layers, the crowd state table, and multi-character QC.
26.16 Ensemble assets are not single images pasted together#
Single-person identity assets answer who someone is. Ensemble assets answer how two people exist together. A two- or multi-person pack stores relative height, common distances, positions, shared lighting, eyeline relationships and occlusion priority. It never replaces the golden singles; each generation still references the individual identities, with the ensemble pack constraining the relationship.
Build ensemble packs by scene relationship rather than producing every permutation of the cast. Lin Yun and Lin Wei need public confrontation, close-range threat and negotiation across a table. Lin Yun and Shen Yan need side by side, protective occlusion and a distance of distrust. Produce only the compositions that are frequent and narratively important, so combinations do not grow exponentially with cast size.
Define identity slots per composition — slot A is always frame-left foreground, slot B right background — and bind each to its references. A slot is a position in this shot template, not a permanent left/right for a character. When moving to the reverse, create a new template and remap; do not smuggle the mole, hair ornament and costume details across with a horizontal flip.
26.17 A coverage protocol for group dialogue#
Divide a three-person dialogue by who speaks and who reacts. For each beat, designate the primary speaker, the person primarily affected, the relational information that must share frame, and the audio that may be off screen. The master conveys position and power. Singles carry long lip sync. Two-shots carry changes of distance. Reactions carry what is not said. A coverage protocol makes shot count serve the storytelling rather than mechanically reversing on every line.
coverage_beat:
beat: B07_authorization_confirmed
speaker: CH_ZHOULAN
spoken_line: The authority is valid. Let her in.
primary_reaction: CH_LINWEI
must_share_frame: [CH_LINYUN, CH_GUARD]
reason: the audience must see security stand down
offscreen_allowed: [CH_ZHOULAN]
lip_sync_owner: CH_ZHOULAN_closeup_or_offscreen
end_state: the doorway is clear
Keep overlapping sound separate from picture ownership. When Lin Wei interrupts Zhou Lan, her voice can enter first and the cut to her close-up follow; there is no need for accurate simultaneous lip sync on two people. The picture always tells the audience who to watch, while sound is allowed to carry conflict across the cut.
26.18 Economic tiers for crowd continuity#
Extras come in three tiers. Narrative extras need stable identity and position across shots. Reaction extras need stability only within a scene. Texture crowds supply headcount, costume color mass and motion density. Each tier gets different assets, review thresholds and rework costs. Managing every background face to principal standard destroys the budget; treating a director with a critical reaction as random texture destroys the causality.
Crowd state records headcount range, occupied zones, costume palette, overall direction of attention and motion level. When the banquet moves from conversation to a public reveal, the critical change is not what each person does — it is environmental motion dropping from 3 to 0 while attention converges on the stage. Have three designated reactors set down a glass, stop applauding and exchange a look, while everyone else simply lowers their motion in unison.
Headcount can vary slightly across shots and must not intrude on the main path, alter a principal's silhouette, or fill an empty table suddenly. Crowd plates, reaction versions and an ambience bed together maintain the scene's density. When background generation is unstable, prefer local variation on one approved plate over re-creating the whole banquet for every close-up.
26.19 Production breakdown of the three-person meeting#
The Lin Yun / Lin Wei / Zhou Lan meeting section starts with a three-person master with no lip sync, locking Lin Yun by the door, Lin Wei near the screen and Zhou Lan at the head of the table. Singles follow for each of them, with backgrounds inherited from the matching camera zone. Once the authorization enters, add a two-shot of Lin Yun with the document and a close-up of Zhou Lan confirming it. Lin Wei's reaction is completed separately rather than chasing synchronized performance inside the three-person frame.
The first version failed on an over-shoulder: Lin Wei's pearl earring migrated onto Lin Yun's ear and the foreground shoulder cut into the document. The fix did not keep retrying the whole two-person image. It used an approved matte of Lin Wei's back and shoulder as foreground, Lin Yun's single close-up as background, and the document as a separate table layer; contact shadow and depth of field were unified and it passed. Feature ownership then marked the pearl earring as a P1 blocking item, so later automated checks search for misattribution first.
The sequence ultimately keeps only two genuine multi-person hero shots; the rest of the performance is carried by singles and off-screen sound. The audience still knows all three are present, because the master, the eyelines, the background anchors and the sound space stay consistent. The sense of an ensemble comes from continuous relationships — not from packing everyone into every frame.
A note on sources#
Crowds and ensembles are a high-failure zone for AI generation. Traditional coverage thinking — space carried by a small number of masters, performance carried by stable singles — remains the more reliable compromise between quality and cost.