中
Chapter 37. Directing AI Performance

Part VII — Generating Images, Video, Performance, and Composites

Chapter 37. Directing AI Performance#

In this chapter
37.1 An emotion label cannot direct a performance37.2 Six performance channels37.3 Performance has an entry and an exit37.4 Directing the eyeline37.5 Micro-actions and habits37.6 Grading intensity37.7 Post-sync and lip sync37.8 Crying, laughing and shouting37.9 Performance takes37.10 From a line's action to visible behavior37.11 Reactions need a cognitive order37.12 The power differential in two-person performance37.13 Sound performance precedes mouth performance37.14 The performance continuity sheet37.15 How to review performance37.16 The suppressed crying scene in Backlit Takeover37.17 SOP for performance37.18 Fault tree37.19 Checklist, exercises and deliverables37.20 The three-layer contract of a performance beat37.21 The take matrix and choosing a performance37.22 Sound-led performance and mouth ownershipA note on sources

37.1 An emotion label cannot direct a performance#

Angry, shocked, sad — these make a model enlarge brows, eyes and mouth while the character has no specific intent. Performance comes from what a character is suppressing, what they want the other person to do, and which parts of the body leak the real state.

37.2 Six performance channels#

Eyeline decides the object of attention. Breath decides rhythm. Mouth and jaw show control. Shoulders and back show tension. Hands show impulse or habit. Distance shows the relationship. Emphasize only a few channels per shot.

performance_spec:
  character: CH_LINYUN
  objective: stop Zhou Lan seeing that the father's name landed
  emotion: shock_suppressed
  intensity: 3_of_5
  eyes: rest on the name for 0.5s, then move to the date
  breath: a short inhale, then held
  mouth: jaw slightly drawn in, mouth closed
  shoulders: kept level
  hands: stops turning the pen; the pen slips
  distance: does not step back

37.3 Performance has an entry and an exit#

A shot enters from the previous emotional state, reaches a peak through a trigger, and leaves through regulation. A model given only the peak will suddenly pull a face. Write "contained — briefly breached — control restored" as a timeline.

37.4 Directing the eyeline#

What a character looks at first, for how long, and when they look away determines how the audience reads the moment. Looking at a document, an opponent or a bystander all mean different things. Eyeline targets must agree with the spatial ledger.

Blinking is not random liveliness. A delayed blink can add control and a rapid one can show cognitive load — without specifying it mechanically in every shot.

37.5 Micro-actions and habits#

The old pen, straightening a cuff, pressing the rim of a glass can all become a character's way of self-regulating. Once repetition establishes them, the action stopping becomes meaningful in itself. Vary how motion signatures are used, so they do not become a fixed animation.

37.6 Grading intensity#

Define one emotion from 1 to 5: faint suspicion, visible defence, suppressed impact, on the edge of control, complete loss of control. A climax is not simply enlarging the features — it may mean the character changes strategy, crosses a distance, or breaks a habit.

Characters in one scene should contrast in intensity. Everyone at 5 leaves the audience no center of attention.

Figure 37-1 Five levels of performance intensity across physical channels for one character

Figure 37-1 Five intensities under one stimulus. Rising numbers are not just a more exaggerated face: eyeline, breath, jaw, shoulders and hands progressively lose control. Identity, camera, wardrobe and light stay fixed so the performance itself can be compared.

You need not generate every level for a take, and you should prepare at least low, standard and high anchors. The director chooses an intensity curve across adjacent shots rather than picking the single best-acted frame. If the first reaction already reaches 5, later scenes have nowhere to escalate.

37.7 Post-sync and lip sync#

Performance with no lip sync is the most stable and suits reactions, narration and off-screen lines. Post-synced audio can let the mouth move gently without strict sync. Accurate lip sync is for short, important frontal lines. Split long sentences into short lip-sync shots with off-screen continuation over inserts.

Lock the final audio before lip sync. After lip sync, check face shape, teeth, perceived age and emotion — not only phoneme alignment.

37.8 Crying, laughing and shouting#

Crying divides into welling tears, a tear falling, broken breathing and full sobbing. Laughing divides into polite, triumphant, concealing and out of control. Write the purpose and the body, not the result. Shouting changes faces very easily; use short lines, profile angles, or cut to the reaction.

37.9 Performance takes#

The same shot can be generated as contained, standard and heightened takes, with the director choosing against the scene's curve. The strongest is not always right. Each take records its spec version, which prevents emotional discontinuity.

37.10 From a line's action to visible behavior#

A performer does not act "I feel wronged," and a model should not either. The director finds the action behind the line first: pressuring, probing, stalling, humiliating, verifying, concealing, luring, refusing, placating. The same line — you finally came — can be a welcome, an accusation, a threat or a confirmation. Different actions produce different eyelines, distances, tempos and physical choices.

Every performance spec writes at least four layers: what the character wants the other person to do right now; what they do not want exposed; the specific information that triggers the change; and what the audience should finally read. When Lin Yun sees her father's signature, it is not generic shock. It is continuing to interrogate Zhou Lan while preventing her from noticing the signature was recognized — with the pen stopping as the leak. That gives the model and the edit an observable difference.

37.11 Reactions need a cognitive order#

A believable reaction usually passes through perception, recognition, interpretation, and then suppression or action. Microdrama's pace can compress those stages; it cannot fully reverse them. A character crying before she can have read the screen looks like a pre-made expression. A character who looks, pauses briefly, then changes her breathing produces the sense of thinking even within one second.

The director can arrange micro-performance on reaction beats: the eyeline landing on the signature and holding still; a first change in pupils and breath; the jaw tightening and the fingers stopping the pen; and finally looking up at Zhou Lan with control restored.

If a model cannot complete that in one clip, split it into an object insert, an eye reaction, the hand stopping and the opponent looking back. Editing can reconstruct the cognitive process — one facial shot does not have to carry every stage.

37.12 The power differential in two-person performance#

A two-person scene should not give both people the same "tense" instruction. Define first who owns the space, who controls time, who is allowed silence, and who looks away first. The higher-status person can usually move less, answer later, occupy the center, or intrude into the other's space; the lower-status person may over-explain, keep checking bystanders' reactions, and retreat physically. These are optional grammars rather than fixed formulas, and a character may invert them deliberately.

When multi-person frames are hard to generate, the performance relationship can still be built through coverage. A's eyeline leaving frame right must be caught by B's reverse at the same height and time. The pause left after A's line determines whether B's reaction reads as struck or as long prepared. Performance continuity therefore belongs not only to the face but to the time between cuts.

37.13 Sound performance precedes mouth performance#

Direct the sound of a key line first: which word takes the stress, where the breath falls, where the pace deliberately slows, whether the sentence closes or pushes at the end. Then choose frontal lip sync, soft profile lip sync, over-shoulder, or off screen, based on the final audio. Generating mouth movement first and forcing the voice to match usually costs the line its natural rhythm.

Split long sentences by meaning, not by a fixed word count. Keep frontal only for the attack or the promise the audience must see delivered; let the rest cross to the other person's reaction, a prop, the space or a memory. That lowers lip-sync risk and gives the line an object to act on. Acceptance for lip-sync shots includes phonemes, identity, teeth, jaw amplitude, eyes and emotion together; the mouth matching alone is not a passing performance.

37.14 The performance continuity sheet#

Establish a baseline, triggers, maximum intensity and exit per scene, so shots do not each choose their own expression:

performance_arc:
  scene: SC_E003_ARCHIVE
  character: CH_LINYUN
  baseline: controlled_investigation_2
  triggers:
    - beat: B04
      event: finds her father's signature
      change: shock_suppressed_3
    - beat: B06
      event: Zhou Lan uses an old family-only name for her
      change: fear_and_suspicion_4
  maximum: 4
  forbidden: [open_crying, shouting]
  exit: controlled_dominance_3
  carry_to_next_scene: left hand still gripping the old pen

The edit chooses takes against this sheet. A single version that is superb but reaches the scene's maximum intensity early should be rejected or moved to the real trigger. The arc matters more than any one shot being fully acted.

37.15 How to review performance#

The first pass is muted, watching body, eyeline and cut points to judge whether the characters appear to be affecting each other. The second is eyes closed, listening for intent, breath and relationship. The third is normal playback, checking whether sound and picture peak at the same moment — when every channel emphasizes at once, it usually reads as overplayed. The fourth attaches one shot either side of the scene to check that emotion does not change out of nowhere.

Review notes must be executable behavior: look up 6–10 frames earlier so she finishes reading the signature first; reduce the mouth amplitude and keep the jaw tension; do not tap the table again because the previous shot already stopped. "A bit more restrained" remains too abstract.

37.16 The suppressed crying scene in Backlit Takeover#

Lin Yun confirms her father's death is connected to the lead she filed. The failed version had her tear up immediately, cover her mouth and shake her head. The emotional information was clear and it contradicted the controlled behavior established for the character. The new version does not generate crying first: she finishes reading the date, and the old pen in her hand stops; in the second shot she asks an excessively technical question; in the third, after the other party leaves, the pen slips from her palm; only at the end of the scene does one tear appear, and she does not wipe it.

That design turns grief into a failure of control rather than a generic crying template. It also leaves layers for sound: dry, close delivery in the first half; one short inhale after the date; room tone dipping subjectively; the pen hitting the floor as the emotional landing; and music held out until the tear.

37.17 SOP for performance#

First, read the character's objective and entering emotional state. Second, write the suppression and the leak. Third, choose the emphasized channels among eyeline, breath, mouth, shoulders, hands and distance. Fourth, set intensity and a timeline. Fifth, decide between no lip sync, post-sync and full sync. Sixth, generate a small number of intensity takes. Seventh, review performance continuity inside adjacent shots. Eighth, write the emotional exit back.

37.18 Fault tree#

Symptom: expressions are exaggerated. There is an emotion label with no objective and no control. Lower the intensity and specify the physical channels.

Symptom: single shots are good and emotion jumps in sequence. emotion_in / emotion_out were not inherited. Choose takes back inside the sequence.

Symptom: lip sync is accurate and the person is wrong. Sync is changing facial structure. Shorten the sentence or move it off screen.

Symptom: the character looks like a puppet. There is no breath, eyeline or micro-motion at all. Add one low-amplitude channel of life rather than setting the whole body moving.

37.19 Checklist, exercises and deliverables#

Check that the performance has an objective; that suppression and leakage are stated; that intensity is graded; that eyelines have objects; that habitual actions carry meaning; that the lip-sync route matches the risk; that takes are chosen in sequence; and that the emotional exit is written back.

Exercise one: write "anger" as three performance specs with different strategies. Exercise two: express an unspoken line with no lip sync. Exercise three: design intensities 2, 3 and 5 for the same crying scene.

Deliverables for this chapter: the performance spec, the emotion timeline, the take matrix, the lip-sync plan, and the sequence acting review.

37.20 The three-layer contract of a performance beat#

Every performance beat divides into public intent, suppressed state and physical leakage. A model most easily plays "sadness" as a generic frown; the director should write "keep control of the meeting, hold down the fear, and after seeing the photo let the breath stop for half a second and the thumb press the phone, then look up." Visible action and timing are far more controllable than adjectives.

The contract also carries prohibitions — do not step back, do not cry audibly, do not glance at the door early. Prohibitions protect the character's strategy and the room to escalate later.

37.21 The take matrix and choosing a performance#

Rank candidates not by which is more emotional but by clarity of intent, identity stability, intensity, accuracy of action, lip sync, and continuity with what surrounds them. A director may choose the take with the most accurate performance that needs local technical repair, over the technically cleanest take with no drama in it.

Review the emotion curve across every shot in a scene side by side. Intensity 5 in one shot may be excellent, and if everything around it sits at 2 it reads as a different performer. The sequence acting review decides whether to change take, add a regulating action, or regenerate.

37.22 Sound-led performance and mouth ownership#

Lock the performance audio first, then decide which words must be seen spoken and which can continue off screen. Let a stable close-up carry the key section delivered to camera and cut the rest to reactions, hands or space; do not bind every syllable to a continuously deforming mouth.

Give one character mouth ownership in a multi-person shot and keep the others in low-amplitude reaction. Interruptions are completed by sound entering early and a cut, which avoids the identity fusion caused by two mouths syncing accurately at once. Even after lip sync passes, judge whether the eyes, breath and body look like they are saying that line.

A note on sources#

Stability in AI performance comes from observable micro-actions and short-duration control, not from stronger emotional vocabulary. Performance selection has to be completed inside the shot sequence.