Part III — Episodes, Scenes, Dialogue, and Cut Points
Chapter 12. Writing Action That Can Be Performed and Generated#
In this chapter
12.1 Visible is not the same as generatable#
"She slaps him" is a visible action to a human crew. To a generative model it contains preparation, contact, transferred force, a head reaction, hand position, and consistency across two faces. Action writing has to satisfy three things at once: it must be narratively visible, performable, and controllable by the model.
A script does not have to become a technical prompt. It does have to supply action state clear enough that shot design can choose to break it up, occlude it, substitute it, or composite it.
12.2 Turn interiority into evidence#
"Lin Yun realizes her father may still be alive" cannot be photographed. Ask first what physical evidence, change of attention and consequent choice the internal change produces.
The payee name appears on the ledger. Lin Yun stops turning the pen. She enlarges the screen and checks the payment date. The pen slips from her fingers and strikes the table. Without looking at anyone, she says to freeze the account.
Nothing there says shocked. Stopping a habitual movement, re-checking, losing her grip and giving an immediate order carry the impact between them. Visible evidence holds up far better than an adjective across different models and different performances.
12.3 The four stages of an action#
Record the before state, the initiation, the completion and the after state of every key action. A great many continuity errors come from writing only the middle.
action_id: ACT_RAISE_AUTHORIZATION
before:
right_hand: document held against the edge of the table
left_hand: gripping the old pen
body: standing beside the door
initiation: right hand pushes the document toward the center of the table
completion: lifts the document to chest height, face turned toward the chair
after:
right_hand: holding the document's top edge
left_hand: pen still in the palm
body: one step forward toward the table
story_effect: everyone sees the title of the authorization for the first time
The after state becomes the next shot's before state. If the action is cut mid-way, the cut point must also carry a shared state.
12.4 One primary change per shot#
"Lin Yun walks into the boardroom, dodges security, throws the file on the table and turns to glare at Lin Wei" contains five classes of change: locomotion, interaction, throwing, turning and expression. Even when a model produces the motion, it will struggle to hold identity, hands and space together.
Break it down: the door is pushed open; Lin Yun steps in and stops; security reaches out; she turns her body past them; the file lands on the table; Lin Wei looks up; Lin Yun turns to her. The edit lets the audience assemble short actions into one continuous event.
One action per shot is not an absolute ban on two motions. Breathing, fabric movement and a slow push-in can be subordinate motion. Keeping the primary change singular makes the causal center easier for the model to understand.
12.5 Action verbs must be concrete#
Avoid handle, respond, express, display, argue fiercely. Use verbs with a visible start and end: push, press down, tear, block, hold, step back, look up, stop, exchange, close, turn to, point at.
Also avoid over-literary phrasing about a gaze cutting the air like a blade. The producible version: she looks at the signature on the report, then raises her eyes to Zhou Lan, without blinking.
Action should also state object and direction. Not "she picks up the document," but "her right hand lifts the authorization from the table, face turned toward the chair." Direction affects downstream continuity and text legibility.
12.6 Performance intensity must be gradable#
"Very angry" invites a model to generate exaggerated features. Break intensity down by physical system: eyeline, brows and eyes, jaw, breath, shoulders and back, hands, and voice.
performance:
emotion: restrained_anger
intensity: 2_of_5
eyes: steady gaze, no avoidance
mouth: lower lip lightly compressed, no shouting shape
breath: a short hold after the inhale
shoulders: kept level
hands: thumb pressing the pen clip
voice: no increase in volume, falling at the end of the line
Intensity 5 should not simply mean every part doing more. Genuine loss of control can show as a break in the rhythm of movement, a change of objective, or crossing a stated line.
12.7 Reactions are not decoration#
Reactions serve four purposes: confirming the information was understood; showing the consequence of a power shift; providing cover for the edit; and giving the audience time to complete the emotion. A crowd being "shocked" should be broken into chosen reactions: laughter stops, one person sets down a glass, two directors look at each other, Lin Wei keeps smiling while her fingers crush the napkin.
Do not answer every reveal with a room full of staring eyes. Design reactions that fit each character's interest: whoever gains observes, whoever is threatened denies, the neutral party withdraws, the one who already knew looks away.
12.8 Design hands and direction for prop actions in advance#
Proof objects are a frequent difficulty in microdrama. A document, a phone, a ring or a glass must not only appear; you need where it came from, which hand holds it, who it faces, and when it is put down.
Prefer static inserts or post-composited text for anything important. The performance shot can show a character raising a report; the next shot cuts to a separately produced close-up of the report; then the reaction. Do not ask a moving hand, legible text and a stable face to succeed in the same generation.
12.9 Two-person contact#
Handshakes, wrist grabs, embraces, slaps and kisses all involve physical contact, occlusion and identity blending, and they fail often. Decide at the writing stage whether the action truly has to be complete in one frame.
A wrist grab can be split into the hand approaching, a close-up of contact at the wrist, the other body stopping, and the two reactions. A slap can use the initiating swing, the struck person's reaction, the impact sound and a bystander's response, without ever showing full contact. An embrace can cut from a single approaching to a rear silhouette, then complete the emotion with close-ups of hands and faces.
If the contact itself is the story's proof — someone covertly exchanging a key, say — then budget for hand compositing and continuity.
12.10 Controlling attention in crowd scenes#
"Everyone at the banquet stands up at once and surrounds her" produces identity drift and disordered motion. Layer the crowd: foreground principals complete the primary action; two or three designated mid-ground characters react; the background group performs only a low-amplitude unified response.
Establish the group state first, then use single coverage for the key reactions. Extras do not need legible faces. Depth of field, backs turned and occlusion are proper photographic language and they reduce the generation burden at the same time.
Crowd action can also use sound against picture: the room inhales while the frame stays on Lin Wei; then cut to a glass being set down; then a brief wide. The audience will complete the "everyone is stunned" for you.
12.11 Narrative and compliance de-escalation for violence#
Violence has to account for platform rules, safety and generation stability simultaneously. Favor intent, result and emotion over sustained contact detail. An impact heard through a door, a dropped phone, a character staggering into frame — these are frequently stronger than a fully generated assault.
Tag each violent action with necessity, intensity, whether minors are involved, degree of gore, and the alternative representation. Action must not exist merely to raise stimulus; it has to change risk or a character's options.
12.12 Three-layer breakdown for fantasy and combat#
Fantasy action splits into base character action, an effects layer, and environmental response. Let the character complete a stable pose first; separately generate or composite light, smoke, debris and impact; then build force with sound and cutting.
fantasy_action:
base_character: the character raises the sword and holds above the right shoulder
effect: a blue-white arc spreads along the blade toward frame left
environment: sleeve pulled back, dust pushed outward along the ground
timeline:
0-1s: eyes lift and grip tightens
1-2s: the blade cuts outward
2-3s: the arc leaves the blade
3-4s: the opponent blocks and retreats
Do not ask for running, drawing, rolling, casting, striking several people and a collapsing building in one request.
12.13 Intimacy and micro-performance#
Intimacy is not improved by more contact. Change of distance, a held gaze, fingers pausing briefly when handing something over, a drop in speaking volume — all build tension. AI is usually more stable on small, single movements than on complex contact.
Write consent and initiative explicitly, and avoid shots that package coercion as romance. For brand and platform projects, the acceptable level of intimacy should also be settled in advance in the brief and the compliance sheet.
12.14 Action risk tiers#
Low risk: looking up, turning the head, setting down a document, a phone vibrating, a slow push-in. Medium: handing an object over while walking, two people in frame speaking, sitting and standing, a simple wrist grab. High: fast contact, multi-person motion, fighting, complex lip sync, mirror reflections, continuous costume changes.
High risk does not mean forbidden. It means the shot needs breakdown, a candidate budget, compositing and a fallback. The shot ledger should show the risk tier before generation begins.
12.15 SOP for action writing#
First, convert interior lines into visible evidence. Second, write before, initiation, completion and after for key actions. Third, keep one primary change per shot. Fourth, specify hand, object and direction. Fifth, break emotion into physical and vocal intensity. Sixth, design reactions that fit each character's interest. Seventh, layer contact, crowds, violence and fantasy. Eighth, tag risk and the alternative representation. Ninth, have shot design and the generation lead review feasibility together.
12.16 Fault tree#
Symptom: the frame moves but nobody can read it. The action has no object, direction or story result. Go back to the state change.
Symptom: hands and props keep jumping. Before and after states, or the cross-shot handoff, were never written. Build the action state and lock left and right hands.
Symptom: exaggerated expressions, everyone staring. You supplied an emotion label with no physical grading and no interest-based reactions. Rewrite the performance spec.
Symptom: two-person contact keeps failing. You treated full contact as a single-shot task. Split it into initiation, partial contact, result and reaction — or use occlusion.
Symptom: crowd scenes are expensive and chaotic. There is no foreground/mid/background layering. Limit the designated reactors and keep the background to low-amplitude motion.
Symptom: the effect buries the character. Base action, effect and environment were not layered. Approve the performance without effects first, then composite.
12.17 Checklist, exercises and deliverables#
Check that interiority has visible evidence; that key actions carry four stages; that each shot has one primary change; that hand, direction and object are explicit; that reactions match character interest; that important text is composited separately; that complex actions have a breakdown and a fallback; and that risk has entered the budget.
Exercise one: convert ten interior sentences into action. Exercise two: write three versions of a slap — full contact in frame, broken into shots, and without showing contact at all. Exercise three: reduce a ten-person crowd scene to three designated reactors and a sound crowd. Exercise four: write a three-layer timeline for a four-second spell shot.
Deliverables for this chapter: the action state, the performance spec, the action risk table, the contact action breakdown, the crowd layering diagram, the violence substitution table, and the three-layer fantasy timeline.
12.18 Action lines must compile into state events#
A production-grade action line contains subject, starting point, trigger, action, object, world direction, completion condition and state consequence. "She walks over angrily" is missing distance and target. "After Zhou Lan confirms the signature, Lin Yun moves from the door along the south side of the long table to the empty seat, stops, and leaves the document at the center of the table" can be decomposed into spatial and prop events.
Abstract interiority can still live in the performance intent layer, but it cannot replace visible behavior. After each line, ask: what does the keyframe draw, what does the video move, what does continuity update, and what happens in the sound? An action that cannot answer all four has not yet entered production language.
12.19 The minimum narrative form of a complex action#
Find the meaning the action cannot do without, then choose the smallest visible form of it. A glass being smashed may mean a public loss of control, and the audience only needs the glass raised, the sound of breaking, and the room stopping. You do not need to generate the physical trajectory from hand to floor. If the fragments later become evidence, then where they land does have to be clear.
Prepare A/B/C routes for high-risk actions: full performance, elided in the cut, or replaced by sound and consequence. The director approves narrative equivalence first; the producer then chooses by success rate. That prevents a critical meaning being deleted in a panic after generation fails.
12.20 Safety, compliance and performance boundaries#
Violence, intimacy, minors and dangerous action get risk tags, a shooting method and a list of prohibited content at the script stage. AI generation does not remove ethical and distribution responsibility; it can unexpectedly produce stronger imagery than intended. Prompts, references, review and publication all follow the project's policy and the applicable rules.
De-escalating an action does not mean writing the conflict weaker. It means moving the visible emphasis: express harm through preparation, sound, reaction, environmental consequence and the aftermath state. Whenever a real performer, likeness or voice is involved, confirm permissions and protective measures as well.
A note on sources#
One action per shot, short image-to-video durations, approving stills first, and timeline-structured prompts recur throughout Chinese AI video practice as reliable experience. This chapter connects that experience to action continuity, performance and the interface with editing.