中
Chapter 35. Producing Still Keyframes

Part VII — Generating Images, Video, Performance, and Composites

Chapter 35. Producing Still Keyframes#

In this chapter
35.1 The keyframe is the cheapest high-value gate35.2 The input package35.3 Layered prompts35.4 Composition before detail35.5 Candidate strategy35.6 Local repainting35.7 Keyframe phases35.8 Props and text35.9 Upscaling and texture35.10 Keyframe QC35.11 The production questions one keyframe must answer35.12 More references is not better35.13 Three fidelity levels for composition exploration35.14 First and last frames need a reachability review35.15 The authorization shot in Backlit Takeover35.16 SOP for still frames35.17 Fault tree35.18 Checklist, exercises and deliverables35.19 What a keyframe approves is a geometric contract35.20 Candidate diversity comes from hypotheses, not from volume35.21 The keyframe evidence package and downstream readsA note on sources

35.1 The keyframe is the cheapest high-value gate#

Before any video starts, a still already exposes problems with identity, wardrobe, space, props, light, composition and text safe areas. Sending a failed still into video generation only makes the error expensive.

A keyframe is not a concept mood image. It is a producible frame of one shot ID at one phase of its action.

Figure 35-1 Six production gates from composition to a cuttable clip

Figure 35-1 A shot passes through composition, keyframe, first/last frames, motion generation, layered compositing and a cuttable clip. Each step must pass its own gate. When motion fails, return to the keyframe or the shot design — do not pass the error on to compositing and editing.

The chain explains why "generating video" is only one segment of shot production. Composition settles information and space. The keyframe settles identity and static state. First and last frames verify that the motion is reachable. Motion generation completes only the necessary change. Compositing protects text and complex interactions. The edit finally selects the stable usable region. Until the upstream step is approved, the expensive steps that follow should not start.

35.2 The input package#

Each frame reads the style bible, the character's golden references, wardrobe state, the location and camera zone, prop state, lighting state, the shot grammar, the platform safe areas, and the continuity entry state.

keyframe_request:
  shot: E001_S09
  phase: action_completed
  identity_refs: [CH_LINYUN_FACE_RIGHT45@v03]
  wardrobe_refs: [WD_LINYUN_GREY_SUIT@v02]
  location_refs: [LOC_BOARDROOM_TABLE_INSERT@v04]
  prop_refs: [PROP_AUTHORIZATION_OPEN@v03]
  composition: the authorization occupies the lower center; the hand enters from the left
  overlay_safe_area: document title and signature area
  continuity_from: E001_S08_state_out

35.3 Layered prompts#

Assemble project style, asset identity, current state, shot composition, lighting and negative constraints separately. Do not copy forward one ever-lengthening universal prompt.

Ask the model only for a clean, replaceable plane where text is concerned — never for correct content. Load only the negatives relevant to this shot's high risks.

35.4 Composition before detail#

The first round settles character position, shot size, eyeline and prop proportion cheaply. The second locks identity, wardrobe and background. The third repairs hands, faces, text planes and edges locally. Only then come upscaling and texture unification.

Chasing pores and 8K at the outset spends the time on a composition that may be wrong.

35.5 Candidate strategy#

Set an attempt budget per shot. Generate a small number of candidates that differ clearly in composition, choose a direction, then explore local variants. Do not produce dozens of nearly identical faces at once.

Tag candidates with a failure reason. Rejected images must never enter the golden or continuity references.

35.6 Local repainting#

Local repainting suits small areas: fingers, a mole, a collar, a screen plane, background blemishes. Leave a reasonable transition at the mask edge, and check afterwards whether light, texture or identity changed.

When facial structure is wholly wrong, or the character's position or the spatial geometry is wrong, do not chase it with repeated local repairs — return to the whole frame or to the assets.

35.7 Keyframe phases#

Action shots may need a start frame, a peak frame and an end frame. First and last frames must be physically reachable and share identity and space. For key performance, choose the phase that best confirms the state, which is not necessarily the most extreme instant.

Even when video needs only a first frame, imagine the end state in shot design so the action has somewhere to go.

35.8 Props and text#

Generate a clean surface first for documents, phones and screens, then place the approved graphic asset. Shots of a character holding something may leave the text less than fully legible; the separate insert carries the read.

After compositing, check perspective, reflections, occlusion, depth of field and finger coverage, so it does not look like a flat pasted image.

35.9 Upscaling and texture#

Upscaling cannot create correct features or correct text; use it only after the structure is approved. Over-sharpening leaves material from different sources feeling mismatched. Unify skin tone, noise, sharpness and grain while preserving usable dynamic range.

Keep a master with no subtitles and no final compression.

35.10 Keyframe QC#

Check identity, wardrobe, space, props, lighting, composition, safe areas, action entry state and rights. Hero frames must pass every blocking item; background functional frames may record waivers.

Approved frames store their parent assets, the generation record, human modifications and a hash. Video generation references a specific version.

35.11 The production questions one keyframe must answer#

"It looks good" is the last requirement, not the first. A production keyframe must let downstream answer six questions: who this person is; what state they are in; where the action starts or ends; where the important objects sit; which part of the frame carries information; and what the video model may and may not change.

So keyframe approval runs in three rounds rather than everyone commenting at once by feel.

The first round is narrative and directorial: does the frame express the beat accurately, where does the eye go first, is the power relationship readable, does the shot size leave room for performance. The second is continuity: do identity, wardrobe, injuries, props, space, light and direction inherit the correct state. The third is production: do resolution, edges, trackable planes, layer requirements, safe areas, motion headroom and rights provenance permit downstream execution.

One person may run all three rounds, and the conclusions must be recorded separately. Otherwise a visual director approving the atmosphere leaves the generation station believing hands, documents and space were approved too.

35.12 More references is not better#

Once references conflict, the model averages the character. An identity pack needs an explicit hierarchy: one primary golden reference determines facial structure; forty-five degree and profile views supply geometry; expression references supply muscle state without redefining identity; wardrobe references supply only clothing; the previous shot's frame supplies only continuity state.

Every reference states its authority. identity_authoritative can determine face shape and feature proportion. expression_only cannot change age or bone structure. wardrobe_only cannot affect hairstyle. continuity_only can determine only hands, props and orientation. Without declared authority, operators push every image into the model and then cannot explain why the character increasingly looks like a blend of several versions.

When a character passes through a story state change, do not quietly update the golden reference. Wet hair, injuries, fatigue and worn-off makeup are the state layer, derived from a stable identity. Only recasting or a formal visual redesign creates a new identity master version — and that triggers a list of affected shots.

35.13 Three fidelity levels for composition exploration#

The first level can be a sketch, a collage or a low-resolution generation, verifying masses only: is the character left or right, how much of the frame is the document, does the negative space leave room for subtitles and screen graphics. The second verifies photographic relationships: camera height, sense of focal length, perspective, depth of field, eyelines and light sources. Only the third verifies identity, skin, fabric, prop edges and final texture.

Each level needs explicit approval. If the second level shows the boardroom reverse cannot hold the axis, return to layout — do not cling to a composition because the third level has already made Lin Yun's face beautiful. Teams have to resist detail sunk cost deliberately.

For tier-A hero shots, keep a composition branch board: at least one safe option, one stronger power option and one that crops well for campaigns. Approve one baseline for the episode; the other directions can become acquisition creatives without being relitigated during episode production.

35.14 First and last frames need a reachability review#

First/last frame control is often misunderstood as giving the model two attractive images. The real question is whether plausible motion exists between them. If the first frame has the right hand under the table and the last has it beside the face while the document goes from closed on the table to open in hand, then lifting, grasping, opening, turning and holding all occur within three seconds — and any model will happily "complete" that by melting.

A reachability review compares item by item: character position, body orientation, joint angles, hand occupancy, prop form, occlusion relationships, camera position, focal length, background geometry and lighting. When more than one primary action changes, add an intermediate phase or split the shot. When the two frames come from different generation batches, also overlay and compare facial outline, shoulder width, table edge, door frame and horizon; deviations that look small in stills become obvious drift once interpolated.

35.15 The authorization shot in Backlit Takeover#

The original keyframe put Lin Yun, the authorization, the chair and ten meeting attendees in one image. It was rich in information and nothing in it was clear enough: the document text was unreadable, Lin Yun's performance was too small, and background faces kept competing for attention. The team split it into three keyframes with distinct responsibilities: a medium close-up of Lin Yun as the document lands, an insert of the signature area, and a close-up of Zhou Lan losing control of her expression for the first time.

The insert does not ask the model to render text; it asks only for correct paper, fingers, table perspective and white space in the signature area. The approved acquisition document is composited in perspective as a separate graphic asset. Zhou Lan's reaction frame contains no document at all, inheriting the causality through her eyeline toward the table and her fingers stopping their tapping. After the split, each keyframe answers one story question and each enters a stable motion route more easily.

35.16 SOP for still frames#

First, compile the keyframe request. Second, explore composition cheaply. Third, choose a direction and re-anchor identity. Fourth, repair hands, wardrobe, background and props locally. Fifth, composite the text graphics. Sixth, run continuity and safe-area QC. Seventh, approve and lock the version. Eighth, upscale and unify texture. Ninth, quarantine the rejected images.

35.17 Fault tree#

Symptom: beautiful detail, uncuttable. Composition and action phase were not locked first. Return to low-fidelity layout.

Symptom: local repairs make it look less like her. The repaint area is too large or the references disagree. Redo the whole frame and re-anchor.

Symptom: document text is distorted. The model was asked to typeset. Move to graphic compositing.

Symptom: faces go plastic after upscaling. Sharpening and denoising are excessive. Reduce the processing and match the film's overall texture.

35.18 Checklist, exercises and deliverables#

Check that input assets are versioned; that composition precedes detail; that the candidate budget is explicit; that the action phase is correct; that text is composited in post; that local repairs preserve lighting; that approved frames record provenance; and that failures are quarantined.

Exercise one: write a complete keyframe request for E001_S09. Exercise two: classify one failure as composition, identity, hands, space or text. Exercise three: build a first/middle/last three-phase keyframe plan.

Deliverables for this chapter: the keyframe request, the candidate board, the local repair record, graphic composites, keyframe QC, and the approved keyframe manifest.

35.19 What a keyframe approves is a geometric contract#

A keyframe is not "this image looks good." It locks character identity, body pose, scene anchors, prop positions, light direction, composition and action phase. A video model may move within the contract's permitted range; it may not redesign the space. The approval record lists those invariants item by item.

First and last keyframes must also prove physical reachability: can the character reach that position in the target duration, does the hand have a plausible path, does prop ownership change. Two individually correct images with no possible connection between them do not constitute usable first/last frames.

35.20 Candidate diversity comes from hypotheses, not from volume#

The first round of candidates tests composition, angle and performance intensity separately, changing one primary variable at a time. Explore local detail after the direction is chosen. Twenty nearly identical candidates only increase review burden without telling the team which hypothesis is better.

The candidate board tags each image with its variable, input versions, failures and repairable items. Rejected candidates do not enter the reference pool, and their thumbnails and reasons are kept to avoid repeating the attempt. Only a passing frame that fills a clearly identified angle or state gap is considered for promotion to golden assets.

35.21 The keyframe evidence package and downstream reads#

An approved keyframe package contains the original image, a display proxy, alpha/depth or mattes where they exist, referenced assets, prompt and model version, QC, permitted motion, prohibited changes, and a content hash. Video, compositing and editing all read the same package rather than starting from a compressed image sent through a chat app.

After local repainting, the whole package produces a new version and re-runs identity, lighting and space checks. Fixing a hand while changing the face shape cannot inherit the original approval. The version actually used downstream is written into the generation ledger.

A note on sources#

Approving stills first is one of the most consistent lessons in front-line AI video production. This chapter turns it into a formal quality gate and a traceable asset rather than an image casually generated before video.