Part V — Continuity Engineering: Character, Space, Action
Chapter 28. Continuity QC and Drift Scoring#
In this chapter
28.1 Continuity is not one total score#
"Continuity: 80" tells a team nothing about what to fix. Score at least seven dimensions: identity, styling, space, props, action, light and time, and emotion and knowledge. A blocker in any critical dimension cannot be offset by high scores elsewhere.
Hero close-ups, ordinary coverage and background shots use different tolerances. A background face being indeterminate is not a failure; illegible key evidence always blocks.
28.2 Seven-dimension scoring#
continuity_score:
shot: E001_S09
shot_class: hero_closeup
identity: 5
wardrobe: 5
spatial: 4
prop: 5
action: 4
lighting_time: 5
emotion_knowledge: 5
blockers: []
decision: pass
A 1 is a visible break. A 3 is visible but potentially acceptable through editing or local repair. A 5 matches the approved state. Thresholds are defined by shot function; do not put faith in the average.
28.3 Reference truth#
QC has to know what it is comparing against: golden assets, the last approved shot state, the location bible, the timeline, the action handoff and the episode pack. When the references themselves conflict, resolve the source of fact first — a reviewer must not choose by impression.
Every report records the reference versions it used, so that after an asset update you can tell whether an older report still holds.
28.4 What automation is good at#
Automation finds missing IDs, conflicting state fields, swapped hands, changed wardrobe versions, time running backwards, subtitle overlap, missing files, black frames, and some facial similarity and color anomalies.
Visual similarity is good for ranking suspected drift and unsuited to automatically approving profiles, crying, occlusion and low light. A model may also rate a templated face highly similar while missing a failure of character differentiation.
28.5 What humans must judge#
Whether the person is still her, whether the performance is continuous, whether the space feels natural, whether a reaction fits that character's interest, whether a small jump will be noticed at normal speed, and whether a repair is more conspicuous than the error — all need a person.
Reviewers watch three ways: frame by frame for factual errors; at normal speed for perception; and on a phone with the finished cut for importance. A frame-level flaw invisible at normal speed on a phone, which damages no key asset, can be waived on the record.
28.6 Drift trends, not single-shot errors#
Ten consecutive shots each changing slightly may each pass while the character has been replaced across the sequence. Build sequence review: arrange one character's keyframes in time order as a contact sheet and compare face width, eye shape, hair length, wardrobe and color temperature.
Set identity checkpoints every few shots, and force a check after a change of angle, location or costume, and after occlusion clears. A drift trend triggers a hard re-anchor.
28.7 The difference report#
The system compares expected state_in with what is actually observed in the finished shot:
state_diff:
shot: E001_S10
expected:
right_hand: holding_PROP_AUTHORIZATION
left_hand: holding_PROP_OLD_PEN
observed:
right_hand: holding_PROP_OLD_PEN
left_hand: holding_PROP_AUTHORIZATION
category: prop_and_hand_swap
severity: P1
visible_at_normal_speed: true
Observed values can be pre-filled by machine and confirmed by a human.
28.8 Severity#
P0 is a rights, safety or project-wide factual error. P1 blocks the current shot or gate — a changed face, wrong evidence, reversed space. P2 must be fixed before scaling up: mild style drift, a visible background change. P3 does not affect comprehension.
Severity also weighs shot function, visibility, how far it propagates, and repair cost. A wrong amount on a key prop is P1 even if it occupies a single frame.
28.9 Rework routing#
A wrong identity baseline returns to assets. A wrong entry state returns to the continuity ledger. A camera or axis error returns to shot design. Wrong structure in a still returns to the keyframe. Drift during motion returns to video. Text and edges return to compositing. Minor rhythm covering returns to the edit.
route:
symptom: E006_S19 Lin Wei jumps from stage right to stage left after the blackout
root_cause: motion clip ignored locked spatial state
return_to: motion_generation
preserve: approved_keyframe_and_location_state
fix: regenerate from the approved keyframe, restricting character displacement
verify: compare against the end frame of E006_S18 and the establishing shot in S20
Returning to the wrong layer wastes money. Grading cannot fix a light direction, and editing cannot conceal a core character's face changing for long.
28.10 Waivers#
Choosing not to fix something must also be an explicit decision. A waiver records the problem, its visibility, the reason, the risk and the approver. Reasonable waivers commonly cover slight background extra variation, a hand obscured by motion blur, and indeterminate faces in wides.
When the same class of P2 is waived repeatedly, upstream assets or the prompt system are defective and it should be escalated to a systemic issue.
28.11 The order of a continuity review#
Review assets and still keyframes first, then individual motion clips, then adjacent shots, then the whole scene and episode. Watching only the whole episode misses detail; reviewing only frame by frame loses the sense of what matters perceptually.
A recommended set of passes: identity contact sheet; wardrobe and props; space and eyelines; action handoffs; light and time; emotion and knowledge; the final phone viewing. Focus on one class at a time to reduce cognitive load.
28.12 SOP for continuity QC#
First, load the approved reference versions. Second, run automated structural and state checks. Third, generate contact sheets and difference candidates. Fourth, have humans confirm frame by frame. Fifth, watch at normal speed and on a phone. Sixth, score the seven dimensions and identify blockers. Seventh, trace root causes and route. Eighth, after repair, re-review only the affected range plus necessary regression. Ninth, record waivers. Tenth, write recurring problems back into assets, schemas or prompts.
28.13 Fault tree#
Symptom: automated scores are high and the audience still sees a different person. The tool compares only local regions or matching angles. Add sequence contact sheets and human identity judgment.
Symptom: QC produces so many alerts the team ignores them. There is no shot classification and no severity. Set thresholds by function and merge duplicate root causes.
Symptom: the same error is reworked repeatedly. Only the finished shot was fixed, with nothing written back upstream. Count problem types and change the assets or the prompt system.
Symptom: repairing one shot breaks the adjacent shot. There is no dependency regression. Re-review the handoffs either side, the sound and the state.
Symptom: every small flaw blocks. Frame-level perfectionism is ignoring phone perception and commercial value. Use explicit waivers rather than passing things silently.
28.14 Checklist, exercises and deliverables#
Check that seven dimensions are scored; that reference versions are explicit; that automated and human responsibilities are separated; that sequence trends are checked; that differences carry both expected and observed values; that issues are routed to root causes; that waivers are recorded; and that recurring problems feed back upstream.
Exercise one: build an identity contact sheet for ten shots. Exercise two: grade ten problems from P0 to P3. Exercise three: choose the rework station for a changed face, a hand swap, an axis crossing and a lighting error. Exercise four: design the minimum regression scope after one repair.
Deliverables for this chapter: the seven-dimension continuity score, contact sheets, state diffs, the issue queue, rework routing, the waiver log, the regression report, and systemic problem statistics.
28.15 Full inspection, sampling and risk weighting#
Continuity cannot simply mandate frame-by-frame inspection of every shot. Hero close-ups, evidence text, action handoffs, the first appearance of a new asset and paywall cut points require full inspection. Stable backgrounds, low-information wides and validated templates can be sampled. Sampling is not lowering the standard; it puts limited review time where errors are most visible and propagate furthest.
A risk score can combine shot function, asset novelty, headcount, interaction complexity, text importance, maturity of the generation route, and downstream reuse count. High-risk shots automatically enter all four passes: still, motion, adjacent and phone. Medium risk gets at least motion and adjacency. If sampling in a low-risk batch finds one instance of a class of error, widen the sample or escalate to full inspection.
Sampling units must cover batches, models, characters, angles and vendors — never just the most attractive finished shots. Random samples estimate the population of problems; targeted samples hunt known high risks; report them separately. A team that always inspects hand-picked shots before delivery can only demonstrate that its best work is fine, not that production is stable.
28.16 From an issue list to an issue graph#
Ten shots with a face too wide are probably not ten independent problems. They are more likely one wrong identity reference used ten times. An issue graph connects visible symptoms to shots, states, assets, model routes and approval decisions, so the shared root cause is fixed first. Once fixed, the system lists every descendant needing invalidation and regression.

Figure 28-1 Face drift, swapped hands, wrong eyelines, a moved door and a lighting jump look unrelated, and may be produced by a small number of asset or process root causes. QC must establish the scope of impact first, then rework, regress, and close the issue with reviewable evidence.
The five symptom classes on the left correspond to different visual channels, and the issue graph in the middle compresses them into actionable root causes. The closure evidence on the right is not a decorative tick; it should correspond to before/after comparisons, state differences, regression results and approval records. Missing any of those means the problem is temporarily invisible rather than demonstrably solved.
The graph also identifies cascade failures. A wrong location version moves the door in the background, which forces eyeline direction to change, which leads the edit to flip footage horizontally, which finally reverses the mole and the text. Fixing only the mole preserves the whole error chain. A postmortem looks for the earliest departure from authoritative fact, not the most conspicuous symptom.
When closing an issue, record the verification evidence: the new asset version, the replacement shot, the comparison clip, the regression scope and the approver. "Fixed" is not a status button; it is a set of proofs that can be re-checked. An issue with no evidence can only be marked "claimed fixed" and cannot enter the release gate.
28.17 Set acceptance thresholds by shot category#
One set of numeric thresholds does not suit every shot. A frontal close-up of the lead must have clear identity, stable lip sync and correct permanent markers. A fast action wide may leave the face indeterminate while silhouette, wardrobe and direction stay correct. An evidence insert need not show a face at all and must have accurate text held long enough. A dream sequence may break realistic color and still has to obey its own rules.
acceptance_profile:
hero_closeup:
blockers: [identity, age_read, lip_sync, state_visible]
identity_min: 4.5
frame_sampling: every_6_frames_plus_expression_peaks
action_wide:
blockers: [body_count, motion_direction, prop_ownership, spatial_state]
face_status: may_be_not_judgeable
evidence_insert:
blockers: [text_accuracy, orientation, readable_duration]
readable_duration_min_frames: 18
Thresholds are approved by the director, the continuity owner and the client before production starts. They cannot be loosened after seeing how the material turned out. Where a genuine relaxation is needed, use a waiver stating the commercial reason and the risk to the audience.
28.18 How the continuity review room runs#
An efficient review does not have everyone commenting on everything at once. The first pass runs at normal speed and records only where attention breaks, without pausing. The second pass runs by channel: identity, state, space, action, environment and sound. The third pass handles only P0 and P1 root causes and routing. Aesthetic preference, continuity fact and client changes go into different queues.
The review interface shows the current shot alongside the adjacent shots, the approved references, the expected state and the observed difference. Comments must land on a timecode and a category and state the visible fact — "00:37:12 the document jumps from the right hand to the left," not "this looks odd." The meeting assigns the responsible station and the verification standard rather than blind-testing prompts on the spot.
At the end of each day, produce a continuity health report: new issues, closures, reopen rate, distribution by root cause, average rework radius, accumulated waivers and drift trends. A falling issue count with a rising reopen rate means repair quality is poor. A large volume of waived P2s means quality debt is piling up. The goal of continuity QC is not a busier review room — it is the same errors appearing less and less often.
A note on sources#
Human review cannot be removed from AI video production, and it can be structured. Turning "this looks odd" into references, differences, severity and routing is what lets review drive systemic improvement instead of endless regeneration.