Appendices
Appendix F. Fault Atlas and Repair Playbooks#
In this chapter
This appendix starts from what a viewer sees and traces back through material, state, shot design, generation, sound, editing, systems and rights to the earliest wrong layer. Preserve the evidence first, then choose the repair. Do not regenerate a whole shot the moment you see a problem.
F.1 The common fault record#
fault:
fault_id: fault_e001_042
symptom: what is observably happening
timecode: 00:58.440
expected: the correct state per the contract
evidence: [frame, waveform, state_diff, manifest]
severity: blocker
earliest_likely_layer: edit_selection
hypotheses: []
tests: []
selected_route: null
regression_scope: []
close_evidence: []
F.2 "The character looks slightly off"#
Classify first: facial proportion, stable features, age, medium style, angle and occlusion, or attribution in a multi-person shot. Compare against the identity contract and the visible anchors rather than looking only at a composite similarity score.
Root causes in order: a wrong or flipped reference; a missing angle reference; temporary state contaminating identity; too much movement or expression; drift accumulated through continuous derivation; a vendor version change. Repairs run: substitute the correct reference, layer the state, reduce motion, use keyframes, re-anchor to the golden references, repair locally, or change the shot. Regress the current shot, its neighbors, and any later asset referencing that candidate.
F.3 A mole, hair parting or piece of jewelry on the wrong side#
This is a blocker, not a small flaw, because it usually means the whole image is mirrored. Check the watch, the injured side, screen direction, the background door and any text. Repairing only the mole is prohibited.
If an asset was flipped for layout reasons, tag it layout_only so the reference selector cannot use it. Regress every
candidate derived from that asset.
F.4 Two people swapping faces, wardrobe or features#
Check the cast layout, the slots, the distance in reference collages, feature ownership and wardrobe differentiation. Reduce the number of people needing high detail simultaneously and organize prompts by slot; generate separately and composite if necessary.
Writing "don't swap faces" is usually ineffective. After repairing, test swapped positions, occlusion and reverses, to confirm attribution does not depend on a lucky background.
F.5 Age fluctuating#
Check whether fatigue, pallor, crying or rain have been written into the identity description, and check the flashback age state and any skin-smoothing style. The facial structure stays; age is expressed through controlled changes in skin detail, hair and makeup, wardrobe and posture.
If only one shot ages, recompile its state. If a whole batch changes, check the model, the adapter or the reference version.
F.6 Malformed fingers and contact penetration#
Distinguish local form from action state. Extra fingers in a local region can be repaired locally; errors in the holder, the order of contact or a duplicated prop have to go back to action design or generation.
Reduce concurrent action and produce start, contact peak and end frames. Generate insurance coverage before attempting a complex handoff. Check frame by frame; do not approve from a first-frame thumbnail.
F.7 A prop teleporting or duplicating#
Compare the previous shot's end state, the next shot's start state and the event log. It may be generation failure, the editor selecting the wrong take, state that was never committed, or a duplicate event. First check whether an approved candidate already has the correct state — if so, change take; only if not, redo.
Closure evidence includes the timeline state, the prop events and regression on the neighboring shots.
F.8 Crossed axis and wrong eyelines#
Check the floor plan, world coordinates, the axis, camera zones and the target object. A slightly late frame can move the entry point; a wrong direction cannot be fixed by cropping. Before flipping horizontally, check identity, injuries, text and background.
You can add a neutral shot, a visible re-establishing move, or redo the reverse. Write eyelines as world targets, not as "look left."
F.9 The background changing every shot#
Check whether each shot regenerates the location from text instead of using a parent space, camera zones and background plates. Lock the floor plan, landmarks, materials and light sources, and generate per-position references.
Slight texture drift can be stabilized or composited. A change in the position of doors and windows affects spatial logic and has to be redone.
F.10 Wrong light direction or time of day#
Compare story time, weather, the light map and the previous shot. Grading can unify exposure and color; it cannot invent a plausible light source. Recompile the environment state or redo the shot.
Dusk progression, blackout stages and emergency lighting all become events. Regress the view outside, rim light on characters, reflections and the sound environment.
F.11 Rainfall, wetness or wind direction jumping#
Check whether the environment contract and the wetness state accumulate with exposure. Rain can be layered in post; character wetness cannot be random. Pickup shots start from the corresponding state.
Test rain streaks, blacks and skin tones at the target bitrate. Sticker rain is usually missing contact, depth and sound.
F.12 Correct lip sync and a false performance#
Check the vocal intent, the stress on the idea, breathing and physical action, rather than continuing to adjust sync. Mouth sync is only the technical layer. Lock the physical performance first, then sync short lines; split long lines into units of thought or cover them.
If the expression gives the secret away early, adjust the performance and the entry point of the cut. Review with the sound off to verify the intent still reads.
F.13 The voice changing between lines#
Check the voice bible, the anchor, session context, tempo, register, distance and emotional-age vocabulary. Generate in batches by scene, and carry the preceding and following lines and the relationship with each segment.
Assemble long narration by thought unit, unifying room tone without using reverb to hide a drifting timbre. Regress on phone speakers and at low volume.
F.14 Fluent narration that is boring#
Check whether the thought changes, whether the narration restates the picture, whether the picture is only illustrating, and whether music is layered throughout. Build an authority ledger, run a 30-percent-deletion pressure test, and give each unit a cognitive exit.
Schedule a non-narration sound lead every 8 to 15 seconds. Check the two tracks separately by listening without picture and watching without sound.
F.15 Music that feels pasted on#
Common root causes: writing emotion without a narrative trigger; a motif with no meaning; entries placed on cut points rather than on states; wall-to-wall coverage; broken bars after a recut. Return to the music bible and the cue sheet, and use stem additions and removals with legitimate bar exits.
Music does not repair a story nobody follows. Check first whether the music-free version holds up.
F.16 Unintelligible dialogue#
Separate recording, articulation, mix, ambience, music masking, device and subtitles. Deal with articulation and performance first, then spectrum, dynamics and ducking. Do not simply raise overall loudness.
Test on a real phone, on headphones and at low volume. Verify proper nouns and numbers against the lexicon and by human review.
F.17 Subtitles covering faces, breaking badly, or containing typos#
The subtitle source is the line registry; recognition is used only for timing. Check semantic segmentation, the two-line maximum, safe areas, platform buttons and dynamic avoidance. Proper nouns, amounts and dates bind to facts.
Changing a source line invalidates every language's subtitles. Closure evidence includes device screenshots and automated boundary checks.
F.18 Garbled text on a phone or contract#
Do not retouch generated text. Use a blank trackable container plus post-produced graphics. Bind facts, fonts and locale, and run OCR, perspective, blur and device tests.
If changing one number requires redoing the whole shot, the asset layering design has failed.
F.19 "The pacing is slow"#
Do not speed up the whole episode. Ask, section by section, who the audience is watching, what they know and what they are waiting for. Delete pauses after the information has landed, duplicated reactions and repeated explanation. Keep the completion of actions, moments of judgment and the breath in relationships.
Make a shortest-comprehensible version and add back item by item. After repairing, check lip sync, music, subtitles and state.
F.20 Every shot is good and it will not cut together#
Check whether the shot requirements contain only hero shots, with no before-and-after action, reactions, inserts or spatial coverage — and check handles and state interfaces. Missing shots usually go back to shot design, not to the editor's ability.
Generate low-risk insurance: hands, props, reactions, environment or neutral space. Do not paper over a gap in causality with a random transition.
F.21 A cut point converts well and the next episode drops#
Check whether the old debt was repaid, whether the next episode's first shot continues in the same time and place, whether it opens with a recap, and whether the campaign material misled. Analyze the 30 seconds after payment, refunds and comments together.
Fix the next-open contract rather than cutting even earlier.
F.22 Heavy abandonment after an ad#
Check whether the ad interrupted an action peak, whether the restored question is concrete, whether music and loudness jump, and whether the platform's actual insertion leaves black frames. Reselect a local completion point, write a resume contract, and screen-record it in the platform sandbox.
F.23 Too many automated QC false positives#
Check the scope of applicability, visibility, base rates, thresholds, evaluator version and input distribution. Stratify by shot type, and route low confidence to humans. Do not switch rules off to make the red lights go away.
Recalibrate against a golden set, shadow-run, then stage the rollout. Record the cost of false positives against the consequence of misses.
F.24 Automated QC missing a blocker#
Determine whether a deterministic rule was overridden by a composite score, whether the evaluator treated low confidence as a pass, and whether a key fact lacked an authoritative input. Blockers gate independently and are never averaged.
Escaped samples enter the regression set. Check historical assets of the same class, and assess whether publishing should pause.
F.25 The same task billed twice#
Check for duplicate events, worker crashes, call intent, idempotency keys and vendor jobs. Stop the fan-out first, query the remote side, and preserve the evidence. The fix persists the intent before calling and adds a recovery algorithm.
A unique row in your database does not mean a unique external call. Regression covers lost responses and process restarts.
F.26 An old asset contaminating a new batch#
Check the cache key, lifecycle, the dependency invalidation consumer, and the version at submission. A late-arriving old result must be quarantined. Add a version checksum to the cache key, and have asset promotion actively invalidate context.
Scan every candidate and published derivative — not only the one shot where it was noticed.
F.27 The human approval queue exploding#
Order by irreversibility, critical path, budget and deadline. Aggregate common root causes. Sample low-risk rules. Apply upstream backpressure. Do not automatically loosen the gates or let more agents keep generating.
Compute the human load, and improve the differences and evidence in approval packages so people do not reopen the same material repeatedly.
F.28 The client says "make it feel more premium"#
Convert the feedback into observable objectives: material, lighting ratio, camera movement, typography, pacing or brand presence. Ask for a timecode, a reference and a commercial purpose. If the scope grows, use a change request.
Do not substitute a whole-episode LUT, more slow motion and a more expensive model for defining the problem.
F.29 Missing rights evidence#
Quarantine the asset immediately and stop new derivations. Check provenance, licensed use, territory, term and vendor terms. "Found it online," "it's AI generated" and a payment screenshot are not a rights chain.
If it cannot be cleared, replace the asset and regress every release package along the dependency graph. Keep the disposition record.
F.30 Publishing the wrong version#
Stop traffic, preserve the platform receipt and the bad file, confirm the scope, roll back to the signed release candidate, and notify the responsible parties. The root cause is usually publishing that bypassed the immutable manifest, checksums or the two-person gate.
Do not delete evidence first. After recovery, verify at low volume before scaling back up.
F.31 The expected-total-cost selector#
expected total cost = direct repair cost
+ probability of failure × cost of repairing again
+ regression cost
+ schedule impact
+ narrative and rights risk
A local fix looks cheap, but if it leaves a hidden state error that contaminates several episodes, its expected cost is higher. Regeneration is neither the default nor the last resort; it is one route among several.
F.32 Closure conditions#
Closing a problem requires the fix version, evidence, regression results, an approver and a time. A systemic problem also needs a control improvement and a regression test. A problem that cannot be reproduced but carries high risk goes into monitoring — not "not observed, closed."
F.33 Fault drills#
Every month, choose one fault each from identity, state, sound, systems, rights and release, and inject it in an isolated environment. Record detection, containment, recovery and improvement. The drill passes not when things end up fine, but when there was no overreach of authority, no duplicate spend, no lost evidence and no inexplicable state.
F.34 Final checklist#
- Symptoms are written as observable facts with a timecode.
- Expectations come from a contract, not from someone's impression.
- Root causes are traced to the earliest responsible layer.
- Repair routes are compared by expected total cost.
- Regression scope comes from the dependency graph.
- Blockers are not offset by composite scores or client requests.
- Closure has a version, evidence and an approval.
- Systemic incidents are written back into tests and runbooks.
Exercise: choose ten fault classes from this appendix and produce a reproducible sample, a diagnosis tree, three routes
and closure evidence for each. Deliver fault_cases/, diagnosis_trees/, repair_decisions.csv,
regression_reports/ and fault_drill_postmortem.md.