中
Appendix L. Twelve One-Page Duty Runbooks

Appendices

Appendix L. Twelve One-Page Duty Runbooks#

In this chapter
L.1 Sudden identity drift on a leadL.2 Wrong prop holder or stateL.3 Boardroom space and eyelines breaking downL.4 Long narration losing its voice or structureL.5 Music needs recutting after picture lockL.6 Prompt, model or vendor migrationL.7 Budget overrun or a spike in cost per approved secondL.8 Likeness, voice or music rights revokedL.9 A wrong version is already publishedL.10 The human approval queue is congestedL.11 Campaign metrics fall sharplyL.12 State database or workflow disaster recoveryL.13 General runbook disciplineL.14 Drill record

Runbooks exist to handle known incident types safely under time pressure. The person on duty protects facts, rights, budget and release state first, then restores production. If the situation exceeds a runbook's scope, stop in a safe state and escalate — do not widen your own authority.

L.1 Sudden identity drift on a lead#

Trigger: the lead's stable anchors change, the automated identity score falls below the threshold, or a human judges the character unrecognizable.

Immediate actions: pause new generation on that identity version; preserve candidates, inputs, prompt manifests and evidence frames; query every task in the same batch and using the same references; do not overwrite or delete the failed assets.

Diagnosis order: confirm whether the image is mirrored; check reference lifecycle and angle; check whether fatigue, age or rain contaminated identity; check the range of motion and expression; check for continuous derivation from generated end frames; confirm model and adapter versions.

Recovery: recompile context from the golden references; keep the state and reduce identity conflicts; validate statically first, then a small batch of video; restore only the affected tasks after it passes.

Closure evidence: identity anchors, human review, the regression set, zero remaining affected objects, and the reference selector fixed. If a model upgrade is involved, keep the old route as a rollback.

L.2 Wrong prop holder or state#

Trigger: a prop teleports, duplicates, opens or closes inconsistently, appears in the wrong hand, or its ownership is misused.

Immediate actions: freeze edit promotion for the affected shots; read the previous shot's end state, the next shot's start state and the events; determine whether the error came from generation, take selection, an uncommitted state or a duplicate event.

Diagnosis: if an approved take already has the correct state, route it to editing for replacement. If every candidate is wrong, go back to the action contract. If the database state is wrong, create a correcting event rather than overwriting history. If the event was duplicated, check idempotency.

Recovery: use the correct take, add insurance coverage, or redo start–peak–end; commit the prop event; recompute dependencies.

Closure evidence: zero state difference between adjacent shots, frame-by-frame contact passing, a correct prop timeline, and completed music and duration regression.

L.3 Boardroom space and eyelines breaking down#

Trigger: doors and windows swap position, a character looks the wrong way, reverses face the same direction, or the axis is crossed with no explanation.

Immediate actions: stop using a horizontal flip as a quick fix; open the location floor plan, seating, world coordinates and the axis; mark the bad shots and the establishing shots either side.

Diagnosis: check whether the shot came from the wrong location version; whether the camera sits in an approved zone; whether the eyeline was written as a screen direction instead of a target; whether character movement updated the coordinates; and whether background plates were mixed.

Recovery: substitute a correct candidate first; then add a neutral insert or a visible re-establishing move; redo the reverse if necessary. Flipping is used only when identity, text, injuries and background all permit it.

Closure evidence: the muted spatial test, target eyelines, the axis diagram, background landmarks and regression on adjacent shots.

L.4 Long narration losing its voice or structure#

Trigger: a section sounds like a different person, the emotion only ever escalates, joins are abrupt, or listeners cannot restate how the judgment changed.

Immediate actions: preserve every segment and its context; do not regenerate the whole passage; label voice, register, tempo, distance, breath and cognitive exit by unit of thought.

Diagnosis: check whether it was segmented mechanically by word count; whether emotional descriptions changed the identity's apparent age; whether every segment used the same voice anchor; whether the narration restates the picture; and whether music only ever adds.

Recovery: redo only the failed units, carrying the thought before and after; unify room tone and space without using reverb to hide timbre drift; make a version with 30 percent removed; redistribute authority between picture, narration and music.

Closure evidence: four passes — narration alone, with music, with picture, and at low volume on a phone — with the target listener able to restate the arc of understanding.

L.5 Music needs recutting after picture lock#

Trigger: a shot lengthens or shortens, a cue's bars break, music masks dialogue, or the theme's meaning no longer fits.

Immediate actions: record the picture lock change and the affected cues; do not stretch the whole piece; read the theme, tempo, key, bar and stem state.

Diagnosis: determine whether the change is local duration, scene structure or emotional meaning; find legitimate loops, sustained notes and exit bars; check dialogue, Foley and the next cue.

Recovery: prefer moving the bar exit, adding or removing texture bars, extending a sustain, or rearranging stems. Rewrite the cue when the meaning has changed. Any replacement music clears rights first.

Closure evidence: a new cue sheet version, intelligible dialogue, thematic continuity, phone-speaker playback, whole-episode and continuous-playback regression, and an updated rights record.

L.6 Prompt, model or vendor migration#

Trigger: a model upgrade, a service shutting down, a price or terms change, or an adapter rewrite.

Immediate actions: freeze the current capability cards and the old components; list affected tasks and unreleased assets; do not simply swap an API name in production.

Diagnosis: run the project's private evaluation set and compare identity, state, action, performance, cuttability, cost per approved second, latency and rights; check whether unsupported capabilities are being silently dropped.

Recovery: update the adapter and run prompt regression; shadow first, then 10 percent staged; check critical shots fully by hand; keep the old route and its rollback conditions. If the terms are incompatible, do not migrate even if the images are better.

Closure evidence: capability cards, blind evaluation, the staged rollout report, rights review, a rollback drill and lineage across old and new assets.

L.7 Budget overrun or a spike in cost per approved second#

Trigger: a batch forecast exceeds the cap, failed attempts cluster, human repair surges, or a vendor raises prices.

Immediate actions: pause low-priority calls and calls with no hypothesis while preserving the critical path; reconcile reservations, actuals, duplicate charges, human minutes and approved seconds.

Diagnosis: classify failures by shot type; find the shared asset, model, action or approval bottleneck; distinguish price increases, falling pass rates, propagating rework and excess work in progress.

Recovery: cut shots with no function, move to lower-risk coverage, fix the common root cause, reduce concurrency, switch to a compatible route, or request new budget. Never silently lower resolution, identity thresholds or rights standards.

Closure evidence: a new forecast inside the cap, approved seconds recovered, stop rules enforced, and changes authorized by production.

L.8 Likeness, voice or music rights revoked#

Trigger: a licence expires, a revocation notice arrives, terms change, or provenance cannot be proven.

Immediate actions: quarantine the affected assets and stop new tasks and publishing; preserve the notice and the existing evidence; revoke access and generation capability without deleting historical records.

Diagnosis: query every candidate, cut, language, creative, platform and backup by rights ID; separate unreleased, released, and evidence the contract permits you to keep.

Recovery: replace the voice, music or visual asset; redo the derivatives; take down or update published content per legal advice; notify the client and the platforms. Process vendor deletion requests.

Closure evidence: zero results for new use, replacement versions passing, platform status confirmed, the basis for deletion or retention archived, and an updated rights snapshot.

L.9 A wrong version is already published#

Trigger: old subtitles, a wrong graphic, the wrong episode, unapproved cover art, or a file that does not match the manifest has gone live.

Immediate actions: stop or limit traffic; preserve the platform receipt, the bad file, the remote ID and the time of discovery; notify the release owner — do not delete evidence first.

Diagnosis: confirm the platform, territory, language, user scope and caching; compare the remote hash with the signed manifest; find which layer it escaped — selection, upload, verification or the gate.

Recovery: roll back to a known approved version; verify playback, sound, subtitles and events at low volume; then scale up. If no rollback is available, stop distribution first.

Closure evidence: the correct remote hash, the impact list, notifications, the root cause, a repaired two-person gate and a regression test.

L.10 The human approval queue is congested#

Trigger: waiting exceeds the SLO, release-critical items are buried under ordinary candidates, or upstream keeps producing until versions go stale.

Immediate actions: reorder by irreversibility, critical path, budget, deadline and safety; apply upstream backpressure; aggregate common root causes and duplicate candidates.

Diagnosis: check whether approval packages lack differences and evidence; whether rules can safely pre-filter; whether one person carries too many approvals; and whether batch parallelism exceeds capacity.

Recovery: review key facts and release items individually; sample low-risk environment shots against calibrated rules; add trained reviewers or reduce parallelism; improve the approval interface. Automatically lowering thresholds is prohibited.

Closure evidence: waiting restored, no escapes on critical items, a sustainable human load, and updated backpressure rules and capacity model.

L.11 Campaign metrics fall sharply#

Trigger: a significant change in clicks, first-episode completion, continuation, payment or ad value.

Immediate actions: check event completeness, definitions, versions, platform, bidding, audience and landing point before changing anything in the story; freeze the interpretation of this round's experiment.

Diagnosis: locate the funnel position. Low clicks: check the creative. High clicks with immediate exit: check the landing. Low completion: check comprehension and pacing. Low next-episode entry: check the cut point's payoff. Low payment: check value and friction. High refunds: check for misleadingness.

Recovery: fix deterministic data or landing problems. Turn creative problems into issues with timecodes and a minimum test. Slow variables enter a batch change request rather than daily edits to the series bible.

Closure evidence: data quality passing, a complete experiment registry, the fix version and its control, no degradation in guardrails, and the recorded scope of applicability.

L.12 State database or workflow disaster recovery#

Trigger: the database is unavailable, materialized views are corrupt, event consumption has stalled, or task state and budget disagree.

Immediate actions: stop new writes and irreversible publishing; preserve logs, event offsets, object storage and the most recent snapshot; confirm whether read-only operation is safe.

Diagnosis: distinguish authoritative events, snapshots, materialized views and media objects; establish the last consistent point, the RPO, duplicate messages and unsettled calls. Do not guess state from a cache.

Recovery: restore the snapshot in an isolated environment and replay events; rebuild the views; reconcile task, budget, asset, approval and release invariants; sample media checksums; restore workers at low volume.

Closure evidence: RPO and RTO, a recovery comparison, safe duplicate handling, disposition of pending calls, a replay by someone who did not build the system, runbook revisions and the next drill date.

L.13 General runbook discipline#

  1. Describe observable facts first; do not guess at root causes.
  2. Stop the spread first; do not optimize for speed of recovery.
  3. Preserve inputs, logs, versions, remote receipts and the timeline.
  4. Confirm who holds the decision authority; the person on duty does not exceed it.
  5. Fix the earliest wrong layer, not only the final file.
  6. Derive the regression scope from the dependency graph.
  7. Verify at low volume or in isolation before restoring.
  8. Closure requires evidence, an owner and a follow-up test.

L.14 Drill record#

drill:
  drill_id:
  runbook_version:
  scenario:
  injected_fault:
  participants: []
  detected_in_s:
  contained_in_s:
  recovered_in_s:
  evidence_preserved: true
  unauthorized_actions: 0
  duplicate_cost: 0
  data_loss:
  missed_steps: []
  runbook_changes: []
  regression_tests: []
  next_drill_date:

Drill at least two categories every month, and cover rights, release and disaster recovery each quarter. A drill is not a performance of how fast the team is. It finds where controls, evidence and documentation still depend on an individual.