中
Chapter 87. Implementing Automated QC: Rules, Model Judgment, and Human Evidence

Part XVI — Engineering Implementation and Production Infrastructure

Chapter 87. Implementing Automated QC: Rules, Model Judgment, and Human Evidence#

In this chapter
87.1 Check timing, the rule registry, and deterministic verification87.2 Media, model, and timeline evaluation87.3 Issue clustering, human review, and threshold calibration87.4 Quality-system acceptance and evaluator red-teamingA note on sources

The value of automated QC is not a score on the finished cut. It is catching determinate errors early, narrowing where humans have to look, and routing each problem to the right workstation. The system must keep rule checks, signal analysis, model-assisted judgment and human taste distinct, so no single layer can impersonate quality as a whole.

87.1 Check timing, the rule registry, and deterministic verification#

What gets checked, and when.

Before an episode pack locks, check facts, durations and dependencies. Before an asset locks, check identity, rights and state. After candidates are generated, check media, identity, action and continuity. After the edit, check subtitles, graphics, audio and pacing. Before release, check technical specification, rights, manifest and approvals.

The earlier an error surfaces, the smaller the rework radius. A contract amount that conflicts at pack stage should block graphics production — not wait for OCR to find it in the release file.

The rule registry.

qc_rule:
  rule_id: CONT_PROP_HOLDER_001
  version: 2.1.0
  layer: continuity
  applies_to: shot_transition
  description: the end-state holder of the outgoing shot must equal the start-state holder of the incoming shot
  severity: blocker
  input_schema: shot_state_pair.v2
  evaluator: deterministic
  failure_route: continuity_supervisor
  regression_set: reg_prop_handoff_v04

Every rule has a subject, a version, its evidence and its routing. Changing severity counts as a version change, because it changes release behavior.

Deterministic checks.

Directly computable items include schema, missing dependencies, IDs, versions, state interfaces, timecodes, subtitle boundaries, exact OCR fields, resolution, frame rate, color gamut, audio channels, loudness, peaks, black frames, frozen frames, file checksums and rights expiry.

A deterministic failure must never be outweighed by a high composite score. If the amount is wrong, no quality of identity or performance makes it releasable.

87.2 Media, model, and timeline evaluation#

Media signal analysis.

Video analysis detects duplicated frames, flicker, anomalous motion, edge tearing and sudden color shifts; audio detects clipping, silence, noise-floor jumps, phase problems and dialogue intelligibility. Thresholds vary by shot type: a static reaction shot legitimately has low motion, while a long frozen frame in an action shot is probably a fault.

Signal detection reports a phenomenon; it does not infer a story cause. Low brightness in a rainy-night shot is not an exposure error, and judging it requires the scene state.

Model-assisted evaluation.

Identity, hands, action state, performance intent and narrative comprehension can all be assisted by models. Inputs must carry the contract and the visibility conditions; outputs return per-dimension results, confidence, evidence frames and failure characteristics. A model evaluator has its own version, calibration set and statement of applicability.

qc_finding:
  finding_id: find_015_07
  evaluator: identity_eval@3.2
  result: review
  confidence: 0.64
  observed: right-profile jaw width appears to have drifted
  evidence_frames: [42, 57]
  anchors_visible: [mole, hair_part, jawline]
  route: human_identity_review

Low confidence goes to a human. It does not randomly resolve into a pass or a fail.

Timeline-level checks.

Correct shots do not guarantee correct adjacencies. The system extracts each shot's first and last state and checks holder, injury, wardrobe, light, eyeline, musical bar and dialogue knowledge. After an edit version changes, recompute only the affected adjacencies and the overall duration, rather than indiscriminately re-running every expensive evaluation.

The timeline also checks audience knowledge: if a fact is not shown until 00:40, a reaction at 00:32 cannot be grounded in it.

87.3 Issue clustering, human review, and threshold calibration#

Deduplicating and aggregating issues.

One root cause can produce thirty findings. A wrong identity reference will fail several shots at once, so the system clusters by entity, rule and shared dependency into a single incident and shows the impact list. Closing the root issue triggers regression on the related shots; the operator is not asked to dismiss duplicate notifications one at a time.

Deduplication must not hide independent errors. An OCR error and a subtitle occlusion at the same timecode belong to different responsibility layers and are handled separately.

The human review interface.

By default the interface shows the current issue, the expectation, the evidence frames, the shots before and after, the relevant contracts, a candidate comparison and the downstream impact. A reviewer can pass, fail, mark a false positive, create an exception or change the routing. Keyboard shortcuts serve repetitive work, but high-risk approvals require explicit confirmation.

Human decisions record a reason code and a note, which feed evaluator calibration. The system does not assume humans are always right: significant disagreements can be arbitrated, and reviewer consistency is auditable.

Threshold calibration.

Use samples with human golden labels to compute false positives, false negatives and behavior at different base rates. Blocker rules prioritize reducing false negatives; low-risk aesthetic hints can tolerate more misses in exchange for less review fatigue. Thresholds are stratified for principal close-ups, wide shots and background performers.

An evaluator version upgrade runs in shadow first, without affecting the gate. Compare old against new, against human results and against cost, then enable it gradually.

87.4 Quality-system acceptance and evaluator red-teaming#

Fault tree.

Many issues and nobody acting on them: routing and clustering are missing. High automated scores with wrong facts: deterministic rules were overridden by a composite score. False positives suddenly rising: input distribution or thresholds have drifted. Fixing one shot breaks another: regression scope was not derived from dependencies. Humans closing findings too fast: the interface lacks evidence, or the queue is overloaded.

SOP, checklist and deliverables.

Establish the rule registry. Run checks by production stage. Do deterministic checks before model-assisted ones. Preserve evidence. Cluster by shared root cause. Route to humans. Record decisions. Regress by dependency. Calibrate thresholds. Release evaluators through shadow runs.

  • A blocker cannot be offset by a composite score.
  • Every finding has evidence and a responsibility layer.
  • Low-confidence evaluations go to a human.
  • Every evaluator has a calibration set and a scope of applicability.
  • Closing a fix requires regression evidence.

Exercise: implement ten rules and two model evaluator interfaces, delivering rule_registry.yaml, qc_runner/, findings.jsonl, review_ui_spec.md, calibration_report.csv and regression_report.md.

Red-teaming the evaluators.

Feed the identity evaluator three kinds of adversarial sample: a very similar face with the small mole moved to the other side; a correct identity under a heavy stylistic filter; and two people in frame with features swapped between them. Watch whether overall similarity fools it. Feed the narrative comprehension evaluator passages whose surface keywords are all present but whose causal order is wrong, and check whether it is only doing text matching.

Every escape enters the calibration set — but do not keep adding special cases for individual examples forever. First decide what is actually missing: the input contract, visibility, model capability, or the threshold. If an evaluator cannot handle multi-person occlusion reliably, narrow its stated scope and force human review, rather than continuing to claim "full coverage of identity consistency" in the registry. An honest capability boundary is itself quality control.

Go-live evidence includes, at minimum, golden-set performance, false positive and negative rates, applicable shot types, shadow-run results and a rollback switch. Missing any one of these, an evaluator may advise — it may not control the release gate.

A note on sources#

This chapter describes a layered checking architecture rather than a particular QC product. What transfers is running deterministic checks before model judgment, attaching evidence and a responsibility layer to every finding, and treating a declared capability boundary as part of the quality system.