中
Chapter 59. Automated Evaluation and Human Judgment

Part XI — The Agentic Production System

Chapter 59. Automated Evaluation and Human Judgment#

In this chapter
59.1 Evaluation layers, test suites, and the evidence interface59.2 Human authority, evaluator drift, and the operating flow59.3 A worked checklist, thresholds, and golden evaluation sets59.4 Responsibility boundaries, failure cases, and cost calibration59.5 Evaluator governance, release decisions, and red-teamingA note on sources

59.1 Evaluation layers, test suites, and the evidence interface#

Automation checks determinate facts first.

Automated checks are good at schema, IDs, missing files, references, overlapping time, black frames, silence, subtitle overflow, loudness, peaks, resolution, hashes and some similarity measures. They should not pretend to know whether a crying scene is moving.

Three layers of evaluation.

Deterministic rules pass or fail directly. Statistical models supply anomaly candidates. Humans judge aesthetics, performance, humor, emotion, ethics and commercial trade-offs. The interface must show which layer a conclusion came from.

The automated test suite.

Project validation: schema, stable IDs, asset references, legal states. Picture: black frames, freezes, dimensions, subtitle safe areas, facial similarity candidates. Sound: missing tracks, loudness, peaks, sync. Business: whether the cut point and next-episode repayment exist. Rights: whether key assets carry records.

Confidence and evidence.

Anomaly models output a confidence, the reference version and visual evidence. They never delete material automatically. Low confidence enters sampling; high-confidence blocking candidates still need a rule or a person to confirm.

The review interface.

Show the finished picture, timecode, shot and beat, expected state, reference assets, issues, the repair station and the shots either side simultaneously. A reviewer should not have to hunt for context across folders.

59.2 Human authority, evaluator drift, and the operating flow#

Human authority.

Humans choose the story promise, judge whether a character is compelling, assess performance, decide whether a cut point is honest, judge brand fit, interpret law, and stop a project. The system supplies comparisons, cost and evidence; a single model score cannot replace responsibility.

Evaluator drift.

Updating an evaluation model also requires regression. Keep human-labelled pass and fail sets and monitor false positives and false negatives. An evaluator must never change a release threshold without notice.

SOP.

First, list the rules that can be determinate. Second, treat soft problems as anomaly prompts rather than verdicts. Third, store references and evidence. Fourth, design a review interface carrying context. Fifth, have humans decide and write labels back. Sixth, evaluate the evaluators using those labels. Seventh, run regression on version changes.

Fault tree.

Symptom: automated scores are high and the finished piece is poor. Aesthetic and commercial judgment was disguised as a determinate metric. Restore human approval.

Symptom: reviewers ignore the warnings. Too many false positives with no evidence. Adjust thresholds and rank by risk.

Symptom: the same material passes today and fails tomorrow. The evaluator version drifted. Lock versions and use a regression set.

Checklist, exercises and deliverables.

Check that rules and models are distinguished; that anomalies carry evidence; that human authority is explicit; that evaluators are versioned; that false positives and negatives are measured; and that review decisions are written back.

Exercise one: sort twenty QC items into rules, anomaly models and human judgment. Exercise two: design a single-screen review layout. Exercise three: build an evaluator regression set.

Deliverables for this chapter: the automated test suite, the evaluation registry, the review UI spec, the human decision log, and the evaluator regression report.

59.3 A worked checklist, thresholds, and golden evaluation sets#

A worked automated checklist.

Before E001 enters the release gate, deterministic checks verify that all nineteen shot IDs exist in the storyboard; that the timeline has no overlaps or holes; that every GEN_* source file exists with a matching hash; that subtitles stay inside the safe area; that no audio track is missing; that P0/P1 counts are zero; that the authorization graphic's fields match the evidence ledger; and that music, fonts and audio all carry rights IDs.

Anomaly detection then flags three things: facial similarity for Lin Wei in S11 below the baseline for that angle; an unusual average brightness difference between S18 and S17; a high reading speed on the subtitle in S14. None of those fails automatically. The reviewer checks the surrounding shots and confirms that S11 is an acceptable change of expression, S18 is the story event of a screen lighting up, and S14 genuinely needs splitting across two screens.

False positives, false negatives, and thresholds.

Evaluators track precision, recall and misses by severity. A continuity tool that flags half the shots to avoid missing a face change produces alert fatigue. One that flags only the most obvious errors misses drift accumulating shot by shot.

Different shot classes use different thresholds. Hero facial close-ups are strict; background wides check silhouette only; proof fields use deterministic pixel and text verification rather than the same threshold as ordinary background text.

Golden evaluation sets.

Build a human-labelled golden_eval_set containing clear passes, clear failures, boundary cases and waived cases. Each stores references, the problem category, severity and the reason for the verdict. Any evaluator upgrade re-runs it and compares new false positives, false negatives and handling time.

The set includes deliberately manufactured errors: a mirrored mole, a wound on the wrong hand, time running backwards, missing music licensing, a subtitle covering a document, no repayment after payment. Using only naturally occurring errors misses the low-frequency, high-risk cases.

A worked review interface.

The reviewer sees the current frame, two shots either side, the reference character, the expected and observed states, the model's confidence, and a toggle between normal speed and frame stepping. Clicking a P1 for a prop changing hands makes the interface suggest returning to continuity or keyframes and display the affected downstream.

A human may confirm, downgrade, waive or reassign the root cause — and must write one sentence of justification. That justification enters the training and evaluation data, so the next similar problem can be ranked more accurately.

Active learning, not automatic expansion of authority.

The system prioritizes low-confidence, high-impact and novel problems for human review. With enough labels it can improve anomaly ranking, and it must not acquire release approval authority merely because accuracy rose. A change of authority requires a new governance decision and regression evidence.

59.4 Responsibility boundaries, failure cases, and cost calibration#

The human-machine responsibility table.

Machines confirm files, formats, references and measurements. Machines flag similarity, drift and anomalies. Humans judge appeal, comedy, ethics, performance, honesty at cut points and commercial stop decisions. Legal judgment cannot be replaced by a model summary; the model can only organize the questions and sources awaiting verification.

A failure case: metrics hijacking the work.

If a team requires emotional intensity of at least four in every shot and a reversal every fifteen seconds, models will manufacture sustained exaggeration and meaningless turns. Metrics are diagnostics, not creative targets. A human can approve a low-intensity long pause when it serves the scene; the system records the reason for the waiver rather than forcing it back to an averaged template.

Sampling strategy.

Stable low-risk batches can reduce item-by-item human review, while first shots, hero close-ups, cut points, rights, critical text and any new model route are always reviewed in full. When sampling finds a serious problem, widen to the whole batch and trace the shared input. Sampling ratios follow historical defect rates and impact, not a desire to save labor alone.

Evaluation objects bind to gates.

The same metric means different things at different stages. Keyframe identity similarity stops errors entering video. The motion stage checks drift at first, middle and last frames. The rough cut asks whether the audience perceives adjacent shots as the same person. One global "character consistency: 87" cannot replace three specific gates.

Evaluation records carry the artifact type, version, reference, threshold source and permitted actions. Failing a keyframe threshold means redoing the still. Finding the same class of error only at the rough cut means tracing the existing scope, not merely tagging the current timecode.

Calibration, base rates, and cost.

Before judging an evaluator, know the defect base rate. If ten of a thousand shots have identity errors and the evaluator raises fifty warnings while catching eight, the recall looks strong and it creates forty-two human re-reviews. Compute the cost of misses and the cost of false alarms separately by severity, and choose thresholds accordingly.

P0 and P1 tolerate more false positives to reduce misses. P3 polish issues need higher precision so reviewers are not buried. Record the change in human minutes after each threshold adjustment rather than looking only at offline accuracy.

59.5 Evaluator governance, release decisions, and red-teaming#

Several evaluators is not majority voting.

Visual similarity, rule validation, a vision-language model and a human may reach different conclusions. Keep each one's evidence rather than averaging scores. When a rule finds a wrong amount on the authorization, it blocks even if two aesthetic models scored it highly. When a model flags face drift and a human confirms it is a plausible expression, record it as a boundary sample.

High-impact conflicts go to human review; low-impact conflicts follow a preset priority. Persistent disagreement between evaluators means checking references, shot class and thresholds — not continually adding new judge models.

Release decisions are not made by a composite score.

A composite score is useful for ranking review priority and unsuitable for release. A work averaging 95 across story, picture and sound still cannot ship with a missing music licence, and ten minor P3s must not offset one wrong amount.

The release gate uses blocking rules, required evidence and a human signature. The evaluation system prepares complete evidence, ranks anomalies clearly and surfaces similar historical problems; final responsibility remains with the release owner.

An evaluator registry with declared applicability.

Each evaluator registers its input type, output definition, training or rule source, applicable shot classes, known blind spots, threshold version, calibration set and owner. An identity similarity model that performs well on frontal close-ups must not be applied by default to profiles, occlusion and wides. It must report "indeterminate" rather than producing a falsely precise score outside its applicability.

Evaluator upgrades follow regression and staged rollout, exactly like prompts. A new version that reduces false positives while missing P1s cannot ship on average accuracy alone. Historical RCs retain the evaluator version in force at the time; audits do not rewrite old decisions with a new model.

Consistency and arbitration among human reviewers.

Have two reviewers independently judge a small sample, compute disagreement, and discuss where the rule boundaries lie. The goal is not identical aesthetics; it is calibrating blocking facts, severity and evidence requirements. Whether two faces are the same person, whether text is correct, and whether a licence exists should show high agreement; the appeal of a performance can remain the director's judgment.

High-impact disagreements go to arbitration, where the arbiter sees the original evidence without first seeing who argued which side, reducing authority bias. The final decision, its reasoning and the boundary sample enter the golden evaluation set for later joint calibration.

The evaluation system needs red-teaming too.

Deliberately construct samples designed to fool the evaluators: wrong facial structure with similar coloring; an authorization amount wrong by one digit; missing music rights with a perfect image; subtitles partly covered by platform UI; a character knowing something early while the dialogue sounds natural. Test whether a high composite score can conceal a blocking error.

Also test tolerance for deliberate creative exceptions: a long intentional pause, subjective defocus, a dream sequence crossing the axis, a silent cut point. The evaluator should flag the deviation and request an explanation rather than automatically reverting to the template. Good evaluation catches disguised errors while permitting approved artistic choices.

A note on sources#

Automated evaluation earns its place by preparing evidence and ranking risk, not by replacing judgment. The layer a conclusion comes from — rule, model, or person — has to remain visible at the moment someone signs.