Part VII — Generating Images, Video, Performance, and Composites
Chapter 34. Evaluating Model Capability and Choosing for Production#
In this chapter
34.1 Do not choose a production model from a demo reel#
Vendor demos show best-case samples. A project needs stable output on ordinary shots. Selection has to answer: how many of this show's shots are singles, two-handers, hands, text, low light, lip sync and motion — and for each class, what is a route's pass rate, average attempts, repair time and rights position.
The strongest model need not carry every shot. Stable stills, character identity, video motion, lip sync and upscaling can come from different tools, provided the color and asset interfaces can be unified.
34.2 Build a project test set#
The test set comes from the real shot ledger, not from arbitrary landscapes. It should include at least: the lead frontal, in profile and crying; two-person dialogue; a document in hand; an editable screen; interior reverses; low light; walking; simple contact; long fabric; background crowds; lip sync; and camera movement.
Take one ordinary and one stress shot per class. Every candidate uses the same input assets, the same goal and the same attempt budget.
34.3 Evaluate stills and motion separately#
Stills are judged on identity, composition, hands, room for text, materials and light. Motion is judged on identity retention, causality of movement, first and last frames, camera stability, recovery from occlusion, and how much of the clip is cuttable. A good still does not predict video capability.
Supporting first/last frames, multiple references, regional control or lip sync says nothing about actual quality. Verify each feature against your own shots.
34.4 Pass rate and usable seconds#
shot pass rate = candidates reaching the current gate / total attempts
cost per usable second = (generation + human selection + repair + retries) / final usable seconds
Log failures that look good and cannot be cut. If only two seconds of a five-second clip are stable, cost it at two seconds.
34.5 The scoring matrix#
model_test:
model_route: VIDEO_ROUTE_A
shot_class: single_character_micro_action
attempts: 20
pass: 13
usable_seconds: 41
scores:
identity: 4.4
motion: 4.1
camera: 3.8
hands: 3.6
editability: 4.2
labor_minutes: 95
commercial_rights: verified
best_use: single-character micro action, 2-4s
avoid: two-person contact, text during motion
34.6 Route work rather than crowning a winner#
Build a mapping from shot class to production route: static document inserts go to graphic compositing; single-character micro action goes to the identity-stable video model; crowd wides go to a low-detail route; complex effects go to base action plus compositing; long dialogue goes to short lip sync with off-screen coverage.
The showrunner routes by shot risk, budget and deadline, so creators are not choosing tools ad hoc each time.
34.7 Version regression#
A model update can improve motion while breaking faces or color. Run the fixed regression set on any version change and compare against the current production version. A passing new version affects only new tasks; whether shots in progress migrate is decided separately.
Record model, version, parameters, region, date and account licence. A result that cannot be reproduced cannot serve as a reliable production baseline.
34.8 Rights, privacy and availability#
Selection also checks commercial rights, how uploaded assets are handled, retention periods, region, generation labelling and service stability. When a client's likeness or unreleased IP is involved, data terms can block a cloud route outright.
Prepare an alternative vendor and an intermediate format, so project state is not locked inside one platform's project file.
34.9 Testing must control its variables#
The easiest mistake in model testing is giving candidate routes different amounts of help: route A gets retouched reference images while route B gets a rough screenshot; A gets ten retries and B gets three; A is operated by someone who knows the tool and B by a first-time user. What you measure then is input quality, operator experience and budget — not the models.
A formal comparison freezes seven conditions: input asset versions, target duration, aspect ratio and resolution, permitted attempts, the ceiling on human intervention, the pass criteria, and the reviewers. If a tool needs special inputs to show its capability, build a separate best-practice route for it — and count the hours of producing those inputs in the total cost.
Blind review is more reliable. Hide the model and the operator from the candidates and show only shot ID, goal, version and result. Judge usability in the first round and analyse routes in the second. Otherwise brand preference, price expectation or one memorable success quietly moves the standard.
Write the pass criteria before generating. For a single close-up, the blocking items might be: identity below 4/5, wrong primary eye direction, mouth moving without dialogue, hand structure errors persisting past four frames, or fewer than 1.8 seconds of continuous usable material. Deciding the criteria after review turns "I like this shot" into "the model passed."
34.10 Let risk tiers set the size of the test#
Not every shot deserves equal evaluation. Split by consequence into three tiers.
- Tier A is hero shots, paywall cut points, a character's first appearance and key evidence close-ups. Failure damages story or conversion directly, so sample more and retain the full generation record.
- Tier B is ordinary dialogue, reactions and action connections. It is the largest population, judged mainly on stable pass rate and human throughput.
- Tier C is environment, transitions, rear views and functional shots that stock material could replace. Judge on speed, price and ease of substitution.
Suppose a season has 480 shots: 48 in tier A, 312 in B, 120 in C. Route selection cannot chase maximum quality on tier A alone. Ten extra minutes of labor per shot in tier B adds 52 hours; twenty extra minutes in tier A adds 16. Model evaluation must answer both how the best shots are made and how the most numerous shots are made.
34.11 The production route card#
After evaluation, do not publish a model name. Issue an executable route card:
route_id: ROUTE_SINGLE_REACTION_V03
applies_to:
shot_classes: [single_closeup_reaction, silent_insert]
risk_tier: [A, B]
required_inputs:
- approved_keyframe
- identity_reference_set
- state_in
default_settings:
generated_duration: 4s
target_usable_duration: 1.8-2.8s
camera_motion: locked_or_micro_push
attempt_policy:
first_batch: 4
hard_cap: 10
human_review_after: 4
pass_gates:
identity_min: 4.2
continuous_usable_seconds_min: 1.8
fallbacks:
- simplify_to_breath_and_eyeline
- cut_to_over_shoulder
- use_still_alive_frame
owner: generation_supervisor
approved_at: 2026-07-05
The value of a route card is removing improvisation. It states scope, required inputs, default parameters, attempt policy, pass gates and de-escalation. When a new model appears, you do not overturn the workflow — you let it challenge one route card and run the same regression set.
34.12 Shadow production and migration rules#
Passing an offline test does not mean switching a season already in production. Take a small batch of real, undelivered shots and run shadow production: the old route completes them as normal while the new route generates in parallel without affecting schedule. Compare more than image quality — queue time, upload failures, review speed, export formats, color shifts, metadata completeness and turnaround on fixes.
Set explicit migration conditions: two consecutive tier-B batches improving pass rate by at least 15 percent, total cost per usable second falling 10 percent, no increase in identity blocking defects, and a verified backup route. One prettier hero shot does not justify changing the texture of every ordinary shot mid-season.
Shots already locked in the edit do not migrate as a rule. Where appearance, grain or motion rhythm change noticeably, start the new route at a scene, episode or season boundary and note it in the change log. A tool upgrade is not purely a technical event; it is a visual continuity event.
34.13 Route decisions on Backlit Takeover#
Testing found route A stable on Lin Yun's single close-ups, with only a 20 percent pass rate on hands and paper in the two-person document handoff. Route B produced slightly softer faces and completed medium-shot walking more reliably. Precise document text failed on both. So the team did not pick an overall winner; it split the work. Lin Yun's reaction close-ups went to A, two-person mediums to B, the handoff was broken into the initiating hand, the other party's reaction and an insert of the authorization, and the text was composited from approved graphic assets.
The original shot list had Lin Yun walking into the room, opening the authorization and delivering a twelve-word line simultaneously. After repeated stress-test failures, the director did not keep swapping prompts. She rewrote it as four shots: stopping at the door, the room turning, the document landing, and a short line from Lin Yun. Single-shot pass rate rose and the cutting rhythm gained more authority. Model evaluation here is not a procurement exercise; it reshapes the directing language into something producible.
34.14 SOP for model selection#
First, count shot classes. Second, sample a test set from the real shot list. Third, give each model an equal attempt budget. Fourth, evaluate stills and motion separately. Fifth, record pass rate, usable seconds and human hours. Sixth, verify rights and privacy. Seventh, build shot routing rather than seeking a single winner. Eighth, store the baseline and regression set. Ninth, re-check real cost after a small production batch.
34.15 Fault tree#
Symptom: the demo was stunning and the project failure rate is high. The test set does not represent this show's shots. Rebuild it from real shot classes.
Symptom: the per-generation price is low and the total cost is high. Retries and repair hours were ignored. Compute cost per usable second.
Symptom: characters drift after a new version ships. There was no regression testing. Freeze the production version and compare fixed shots.
Symptom: the team argues about tools on every shot. There is no routing matrix. Preset a primary and a fallback per shot type.
34.16 Checklist, exercises and deliverables#
Check that testing comes from the real project; that stills and motion are separated; that failures are logged; that cost per usable second is computed; that rights are verified; that versions are reproducible; that fallback routes exist; and that model updates run regression.
Exercise one: extract a twelve-shot test set from E001. Exercise two: compare cost per usable second across two routes. Exercise three: set a primary and fallback for six shot classes. Exercise four: design the model version regression sheet.
Deliverables for this chapter: the model test set, the evaluation matrix, the shot routing table, the cost benchmark, the rights review, the regression suite, and the fallback plan.
34.17 The unit of evaluation is a production task, not a prompt demo#
Organize the test set by shot class: frontal lip sync, profile reaction, two-person relationship, hands and props, on-screen text, low light, action handoff and crowds. Use real assets, real target durations and real acceptance criteria, and record the full cost from request to approval. A model that produces a beautiful first frame with a low motion pass rate is not suited to carry the video route.
An evaluation report contains at minimum first-pass rate, average attempts, proportion of usable seconds, human repair minutes, latency, price, the distribution of failures, and rights limitations. Single examples are never used for ranking; repeat the same task at different times and versions to observe stability.
34.18 A routing table needs a primary, a fallback and an exit condition#
Choose a primary, a fallback and a manual route for each shot class, and state the conditions that trigger each. Frontal lip sync may use model A, complex hands may use static compositing, environmental wides may use model B. There is no single best model for a whole show.
route_card:
shot_class: two_person_prop_handoff
primary: layered_keyframe_plus_short_motion
fallback: reaction_insert_and_sound_bridge
manual: hand_plate_composite
stop_primary_after: 3_failed_attempts
blockers: [identity_swap, prop_ownership_error]
Exit conditions stop the team from retrying indefinitely on sunk cost. Validate the fallback before production rather than attempting it for the first time when the primary service fails.
34.19 Equivalence and exploiting difference during vendor migration#
Migration does not require a new model to reproduce the old one pixel for pixel. It requires compatibility in identity, state, shot function and style baseline. Compare key metrics on a fixed regression pack, then decide which locked batches stay on the old route and which new batches adopt the new one.
Differences can also be exploited: a model better at environmental motion can take only environmental shots, while lip sync stays with the previous service. The migration log stores model versions, parameters, terms and the switch-over batch, so historical release candidates remain explicable.
A note on sources#
Tool capability and pricing change quickly, so this chapter is not tied to any single product. The stable method is continuous evaluation against your own shots and your own real cost of failure.