Part XVI — Engineering Implementation and Production Infrastructure
Chapter 83. Model and Vendor Selection: Buy Shot Capability, Not Leaderboard Rank#
In this chapter
Model selection is not finding one strongest model to carry a whole series. Different shots make different demands on identity, action, text, camera movement, lip sync, speed, cost and rights. The engineering goal is capability routing: every task lands inside a validated stable region, and migration stays safe when vendors change.
83.1 Build the selection baseline from task risk#
Build the task taxonomy before the vendor table.
The project divides shots into nine classes: single-character static close-up, single-character speech, two-person reverse, multi-person shared frame, prop contact, character movement, environment establishing, graphic screens, and effects compositing. Audio divides separately into dialogue, long narration, singing, music themes, music extensions, Foley and noise reduction.
task_profile:
id: video_two_person_prop_handoff
required:
identity_count: 2
prop_state_transition: exact
start_end_frame_control: true
duration_s: 2_4
commercial_rights: required
preferred:
camera_motion: subtle_push
deterministic_seed: useful_not_required
forbidden:
generated_readable_contract_text: true
acceptance_set: eval_handoff_v05
Only with stable task definitions are vendor results comparable. Testing every model with one generic prompt about an executive walking into a meeting room compares first aesthetic impressions and cannot support production routing.
The evaluation set must come from real risk.
Select twenty-four representative shots: the lead frontal and in profile, speaking, strong backlight, two-person occlusion, the document handoff, movement in rain, glass reflections, a phone screen, and a longer take. Each sample has frozen inputs, acceptance dimensions, a human golden judgment and failure tags. Keep the evaluation set in two parts — a public benchmark and a private project set — so the team cannot tune only to the test.
On a model upgrade, run the same inputs without demanding pixel identity. Compare narrative state, identity, action, cuttability, cost and elapsed time. A new version that is prettier overall while lowering the handoff pass rate is not a global replacement.
83.2 Turn blind evaluation into model routing#
The provider capability card.
provider_capability:
provider_id: provider_video_a
model_version: model_2026_06
validated_tasks:
single_closeup: {pass_rate: 0.91, median_cost_usd: 3.2}
two_person_handoff: {pass_rate: 0.58, median_cost_usd: 8.7}
environment_motion: {pass_rate: 0.88, median_cost_usd: 2.9}
unsupported: [exact_text, three_person_contact]
operational:
p95_latency_s: 168
rate_limit: 20_per_minute
idempotency: false
rights_profile: rights_provider_a_2026_06
last_validated_at: 2026-07-05
Pass rates must carry sample size and a confidence interval. Three samples all passing is not 100 percent. Cost uses cost per approved second, not the price of one call.
Blind and pairwise comparison.
The review interface hides vendor and parameters and randomizes left/right order. Judge blockers first, then make pairwise preference judgments, so a composite aesthetic score cannot mask a state error. Identity, action and fact are thresholds; performance, composition and texture are compared only above them.
At least two reviewers judge independently, with disagreements arbitrated. When one model always scores highly with the visual director and fails with the continuity lead, the issue is not who understands more — it is different metric weights and different use cases.
Routing policy.
Low-risk environment shots can route cost-first. Principal close-ups route identity-first. The document handoff uses the most stable model and automatically generates insurance coverage. Urgent pickups can take a low-latency route with an extra human gate. The router outputs a recommendation and its reasons; it does not approve a final take.
routing_decision:
task_id: sh_e004_032
selected_provider: provider_video_b
policy: identity_first_closeup
evidence: {identity_pass_rate: 0.94, latency_within_deadline: true}
fallback_order: [provider_video_a, composite_still_motion]
requires_human_gate: true
83.3 Treat commercial terms as part of availability#
Terms are also capability.
Record whether inputs are used for training, commercial rights in outputs, likeness and voice restrictions, data territory, deletion mechanisms, retention periods, incident notification and the terms version. A model that passes on visual capability with rights unsuited to the project routes as unavailable.
A change of terms triggers impact analysis: which unreleased assets that model generated, whether re-licensing is needed, and whether later derivatives are permitted. Legal review cannot wait until a season is finished.
Migration and de-escalation.
Every task class keeps at least one narrative de-escalation route: when the video model is unavailable, use keyframes with light motion; when precise lip sync fails, move to reaction shots and off-screen voice; when multi-person contact fails, split into coverage. A de-escalation states what it sacrifices rather than quietly lowering quality.
Migration runs the evaluation set first, then one complete previsualization, then a small staged rollout. A model change affects prompts, reference weights, durations, color and cost — you cannot simply change an API name.
83.4 Selection acceptance and a committee exercise#
Fault tree.
High on a leaderboard and failing on the project: the test tasks do not represent real shots. Cheap calls with high total cost: low approval rate or heavy human repair. Faces more stable after an upgrade and the edit harder: motion and handles were not part of acceptance. The backup model cannot take over: context and outputs are locked into a vendor's formats. Rights risk appearing suddenly: the terms version never entered the dependency graph.
SOP, checklist and deliverables.
Define the task taxonomy. Build the private evaluation set. Run blind evaluation. Compute cost per approved second. Produce capability cards. Add rights and operational metrics. Write routing and de-escalation. Run a migration drill. Re-test monthly or after any significant version change.
- Selection samples come from the project's real risks.
- Threshold metrics are separated from aesthetic preference.
- Cost includes failures, labor and waiting.
- Terms are locked together with the model version.
- Every critical task has an explicable fallback.
Exercise: compare three candidate services across twelve shot tasks, delivering task_taxonomy.yaml,
evaluation_set/, provider_cards/, pairwise_results.csv, routing_policy.yaml and migration_report.md.
A selection committee exercise.
Have visual, continuity, production, legal and engineering each cast one vote — and require each to write their veto condition before voting. Visual may prefer performance, continuity may veto on state errors, legal may block on usage terms, production watches cost per approved second, and engineering watches stability and observability. The outcome is not a majority vote: every hard threshold must pass first, and the trade-off happens only among the remaining feasible options.
Include in the exercise one vendor with the most attractive imagery whose terms prohibit a class of commercial campaign, and one with the lowest per-call price and the highest failure rate. If the team still selects on impression, the capability cards are not genuinely participating in the decision. Minutes record the reasons for rejection, applicable tasks and the conditions for re-evaluation, so the same argument does not restart from scratch three months later when someone changes role.
The evidence that capability routing passed acceptance is that the same batch of tasks yields an explicable, repeatable choice under the rules — with reasons recorded whenever a human overrides — not that the router happened to recommend the model the panel liked.
Have someone outside procurement re-check the original samples and the calculation basis.
A note on sources#
Vendor capability and pricing move quickly, so this chapter avoids naming products. What transfers is deriving evaluation from your own shot risk, treating terms as capability, and keeping a rehearsed migration path.