中
Chapter 39. Managing Failed Generations and Controlling Cost

Part VII — Generating Images, Video, Performance, and Composites

Chapter 39. Managing Failed Generations and Controlling Cost#

In this chapter
39.1 A failure is not waste — it is production data39.2 The generation ledger39.3 Classifying failures39.4 Attempt caps39.5 Four rework routes39.6 True cost39.7 Ranking candidates39.8 Stop signals in batch production39.9 Feeding failure data back39.10 Attempts must form experiments, not gambling39.11 How to allocate the attempt budget39.12 Daily, per-shot and season costs39.13 Failure clustering and upstream repair39.14 Kill, hold, fallback39.15 The failure ledger on Backlit Takeover39.16 SOP for cost control39.17 Fault tree39.18 Checklist, exercises and deliverables39.19 Every attempt must answer one question39.20 Value-weighted stop rules39.21 Recovering value from failed assetsA note on sources

39.1 A failure is not waste — it is production data#

If a team keeps only successful files, it cannot know which shot types are most expensive, which model keeps failing, or whether a prompt change actually helped. Every attempt should carry a shot ID, input versions, parameters, cost, elapsed time, result and failure tags.

39.2 The generation ledger#

generation:
  id: GEN_E001_S09_004
  shot: E001_S09
  route: VIDEO_ROUTE_A
  model_version: 2026_06
  inputs: [KF_E001_S09@v04, CH_LINYUN@v03]
  prompt_version: MOTION_PROMPT_12
  seed: 18422
  duration_generated: 5s
  usable_range: 1.1-3.4s
  compute_cost: 2.80
  review_minutes: 3
  repair_minutes: 6
  result: usable_after_composite
  defects: [background_text_drift]

39.3 Classifying failures#

Tag identity, hands, props, space, action, physics, lighting, camera, lip sync, text, style, compliance and service errors separately. Then tag the root-cause layer: assets, shot design, prompt, model capability, parameters, compositing, or the requirement itself.

"It doesn't look good" cannot support learning.

39.4 Attempt caps#

Set a cap per shot by value and risk. A low-value transition that fails two or three times should be simplified. A hero shot can absorb more, provided each round changes an explicit variable. Repeating the same request is not an experiment.

Reaching the cap triggers a review: local repair, new keyframe, change model, change shot design, change the action, or change the script. Never retry automatically without limit.

39.5 Four rework routes#

Local blemishes go to retouching or compositing. A wrong static foundation goes back to the keyframe. A mismatch in motion capability goes to a different route or a shot split. A requirement that is simply not producible goes back to shot design or the script. The earlier you return to the right layer, the cheaper it is.

Covering with the edit suits only brief errors that do not affect fact. You cannot push every failure into post.

39.6 True cost#

Cost includes API or subscription fees, human selection, repair, transfer, waiting, asset production, failed retries and opportunity cost. Account for it by shot and by final usable second.

true shot cost = all attempt fees + human minutes x rate + compositing + QC + rework
cost per usable second = true shot cost / seconds adopted in the cut

Record cycle time too. A cheap route with a two-day queue can still miss delivery.

39.7 Ranking candidates#

Rank blocking items first, then compare narrative clarity, continuity, performance and aesthetics. A candidate with flashier motion and a character who does not look like herself must not win.

Keep the distinction between "best usable" and "best direction, needs repair," so the team does not mistake repair potential for completion.

39.8 Stop signals in batch production#

Persistently low pass rates in one shot class, cost per usable second over budget, human selection becoming the bottleneck, one asset causing drift across many shots, or audience tests showing indifference to a set piece — any of these should pause scale-up.

Efficiency optimization cannot rescue content nobody wants. Prove the story and the shot's value first, then optimize throughput.

39.9 Feeding failure data back#

If a document-in-hand shot fails repeatedly across twenty shots, change the shot template and the action rules rather than retrying each one. If a costume bleeds color in low light, update the asset. If multi-person lip sync fails often, adjust the shot routing.

Summarize failure types, pass rates, cost and upstream root causes weekly, and update schemas, prompts and model routing.

39.10 Attempts must form experiments, not gambling#

Before each retry, write the hypothesis, the variable being changed, and how the result will be judged. If you change the reference image, prompt, duration, model and camera movement simultaneously, a success tells you nothing about why. Exploration may widen differences to find a direction; in production, change one primary variable per round.

retry_decision:
  previous: GEN_E001_S09_004
  observed_defect: an extra finger appears on the right hand after the document slides
  hypothesis: motion amplitude and hand occlusion are jointly causing structural drift
  change:
    motion_distance: 12cm_to_7cm
    all_other_inputs: frozen
  success_condition: five stable fingers throughout, document crosses the table midline
  result: failed_same_defect
  next_action: route_to_hand_composite

Two consecutive changes to the same variable with no improvement should lower your confidence in that hypothesis. Three similar failures should trigger a pause rather than continued luck. For a production system, knowing a route does not suit the task is itself a valuable result.

39.11 How to allocate the attempt budget#

A shot's budget is set by four factors: narrative value, probability of failure, available alternatives, and downstream impact. A hero shot with high value but a reliable live-action alternative does not necessarily warrant a high cap. An ordinary connecting shot with low value that could become a sound bridge should not consume a dozen attempts.

A simple decision score helps:

value of another attempt = narrative gain if it succeeds x probability of success this round
                          - cash and labor cost this round
                          - the impact of delay downstream

This is not about pretending to be precise. It forces the team to compare the future cost of "one more try" against "change the shot design." Sunk cost does not enter the formula. Twenty attempts already spent is not a reason for the twenty-first.

39.12 Daily, per-shot and season costs#

Daily cost controls cash burn and staff load. Per-shot cost compares routes. Season cost identifies structural risk. Do not merge the three into one average.

One hero shot may be very expensive and appear three times, which the season can absorb. An ordinary reverse that overruns slightly on each of two hundred shots creates the largest gap. The production dashboard should show attempts, pass rate, human minutes, wait time and usable seconds aggregated by shot class — not just total generation credits.

Also separate "generation complete" from "deliverable complete." The former may not yet include keying, lip sync, graphics, grading and QC. If route A costs 30 to generate and 90 in post while route B costs 60 and 20, reading the platform invoice alone produces the wrong choice.

39.13 Failure clustering and upstream repair#

Cross-tabulate weekly by shot type, character, wardrobe, location, action, model and failure tag. When one error repeats across many shots, look for the shared parent first. Drift in the nose bridge across all of Lin Yun's profiles may mean the profile golden reference is insufficient. Black suits swallowing the silhouette in every night scene may be a wardrobe and lighting design problem. Hand failures in every document handoff mean the action template needs updating.

After an upstream repair, list the affected scope and a regression sample. Do not fix only the new shots and leave the ones that passed while using the old faulty asset. The problem ledger records "root cause fixed" and "instances fixed" as separate states, so nobody believes repairing one image closed a systemic issue.

39.14 Kill, hold, fallback#

Reaching the attempt cap requires an explicit decision.

  • KILL: the shot or effect is no longer made, and its story function is carried elsewhere.
  • HOLD: paused pending assets, model capability or a client decision — with no further budget consumed.
  • FALLBACK: switch immediately to a predefined de-escalation route, such as a living still, off-screen sound, an insert, a rear view, compositing, or a shot design change.

"Try again later" is not a state. It leaves unfinished tasks in the queue consuming attention. A hold must have a release condition and a date. A fallback must have an owner and new acceptance criteria. A kill must confirm that the story information was not deleted along with the shot.

39.15 The failure ledger on Backlit Takeover#

The pilot's most expensive shot was originally Lin Yun walking beside the long boardroom table while meeting the eyes of six people. The team generated it 27 times across three models and still saw identity drift, the table stretching, and background faces changing. Clustering the failures showed this was not a prompt problem. One shot contained long-distance movement, multiple identities, a reflective table surface and a complex camera retreat simultaneously.

The producer triggered a kill-or-redesign review. The director kept the story function and deleted the technical form: a heel stopping at the door, three of the six reactions, a static medium close-up of Lin Yun, and an insert of the document on the table. The new approach passed within 14 generations, and the finished sequence has better rhythm than the original virtuoso long take. The earlier 27 attempts were logged as route learning rather than amortized into the new shot — which would have created a false per-shot price — while still entering pilot R&D cost and the baseline for later budgets.

39.16 SOP for cost control#

First, establish pass-rate baselines per shot class. Second, set a budget and attempt cap per shot. Third, record every attempt. Fourth, use specific failure tags. Fifth, change one explicable variable per round. Sixth, trigger a rework route review at the cap. Seventh, compute true cost and usable seconds. Eighth, aggregate recurring root causes and fix upstream. Ninth, scale up only what has been shown to work.

39.17 Fault tree#

Symptom: the budget empties fast. There are no attempt caps and no failure classification. Stop automatic retries and review root causes.

Symptom: generation is cheap and the team is slow. Selection and repair hours are uncounted. Compute true cost.

Symptom: the same class of error appears every episode. Only finished shots are being fixed, with no update to shot design, assets or routing. Run failure clustering.

Symptom: nobody dares abandon a failing shot. Sunk cost is driving judgment. Compare the future cost of repair against redesigning the shot.

39.18 Checklist, exercises and deliverables#

Check that every attempt is linked to a shot ID; that failures are specific; that attempt caps exist; that each round changes a variable; that real labor cost is included; that accounting is by usable second; that recurring problems feed upstream; and that scale-up has stop signals.

Exercise one: build a ledger for ten failures. Exercise two: compare the cost of regenerating, compositing and redesigning. Exercise three: find the systemic root cause among fifty failures. Exercise four: set different caps for hero shots and transitions.

Deliverables for this chapter: the generation ledger, the failure taxonomy, attempt budgets, candidate ranking, the true cost report, rework decisions, and the weekly failure review.

39.19 Every attempt must answer one question#

An attempt record states the hypothesis, the single primary variable, input versions, the expected improvement and the stop condition. Changing model, prompt, references, duration and composition at once makes even a success unreusable. A failure is not "one more roll" — it updates your judgment about the root cause.

Figure 39-1 Generation attempts as stoppable experiments with hypotheses, single variables and frozen inputs

Figure 39-1 Each round changes one primary variable while the other inputs stay frozen, and candidates face the same pass gate. When the attempt cap is reached without passing, the route stops and moves to an approved fallback. The dice at the bottom represent the undocumented gambling that is forbidden.

The ledger in the lower left stores each round's variable, candidates, result and cost. Even when a round passes by chance, judge whether it supports the original hypothesis; an unexplained lucky result may be used for the current shot without being promoted into a general method.

attempt:
  hypothesis: profile drift is caused by the missing angle-matched identity reference
  change: add_CH_LINYUN_LEFT45_v03
  frozen: [model, seed_family, wardrobe, lighting, motion]
  success: identity_score_and_human_pass
  stop_after: 2

39.20 Value-weighted stop rules#

The attempt cap is set jointly by narrative value, alternative routes, cost already spent, probability of success and delivery time. A hero cut point may permit more rounds; a background transition moves to fallback quickly. Sunk cost does not raise the probability of the next success and is never a reason to continue.

Kill means this route stops. Hold means waiting for new assets or model capability. Fallback means adopting an approved alternative. Each state retains its evidence, so that next week someone else does not repeat the same failure from scratch.

39.21 Recovering value from failed assets#

Rejected candidates can train internal evaluation, build fault examples, test repair tools and estimate route risk — while staying quarantined from golden references and observing rights and data limits. Wrong faces, garbled evidence and fused hands go into separately tagged failure sets.

Cluster frequent failures weekly and estimate the return on a systemic fix. If 30 percent of retries come from one missing rear-view costume asset, adding the asset is more economical than continuing to optimize prompts. A failure ledger should ultimately reduce future failures, not merely explain this week's overspend.

A note on sources#

The per-episode costs creators publish often omit failures and labor. Accounting in usable seconds, attempts and repair hours is what reveals whether a workflow actually scales.