中
Chapter 58. Orchestration, Budget, and Failure Recovery

Part XI — The Agentic Production System

Chapter 58. Orchestration, Budget, and Failure Recovery#

In this chapter
58.1 The workflow is a DAG, gates outrank throughput58.2 Budget, recovery, de-escalation, and human takeover58.3 Diagnosis, a worked DAG, and crash drills58.4 Priority, circuit breakers, dead letters, backpressure, and fan-out58.5 Compensating transactions, SLOs, and real recovery capabilityA note on sources

58.1 The workflow is a DAG, gates outrank throughput#

The workflow is a directed acyclic graph.

Tasks form a DAG by dependency. Video depends on approved keyframes; keyframes depend on shot design and assets; subtitles depend on final audio; release depends on the master, rights and QC. The scheduler runs only tasks whose dependencies are satisfied.

Gates outrank throughput.

Parallel generation burns through mistakes faster. While script, assets and keyframes are unapproved, expensive tasks stay blocked. Before widening a batch, complete a small number of canary shots to confirm the model route and the cost.

The task record.

task:
  id: TASK_MOTION_E001_S09
  type: motion_generation
  dependencies: [KF_E001_S09_APPROVED, CONT_E001_S09_VALID]
  input_hash: abc123
  budget: {attempts: 4, cost: 20, wall_time: 30m}
  retry_policy: technical_errors_only
  status: ready
  idempotency_key: E001_S09_motion_v04

Concurrency limits.

Limit concurrency by vendor quota, network, human review capacity, and shared-asset risk. Generating a hundred shots at once with one reviewer simply queues up errors; pass a small batch first, then widen.

Timeouts and retries.

Network timeouts and rate limiting can retry with exponential backoff. A content failure must not retry unchanged. Every retry records the reason and what parameters changed. Reaching the cap creates an issue.

58.2 Budget, recovery, de-escalation, and human takeover#

Budget.

Budget divides by project, stage, episode, shot and task. Reserve before scheduling and settle real cost on completion. High-value shots can request more; low-value shots that exceed the cap go automatically to a shot-design review.

Resuming from a break.

Every task output is written atomically: files are staged, verified, then marked complete. After an interruption, the scheduler reads task records and does not repeat work already complete with a matching hash.

Vendor unavailability.

The routing table stores primary, fallback and any input conversion. A fallback must have passed the regression set in advance. After switching, generate a canary first rather than migrating the whole queue.

Human takeover.

Repeated failures, factual conflicts, rights blocks, budget overruns and aesthetic judgment enter the human queue. A person may re-route, change inputs, approve a waiver or stop — never edit the database directly to bypass events.

SOP.

First, generate the DAG from the dependency graph. Second, set budgets and idempotency keys. Third, schedule only tasks past their gate. Fourth, widen concurrency after a small canary. Fifth, separate technical from content retries. Sixth, write atomically and verify. Seventh, prepare fallback routes. Eighth, have human takeovers write events back.

58.3 Diagnosis, a worked DAG, and crash drills#

Fault tree.

Symptom: duplicate charges after recovery. There is no idempotency key or output hash. Deduplicate by task identity.

Symptom: an asset error is discovered only after a whole batch completes. There was no canary and no gate. Validate a small batch first.

Symptom: one vendor outage stops everything. The fallback was never tested. Maintain an alternative route that passes regression.

Checklist, exercises and deliverables.

Check that the DAG is acyclic; that budgets are tiered; that retries are classified; that tasks are idempotent; that outputs are atomic; that fallbacks are verified; and that human takeovers leave logs.

Exercise one: draw the DAG for E001 from pack to release. Exercise two: design recovery from a power failure. Exercise three: simulate a canary switch when the video vendor goes down.

Deliverables for this chapter: the task DAG, the scheduler policy, budget reservations, retry rules, the recovery runbook, fallback routing, and the human queue.

The E001 task DAG.

E001_PACK_APPROVED
├─ ASSET_RESOLVE
│  ├─ CHARACTER_REFS_READY
│  ├─ LOCATION_READY
│  └─ PROP_GRAPHICS_READY
├─ CONTINUITY_INITIALIZED
└─ STORYBOARD_E001
   └─ STORYBOARD_APPROVED
      └─ KEYFRAMES_S01-S19
         └─ KEYFRAME_QC
            └─ MOTION_S01-S19
               └─ MOTION_QC
                  └─ ASSEMBLY_EDIT
                     ├─ AUDIO_FINAL
                     └─ GRAPHICS_FINAL
                        └─ PICTURE_LOCK → MIX/SUBTITLE/COLOR → RELEASE_GATE

Keyframes can run in parallel once shot design is approved. While a character's new golden version is not yet stable, limit concurrency, or one error is generated into nineteen shots at once.

Budget reservation and settlement.

Before a task enters running, reserve the worst permitted cost from the project allowance; on completion, settle at real cost and release the remainder. A task with no reservation cannot call paid generation. Additional budget must state the previous attempts, the root cause and the alternatives.

Budget is not only money. It covers human review minutes, concurrency slots and deadlines. A route that generates cheaply and needs twenty minutes of human selection may still be demoted by the scheduler.

A crash recovery drill.

Suppose a motion batch crashes at S12. The recovery process reads the task records. S01–S08 have files, hashes and QC and stay complete. S09–S10 have files that were never committed; after verification they move to needs_review. S11 was mid-upload with an incomplete file, so the temporary object is deleted and the task retried. S12 was never called and is not charged.

If the vendor returned success while the local side timed out, query the result by the vendor's request ID rather than calling again. Where that is impossible, record the potential duplicate charge as a cost anomaly.

58.4 Priority, circuit breakers, dead letters, backpressure, and fan-out#

Priority and starvation.

A P1 fix blocking release outranks exploration on a new episode — and long-running exploration must not starve permanently. The scheduler queues by project stage, deadline, degree of blocking and resource type, reserving a small capacity for low priority.

The human review queue is ordered the same way: rights and factual errors first, hero shots above background polish. Shots sharing one root cause are reviewed as a group to reduce repeated judgment.

Circuit breakers and de-escalation.

When a model shows a sustained high error rate, timeouts or cost anomalies, trip a circuit breaker and pause new tasks. The system can de-escalate to a living still, reduced motion, an alternative model or a shot redesign — and any de-escalation must preserve story state and acceptance criteria.

De-escalation does not automatically lower the quality gate. When a core shot cannot be satisfied, it stays blocked for a human to decide.

The dead letter queue.

Tasks past their retry cap, with permanently missing dependencies, or whose output cannot be validated, enter a dead letter queue. Each retains the original task, its attempts, errors, cost and the suggested station. After a human fixes the input, a recovery event is created — you do not clear the history by clicking retry.

Disaster recovery targets.

Define an acceptable data loss window and recovery time. Facts, rights and approved assets use stricter backup; low-value regenerable candidates can use a lower tier. Rehearse database corruption, object storage loss, key rotation and vendor account unavailability regularly.

Successful recovery is not the service restarting. It is rebuilding a specified RC from snapshot and events with matching hashes, subtitles, audio and rights records.

Backpressure.

When keyframe review backs up, the GPU queue congests, or human repair overruns, the scheduler applies backpressure upstream: fewer new shot batches, low-priority variants paused, release-blocking tasks retained. Without backpressure, a local bottleneck expands into hundreds of expired candidates and higher storage cost.

Record the time, reason and owner when backpressure is released as well.

Batch orchestration and fan-out control.

The task graph permitting parallelism does not mean fanning out every shot at once. The scheduler picks a canary per shot class, validates assets, route, cost and evaluators, then expands by approved batch. Shots sharing a new character version especially need limited fan-out, or one identity error contaminates dozens of tasks simultaneously.

Batch size adapts to downstream review capacity. When the keyframe queue waits beyond a threshold, stop new composition exploration. When mixing blocks release, reduce video generation for later episodes and move resources to the batch nearest completion.

58.5 Compensating transactions, SLOs, and real recovery capability#

Compensating transactions.

Media production cannot roll back in place the way a database can. A completed task may already have triggered billing, uploads, human review and downstream compositing. Failure recovery needs compensating actions: mark the wrong asset invalid, cancel dependent tasks that have not started, release reserved budget, retain the logs already produced, and create an alternative route.

Compensation does not delete history. If a bad video already entered the rough cut, the system creates a new edit task replacing the source and invalidates the old export — it does not overwrite a same-named file in a folder. Every high-impact task declares its compensation_action at design time.

Production SLOs and observability.

Monitor at minimum task success rate, queue duration, first-pass rate, retry rate, cost per usable second, human review wait, invalidation propagation time, and RC rebuild success. A service being online does not mean production is healthy: a generation endpoint at 99.9 percent availability with a collapsing pass rate should trip the breaker too.

Order alerts by consequence to audience and project. Expired rights, factual conflicts and a changed release-file hash alert immediately. One background candidate timing out can reschedule itself. Dashboards show trends against a baseline rather than dumping every yellow warning on one producer.

The human queue is a real workflow.

A human takeover item carries a context package, failure history, budget already spent, candidate options and the decision required. Sending "please take a look" forces the person to re-investigate the whole chain. Human decisions are written as events, from which the scheduler resumes, re-routes, adds budget or terminates.

The queue has response deadlines and named backups. A high-risk project must not depend on one director coming online to learn what happened; the on-call person should be able to understand the minimum sufficient context from the approval package while holding no authority over facts that are not theirs.

Workflow instances and state recovery.

The DAG template describes possible paths; a workflow instance records the nodes one episode actually traversed, with input hashes, attempts, cost and decisions. After a restart, the scheduler reads instance state and resumes only nodes that are incomplete or safely retryable — it does not re-execute a generation that was already billed and uploaded successfully.

Node states include at minimum pending, leased, running, waiting_human, succeeded, failed, compensating, cancelled and stale. When running times out, query the vendor's result before taking over; retrying directly can produce duplicate charges and two competing results.

Budget is a workflow resource, not a month-end report.

Reserve before a task starts, settle at actual cost and release the remainder afterwards. Retries, human repair, vendor switches and storage all consume the same production unit. At a soft limit, reduce variants and concurrency; at a hard limit, stop and request approval — an agent must never be able to overdraw indefinitely by trying again.

Budget policy differentiates by shot value. A hero cut point retains more repair attempts; a background transition moves to a simplified approach at its cap. The scheduler weighs money, GPU, human minutes and delivery time together, rather than optimizing API spend while pushing the cost onto the editor.

Fault injection and recovery drills.

Outside delivery peaks, inject controlled faults regularly: model endpoint timeouts, mismatched object hashes, an approver offline, a rights event suddenly blocking, a database restored from snapshot, a failed platform upload. Observe whether alerting, backpressure, human takeover, compensation and RC rebuild happen as the runbook says.

Drills record recovery time, data loss, duplicate charges, orphaned tasks and human cognitive load. A runbook that works only on the happy path is not recovery capability. Change a small number of rules per drill and re-test, so one incident does not add a mass of mutually conflicting automation.

A note on sources#

Orchestration for creative production differs from ordinary batch computing: the expensive resource is not only compute but human judgment, and the failure mode that matters most is confidently producing something wrong at scale.