中
Chapter 85. The Reproducible Generation Runner: Queues, Budgets, Retries, Evidence

Part XVI — Engineering Implementation and Production Infrastructure

Chapter 85. The Reproducible Generation Runner: Queues, Budgets, Retries, Evidence#

In this chapter
85.1 Reproducible tasks and budget as a precondition85.2 Data transfer, verifying results, and classifying retries85.3 Task lifecycle and production observability85.4 Production acceptance, handover, and on-call drillsA note on sources

Once a prompt is compiled, something has to call the external service safely. That runner is not a few lines of API request. It combines task leases, idempotency, budget, upload and download, verification, retries, cancellation, logging and candidate ingestion.

85.1 Reproducible tasks and budget as a precondition#

The task record.

generation_task:
  task_id: gen_e001_sh015_v04
  status: ready
  prompt_manifest: prun_e001_015_a04
  capability_token: cap_gen_015
  budget_reservation: budget_015_04
  priority: high
  deadline: 2026-07-06T04:00:00Z
  attempt_limit: 4
  lease_owner: null
  lease_expires_at: null
  output_contract: candidate_video.v2

A worker takes a short lease; after a crash the lease expires and the task can be re-claimed. State updates use optimistic concurrency so two workers cannot submit simultaneously.

Idempotent calls.

Keep the internal task ID, the attempt ID and the vendor request ID separate. Each change of a controlled variable creates a new attempt; a network timeout re-sends the same attempt under the same idempotency key. Where a vendor supports no idempotency, the runner registers its intent to call first and then queries for a remote job that may already have completed — rather than paying twice immediately.

attempt:
  attempt_id: att_gen015_04
  idempotency_key: gen015:prompt_sha_ddd:attempt04
  changed_variable: micro_eye_shift
  held_constant: [model, seed_family, keyframes, reference_bundle]
  provider_job_id: remote_90211
  status: downloading

Budget is a precondition for execution.

Before calling, check budget at project, episode, scene and task level and create a reservation. Settle at actual cost on return, recording whether failures are billable under the vendor's rules. A call with no reservation is refused technically, and an administrator cannot bypass it with a comment.

When budget is exhausted, preserve the candidates obtained, the evaluations and the logs, and move the task to awaiting_budget_decision. The system does not silently lower resolution or switch to a cheaper model, because that would change the acceptance conditions.

85.2 Data transfer, verifying results, and classifying retries#

Uploads and privacy minimization.

Upload only the cropped references this task needs, name files by anonymized asset ID, and strip unrelated metadata. Original likeness releases, complete scripts and future-episode secrets never leave. Verify rights and data policy before upload; after download, request deletion or record the expiry according to the vendor's retention policy.

Compute checksums for every input so a same-named file cannot be substituted. Signed URLs are short-lived and bound to read-only objects, and logs do not record full sensitive addresses.

Download, verification, and quarantine.

Results download into a quarantine area and are verified for file type, magic number, dimensions, frame rate, duration, corruption, malicious content and checksum. The file extension in a vendor response is not trustworthy. Passing verification means technically readable — not creatively approved.

Candidates record input lineage, model, time, cost and usage limits. Media with missing provenance never enters the asset library, however much a human likes it.

Classifying retries.

Rate limits, timeouts and transient server errors retry with exponential backoff. Schema errors, permission denials and rights blocks are not retryable. Quality failures require an evaluator or a human to choose a new controlled variable first; the runner must not repeat blindly. Retry limits are bounded by count, cost and deadline together.

Consecutive vendor failures trip the circuit breaker. With the breaker open, new tasks fail fast and route to a human or a compatible fallback, which prevents queue pile-up and runaway budget.

85.3 Task lifecycle and production observability#

Cancellation and expiry.

When an upstream episode pack updates, waiting tasks are marked superseded. Running tasks attempt remote cancellation; where cancellation is impossible, the result is quarantined on completion rather than entering review. On arrival, check the input version again, so a late result from an old task cannot contaminate a new batch.

An expired deadline does not mean deletion. Preserve the execution evidence and the reason it did not complete, so the producer can re-schedule, de-escalate the shot design, or cancel the requirement.

Observability.

Metrics include queueing, compilation, upload, vendor processing, download, verification, human waiting, failure rate by class, cost per approved second and budget variance. Traces run from the episode pack through to the remote request and the asset ID. Logs are structured with sensitive fields masked.

An API returning 200, a remote job completing, a file passing verification and a creative acceptance passing are four different success events. The dashboard must display them separately.

Minimum and production implementations.

A small team can start with a single machine, a local database, directory-based object storage and one worker — and still needs task IDs, a state machine, idempotency keys, budget fields and manifests. At scale, migrate to a durable queue, object storage, centralized events and multiple workers; the business semantics stay the same.

Do not start from a complex cluster, and do not let folder polling become a permanent source of truth. The minimum implementation must keep stable interfaces for migration.

85.4 Production acceptance, handover, and on-call drills#

Fault tree.

The same task billed twice: timeout retries lacked idempotency. An old result entering a new cut: the upstream version was not re-checked at submission. A task stuck in running: leases and heartbeats are missing. A downloaded file that opens and cannot be traced: ingestion did not enforce a manifest. Infinite retries on quality failure: the runner exceeded its authority and chose creative variables.

SOP, checklist and deliverables.

Create the task and the reservation. Take a lease. Verify the token and the inputs. Register the intent to call. Call and poll. Download to quarantine. Verify technically. Settle the budget. Submit the candidate. Trigger independent evaluation. Handle cancellation, retries and the circuit breaker.

  • Network retries do not create new creative attempts.
  • Every call carries a budget reservation.
  • Results from old versions cannot submit automatically.
  • Sensitive inputs follow minimal upload.
  • Technical success and creative success are separated.

Exercise: implement a runner against a mock vendor and inject timeouts, duplicate responses and late-arriving results, delivering task_schema.json, worker/, provider_mock/, budget_ledger.csv, trace_samples/ and failure_test_report.md.

Handover and an on-call drill.

Have someone who did not write the runner handle a remote timeout. They should find the call intent from the trace, confirm whether the vendor job exists, judge whether a retry is safe, check the budget reservation, and complete recovery without consulting a developer's private messages. If the original author must explain, the runbook and the observability are not yet adequate.

Then simulate an upstream pack locking a new version while a task is running. The correct outcome is that the old remote job is quarantined even if it succeeds, with its incurred cost preserved, while a new task uses a new idempotency key and a new budget reservation. The on-call person must not rename an old video by hand into the new directory. The drill report records the discovery, the judgment, the action, the recovery and the evidence gaps — and turns the gaps into automated tests.

A runner may enter a real budget environment only once state and ledgers stay consistent across all five conditions: duplicate delivery, process crash, network timeout, budget exhaustion and upstream update.

A note on sources#

This chapter describes an execution pattern rather than a specific stack. What transfers is separating attempts from retries, making budget a precondition, quarantining results, and keeping technical success distinct from creative approval.