Part XI — The Agentic Production System
Chapter 55. An Agent Is Not a Role-Play — It Is a Workstation with Boundaries#
In this chapter
55.1 From station responsibility to least privilege#
A name is not an architecture.
Calling a model "award-winning screenwriter" does not produce a system. A workstation needs explicit inputs, the facts it may read, the outputs it may write, a budget, approval conditions, and prohibited behaviors.
Eleven core stations.
The showrunner manages state and gates. Market analysis generates opportunity hypotheses. Story architecture maintains the promise and the season. The writers' room produces episode packs. The asset director maintains the registry. The shot director produces the shot ledger. Generation production executes routes. The edit director maintains the timeline. QC manages issues. Distribution analysis manages creatives and metrics. Rights and compliance maintain provenance and release restrictions.

Figure 55-1 The agentic production architecture. Agents exchange verifiable information through project state, the event log and the dependency graph; they cannot overwrite each other's facts directly. Candidate results form a release candidate only after QC and human approval.
The center of the diagram is not an all-powerful controller model. It is auditable production memory. Agents read the state they are permitted to read, submit events and candidate artifacts, and the state machine decides whether work may advance. Human approval sits at high-impact gates rather than as decoration at the end of every automated step. That architecture is what allows one model or station to be replaced without stuffing the whole work back into a single chat context.
The agent contract.
agent:
id: AGENT_STORYBOARD
reads: [approved_episode_pack, asset_registry, continuity_ledger, style_bible]
writes: [storyboard_rows, asset_requests, continuity_risks]
may_not:
- modify approved story facts
- invent unregistered characters or props
- start video generation
budget:
max_revision_rounds: 3
advance_when:
- all_assets_resolved
- storyboard_qc_no_P0_P1
human_approval: required_before_keyframes
Least privilege.
The writers' room may request new assets and may not approve them. The generation station may produce candidates and may not mark its own results as finally passed. QC may block and may not quietly edit facts. Separating permissions prevents one agent from both setting the exam and grading it.
55.2 Approvals, failure escalation, and the operating flow#
Human approval points.
Premise, lead casting, styleframes, the first paywall, golden assets, the pilot rough cut and release all require a person. Automation may continue for low-risk format conversion and within already-approved rules.
The approval interface shows differences, evidence, cost and impact — not only accept and reject.
Failure and escalation.
When an agent encounters a missing fact, conflicting versions, a budget overrun or repeated failure, it stops and creates an issue rather than guessing. Three consecutive failures of the same kind escalate upstream instead of retrying.
SOP.
First, list the production stations. Second, write read and write permissions for each. Third, set budgets and stop conditions. Fourth, define output schemas. Fifth, set human approvals. Sixth, test privilege escalation and missing inputs. Seventh, record every action in the audit log.
Fault tree.
Symptom: agents overwrite each other's results. Write permissions are not separated. Establish a single owner per asset.
Symptom: automation burns credits continuously. There is no budget and no failure escalation. Add attempt caps and blocking.
Symptom: results look complete and the facts are wrong. The model was allowed to fill in missing data. Missing data must raise an issue.
55.3 Permission matrices, complete handoffs, and approval packages#
Checklist, exercises and deliverables.
Check that every station has inputs, outputs, permissions, a budget, stop conditions, approvals and prohibited behaviors; that you can trace who changed a fact; and that no station approves itself.
Exercise one: write an agent contract for a music agent. Exercise two: design a privilege-escalation test. Exercise three: decompose "make the whole series" into stations with owners.
Deliverables for this chapter: the agent registry, agent contracts, the permission matrix, the approval map, and escalation rules.
A permission matrix.
| Artifact | May draft | May change facts | May approve | May block |
|---|---|---|---|---|
| Opportunity brief | market agent | showrunner | producer | commercial / QC |
| Series bible | story architecture | showrunner | human creative lead | QC / legal |
| Episode pack | writers' room | story editor | human creative lead | QC |
| Character golden assets | asset director | art director | human art director | QC / rights |
| Shot ledger | shot director | shot director | human director | QC |
| Generation candidate | generation production | may not change upstream facts | may not self-approve | QC |
| Release candidate | edit / production | version composition only | release owner | QC / legal |
"May draft" does not mean owning the fact. A market agent may propose an audience hypothesis and may not write an untested hypothesis as a proven conclusion. Generation production may select a candidate and may not change a character's wardrobe fact because one flawed video looked better.
A complete handoff.
When the writers' room finishes E001_PACK@v09, its output is not a chat transcript. It is the episode pack, the
unresolved questions and the asset requests. The showrunner validates schema, IDs and state, sends
ASSET_REQUEST_E001@v03 to the asset director, and sends a script review task to QC. Only when QC shows no P0/P1
and a human creative lead approves does the episode_pack_approved event publish.
The shot director reads the approved version. If the script requires legible text on the authorization while
GRAPHIC_AUTHORIZATION has not been created, it raises a blocking issue rather than improvising one in a prompt.
Once the asset is complete, the original task resumes under the same task ID.
That handoff prevents three kinds of hidden modification: treating a missing asset as creative licence, treating a candidate as approved, and treating downstream convenience as a reason to change the story.
What an approval package shows.
A human approver should not have to read every underlying log. The package contains the current recommendation, the difference from the previous version, affected artifacts, cost, blocking problems, alternatives, and irreversible consequences.
approval_request:
id: APR_STYLEFRAME_004
decision: choose realistic office route A or illustrated route B
evidence: [STYLE_TEST_A, STYLE_TEST_B, COST_TEST_02]
difference: A gives more natural performance; B has a higher identity pass rate
downstream_impact: characters, locations, keyframes for the first ten episodes
deadline: asset_lock_gate
default_if_no_response: remain_blocked
High-impact decisions must never default to automatic approval.
55.4 Privilege testing, human load, and task packets#
Privilege escalation tests.
Before going live, deliberately supply dangerous inputs: ask the shot agent to create a new character; ask QC to edit the script so it passes; ask the generation agent to bypass an unapproved keyframe; ask the distribution agent to remove the AI disclosure. The system should refuse, name the permission violated, and generate an audit event.
Test soft escalation too: instead of modifying anything directly, the model quietly changes a name, a costume or an amount inside its output. The reference validator must compare against the fact ledger, not merely check that the JSON is valid.
A budget for human load.
Human approval is also a bottleneck. Review high-risk decisions item by item and sample low-risk format conversions; present stable, similar tasks in batches showing differences. The system records average approval time and backlog, so automation output does not exceed human judgment capacity.
When the approval queue grows too long, the scheduler reduces upstream concurrency rather than continuing to manufacture candidates. Production speed is set by the slowest necessary judgment.
RACI and on-call responsibility.
Each gate names who is responsible, accountable, consulted and informed. Agents may be responsible for preparing material; final accountability must land on a named person. When an AI disclosure is missing on release night, "the legal agent said it was fine" cannot be where responsibility ends; the release owner signs against the rights and QC evidence.
High-risk stations have human on-call and escalation deadlines: P0 immediately suspends the related release, P1 is handled before the current batch continues, P2 enters the weekly review. Agents trigger; they do not change severity themselves.
A station receives a task packet, not a sentence of natural language.
The minimum task packet issued by the showrunner contains a task ID, the target artifact, approved inputs, permitted variables, prohibited variables, budget, deadline, output schema, acceptance rules and the escalation target. Natural language explains intent; it does not replace those fields.
task_packet:
task_id: TASK_STORYBOARD_E003_B06
owner: AGENT_STORYBOARD
objective: break the father's signature reveal into generatable shots
approved_inputs: [E003_PACK@v11, CONT_E003_B06@v04]
mutable: [shot_count, shot_size, camera_zone, coverage]
immutable: [signature_fact, character_positions, prop_owner]
budget: {drafts: 2, human_review_minutes: 12}
output: storyboard.schema@v06
stop_when: [missing_asset, fact_conflict, budget_exceeded]
escalate_to: SHOWRUNNER_QUEUE
Without immutable, the shot agent may change the evidence to make a shot look better. Without stop_when, it
invents new facts to fill gaps. Without output, downstream can only parse prose.
Submissions divide into candidate, recommendation and decision.
A candidate is an artifact available for comparison — three keyframes, say. A recommendation is the agent proposing one candidate under the rules, with evidence and risks. A decision is a state change produced by a person or a gate holding approval authority. All three use different event types.
A generation agent may submit candidate_created and candidate_recommended, and may not submit asset_approved.
QC may submit gate_blocked, and may not quietly edit a shot and then declare it passing. Nor may the showrunner
mark a missing licence as cleared while the rights owner is absent.
The distinction prevents a confident model tone from being mistaken by the system for a formal decision. Every approval event verifies the approver's authority and the hash of the candidate it relies on.
55.5 Autonomy levels, capability tokens, and approval quality#
Autonomy is not a global switch.
One system can run different autonomy levels by task risk. L0 is read-only analysis. L1 drafts and a human approves item by item. L2 executes within approved rules with sampling. L3 advances low-risk tasks automatically and stops at thresholds. L4 covers only fully deterministic format and file operations. There is no single level called "fully automatic microdrama."
Casting, changes of fact, golden assets, story cut points, release and rights always retain human accountability.
Transcoding, schema checks, reference resolution and rendering approved templates can be highly automated. Autonomy
binds to task_type × risk_tier rather than to an agent's name.
One safe failure is worth more than one impressive overreach.
Suppose the asset agent finds the script mentions an old photograph of the mother while the registry contains no such prop. The safe behavior is to create an asset request and block the related shots. The dangerous behavior is generating a photograph that looks about right. The latter is faster and may invent the mother's age, location, clothing and relational evidence, contaminating a whole season.
Track a correct stop rate: when facing missing data, conflicts and out-of-scope requests, does the agent refuse and supply minimum sufficient evidence? A production-grade agent's capability includes knowing which tasks it is not currently permitted to complete.
Capability tokens and per-task authorization.
Permissions cannot live only in a system prompt. The scheduler issues a short-lived capability token per task packet, declaring the entities it may read, the artifacts it may create, the tools it may call, its budget and an expiry. When the agent submits results, the system validates the token against the operation's scope. However confidently a model claims approval, without a valid token it cannot write formal state.

Figure 55-2 The agent may read approved inputs and write into the candidate area, while the red boundary blocks it from executing approvals and writing formal state. Candidates must pass an independent approval gate; a model's confident tone is not a substitute for permission.
The formal state store on the right is not open to agents. Even when an agent produces both a candidate and a validation report, it cannot approve itself; high-risk stations keep proposal, evaluation and decision separated.
capability_token:
subject: AGENT_STORYBOARD
task: TASK_E003_B06
read: [EP_E003@v11, CHAR_LINYUN@v03, LOC_MEETING@v05]
create: [storyboard_candidate]
forbidden: [approve_asset, alter_fact, publish]
budget_units: 12
expires_at: 2026-07-06T18:00:00Z
Tokens are minimized per task; a "director agent" never receives permanent write access to the whole library. Emergency expansion is approved by a human with a recorded reason, and revoked automatically on expiry.
A service catalog of stations, and input/output compatibility.
Beyond names and responsibilities, the agent registry declares supported task types, input schema, output schema, model and tool dependencies, SLOs, cost tier and known exclusions. The scheduler matches stations by capability rather than sending everything to one general chat model.
When schema versions are incompatible, the task enters an adaptation or human queue; a model must not guess fields. A new station first runs in shadow mode, reading real tasks without writing production results, and is compared against the existing station on pass rate, human edit volume, cost and safe-stop rate before taking a small share of traffic.
Human approval needs quality standards too.
People are not automatically reliable. Approvers get tired, favor their own proposals, skip evidence, or apply different thresholds across batches. An approval contract specifies the questions asked, the evidence that must be seen, the available actions, conflicts of interest, and the maximum continuous review load. Decisions on key casting, rights and release use two people or separated authority.
The system records approval time, reopen rate and escaped problems, and uses them to improve the interface and the division of labor — not to judge individuals crudely. If one class of candidate always forces the approver to investigate afresh, the approval package is incomplete. If approvals are frequently overturned downstream, the standard needs calibrating or the training samples need extending.
A note on sources#
Naming a model after a job title does not create a production system. What makes agents useful is bounded permission, verifiable handoffs and human accountability at the decisions that matter.