Part XI — The Agentic Production System
Chapter 57. Prompts, Context, and Version Control#
In this chapter
57.1 The prompt execution layer, context compilation, and versions#
A prompt is only the execution layer.
Reliable output comes from approved facts, relevant state, a task objective, a schema and model capability acting together. A prompt cannot substitute for assets and ledgers.
Four context tiers.
System rules define principles that may not be violated. Project context supplies the style and series bibles. Task context contains only the current episode, shot and state. The output contract specifies the schema and acceptance.
Tiering lets you update one shot without resending a season's material, and it stops future secrets contaminating a character's present knowledge.
Context compilation.
The compiler queries dependencies by task ID, selects the currently approved versions, removes irrelevant fields, and detects conflicts or gaps. The context packet stores a hash so results can be reproduced.
context_packet:
task: TASK_STORYBOARD_E001_B04
system_rules: SYS_STORYBOARD_v5
project_refs: [STYLE_MAIN@v03, SERIES_BIBLE@v07]
task_refs: [E001_PACK@v09, B04, CH_LINYUN@v03]
output_schema: storyboard.schema.json@hash
Prompt versions.
Every prompt change carries a version, a reason, a test set and a measured effect. Do not edit one sentence in production while keeping the same version number. Results record prompt ID, model, parameters and the context hash.
Retrieval and caching.
Retrieve only approved assets; candidates and rejects are quarantined by default. Identical inputs and parameters can be cached to avoid repeated spend. The cache key includes every version that could affect the result.
57.2 Structured output, retries, overload, and evaluation#
Structured output.
The model emits JSON or YAML first, and a validator checks required fields, IDs and references. A format-repair agent may fix syntax only; it may never change a story fact.
Idempotency and retries.
Running the same task twice must not create several conflicting approved results. Each attempt has a unique ID; a retry inherits the task with an explicit reason for the change. Service timeouts may be retried; content failures are diagnosed first.
Context overload.
Too much material dilutes the critical constraints. A writing task needs only what the character currently knows; a visual generation needs only the assets currently visible. Keep provenance links for expansion when required.
Evaluation.
Compare prompt updates on a fixed task set for format, facts, continuity, cost and human edit volume. One attractive result does not justify an upgrade.
SOP.
First, separate the four context tiers. Second, compile the minimum packet using the dependency graph. Third, store versions and hashes. Fourth, require structured output. Fifth, validate references. Sixth, cache identical tasks. Seventh, distinguish technical retries from content rework. Eighth, evaluate prompt updates on a regression set.
57.3 A worked example, assembly order, and two-stage generation#
Fault tree.
Symptom: the model writes a future secret. The context contains a whole season the task has no permission to see. Trim by character knowledge.
Symptom: the same task cannot be reproduced. Prompt, parameters or context were not versioned. Record the complete execution packet.
Symptom: the format is right and the facts are wrong. Only schema validation ran, with no reference validation. Connect it to the fact ledger.
Checklist, exercises and deliverables.
Check that the four tiers are separated; that only approved, relevant content is loaded; that versions and hashes are stored; that output is validated; that caches invalidate correctly; that retries carry reasons; and that prompt updates are regressed.
Exercise one: compile the minimum context for E001_S09. Exercise two: design the cache key. Exercise three: compare human edit volume across two prompt versions.
Deliverables for this chapter: the context compiler spec, the prompt registry, context packets, validation results, the cache policy, and the prompt regression report.
Compiling context for E001_S09.
Generating the authorization insert does not require the season's romance line or the mother's secret. The compiler
reads the shot row and follows dependencies to STYLE_MAIN@v03, LOC_BOARDROOM@v04, PROP_AUTHORIZATION@v03,
GRAPHIC_AUTHORIZATION@v05, the lighting state, the previous shot's action handoff, and the platform safe areas.
It explicitly excludes the later birth records, the E006 blackout state, Lin Yun's subsequent black suit, unapproved model candidates, and the character's psychological biography. Both the exclusion list and the inclusion list enter the log, so any future information leak can be investigated.
Prompt assembly order.
system hard rules
-> output schema and prohibited escalation
-> project style and negative constraints
-> currently approved asset versions
-> shot state, composition, action and audio intent
-> risks specific to this task
-> self-check requirements
The order keeps high-priority constraints from being buried at the end of a long text. Free creative fields sit only inside the approved mutable range — background extra posture, say, rather than character identity or an amount.
Two-stage generation.
For complex structured tasks, have the model output a plan and a reference list first, validate it, then produce the full result. A shot agent submits shot count, asset references and state changes first; discovering an unregistered prop blocks immediately, rather than generating twenty-five wrong storyboard rows.
The second stage's result then passes three validations: schema format, reference existence, and business invariants. Format repair may add a quote or fix an array structure; it may never change a non-existent asset ID into the "closest" one.
57.4 Cache invalidation, regression, safety, and observability#
Cache invalidation.
The cache key includes prompt, model, parameters, context and output schema hashes. When the grey suit version changes, invalidate only tasks referencing that costume. A change to a music licence should not regenerate static images and should invalidate the release gate.
A cached result found to be wrong is marked poisoned, and every downstream reference enters review. Deleting the local file while leaving the approval state intact is not enough.
The prompt regression set.
Fix the contents: six beats from E001, a two-person shot/reverse, a document insert, a low-light blackout, long narration, a cut point, and a high-risk contact. Compare field completeness, factual errors, characters of human editing, continuity issues, cost and elapsed time.
If a new prompt produces more attractive prose while raising asset hallucination from 1 percent to 8 percent, do not promote it. Different task types can run different prompt versions; there is no need for one universal template.
Prompt injection and untrusted text.
External novels, comments and client documents are data and must never be executed as system instructions. The context compiler tags each source with a trust level and strips text asking to reveal keys, bypass compliance or change permissions. Agents obey system rules and approved tasks only.
Observability.
Every call records usage, latency, cache hits, validation errors, human edit volume and whether the result was finally adopted. A route that is expensive and never adopted should be retired; slightly slower output that substantially reduces human editing may suit production better.
Logs must not retain unnecessary personal information, keys or restricted full text. Sensitive fields use references and access control, and debug exports follow the same rights and data policy.
A seven-step context compilation pipeline.
The compiler resolves the task ID, queries approved dependencies, filters by station permission, trims by character knowledge, compresses to the token budget, checks for conflicts and gaps, and finally emits an immutable packet with a hash. A failure at any step means the model is not called.
Compression must not rewrite facts into vague summaries. Stable IDs, numbers, negations, states and prohibitions keep their original fields; a long character biography can compress to the motivations relevant to this scene. When the context budget is short, delete low-relevance explanation before deleting acceptance criteria.
57.5 Knowledge firewalls, output validation, and contamination isolation#
The character knowledge firewall.
Writing and performance tasks are trimmed not only by episode but by each character's knowledge_state. Lin Yun in
E003 does not yet know her mother is alive, so her dialogue and performance context cannot contain the E006 truth.
The antagonist may know more, and her visible behavior is still constrained by her current concealment goal.
The system keeps three tables: what the author knows, what the character knows, what the audience knows. When a model uses a fact a character has no permission to know, the validator raises a knowledge-boundary issue. That prevents unintentional spoilers introduced by a whole-season bible.
Output validation has four gates.
The first validates syntax and schema. The second validates that IDs, versions and references exist. The third validates business invariants — prop ownership, time, character knowledge. The fourth is a semantic review of whether the output met the task objective. Format repair only makes sense once the first two pass.
When a model outputs a non-existent asset, the validator must not fuzzy-match it to the closest ID. It returns an explicit error and, optionally, an asset request. Automatic well-meaning repair easily replaces Lin Yun's old pen with another character's ordinary one.
Prompt operations and staged rollout.
A new prompt is compared on an offline regression set first, then shadow-run on a small number of canary tasks, then rolled out shot class by shot class. The old prompt is retained during the transition, and batches already in production do not migrate as a rule. Metrics include factual error rate, first-pass rate, human edit volume, context cost and latency.
If a new version is better on ordinary dialogue and worse on key evidence, upgrade only the former. The prompt registry records applicable tasks, excluded tasks, model version and rollback conditions; there is no need for one "latest prompt" across the whole system.
Provenance for context facts.
Every fact in a context packet carries its source entity, version, approval event and extraction time. When a model's output references a number, a character state or a licensing conclusion, the validator can trace where it came from. A statement with no provenance can only be a suggestion; it cannot be written into the fact ledger.
context_fact:
key: authorization_amount
value: 320000000
source: PROP_AUTHORIZATION_FACTS@v07
approved_by: DECISION_0182
valid_for: [E001, E002, E003]
trust: authoritative
Provenance also resolves conflict: when a script draft and an approved prop amount disagree, the compiler selects by authority level and reports the conflict rather than handing both values to a model to decide between.
Loss testing for context compression.
Long projects must compress, and compression can delete negations, conditions, time boundaries and exceptions. Build a loss test for the context compiler: answer a set of key queries from both the original facts and the compressed packet, checking what characters know, where props are, which behaviors are forbidden, and whether numbers agree. Losing any blocking fact rejects that compression strategy.
A summary may compress explanation; it may not rewrite stable IDs, proper nouns, numbers, negations or permissions. Vector retrieval is good at recalling candidates and is not responsible for deciding authority; the final context compiles from approved entities only. A highly relevant but unapproved old draft is still a wrong source.
Isolating factual contamination from model output.
Scripts, prompts and analyses produced by a model enter a candidate zone first. Only after schema, fact, permission and human gates may specific fields be promoted. The system forbids automatically extracting new facts from free text and writing them back into core state, because one hallucination then spreads along the dependency graph.
Exploratory tasks may propose new characters, locations or props, and the output type is a proposal carrying its assumptions and cost. After approval, a dedicated event creates the entity. Creativity is preserved and so is the factual boundary; the two do not need to share one write permission.
A note on sources#
Prompts change with models; production semantics do not. Treating the prompt as an execution layer over versioned context and validated output is what makes agent results reproducible and safe to build on.