Part XVI — Engineering Implementation and Production Infrastructure
Chapter 89. Security and Rights Red Team: Attack Your Own System Before Release#
In this chapter
An AI microdrama system simultaneously handles unreleased scripts, character identities, voices, client materials, licensing evidence, vendor keys and platform publishing rights. A security red team is not a luxury reserved for large companies; a small team should also rehearse the paths most likely to cause irreversible loss.
89.1 Threat modelling, privilege domains, and prompt injection#
The threat model.
Assets include story secrets, real people's likeness and voice prints, synthetic characters, licence documents, generation inputs, release masters, accounts and budget. Attackers may be outsiders, members exceeding their authority, malicious uploads, poisoned vendor responses, or automation misfiring.
threat:
id: threat_prompt_injection_research
entry: external_research_post
target: rights_documents
path: untrusted_text_to_agent_instruction
impact: confidential_upload
controls: [content_taint, no_network_write, capability_token, egress_review]
test_case: redteam_014
Write the threat model by asset, entry point, path, impact and control. Do not end it with "be careful with data security."
Least privilege and domain separation.
A research agent can read public material and not licence originals. A generation worker can read cropped references and not the full script. The editing system can read approved media and cannot reach billing. The release service can consume only a signed release manifest. Production, test and personal experiment accounts stay separate.
Keys live in a secret manager and are issued short-lived. They are never written into prompts, logs or project files. Departures and project endings trigger permission revocation and an access audit.
Prompt injection testing.
Plant malicious instructions in research pages, subtitle files, image metadata and vendor documentation. The system should treat them as content data and never promote them to system instructions. The context compiler tags provenance, and the agent contract states explicitly that external content carries no authority.
A prompt telling the model to "ignore malicious instructions" is not enough. The real control is that the agent has no permission to read the rights library, send files out, or publish — so even a deceived model cannot complete the attack chain.
89.2 Supply chain, likeness and voice, and content safety#
Files and the supply chain.
Uploaded files are checked for type, size, magic number, macros, compression bombs, path traversal and malicious payloads; image metadata is stripped; archives are expanded in a quarantine area. Third-party fonts, plugins, models and scripts record provenance, version and licence.
Build dependencies are version-locked and produce a software bill of materials. Urgent upgrades are tested in an isolated project first, never applied automatically to a project mid-delivery.
Likeness, voice, and consent.
Material from real people is managed with explicit purpose, territory, platform, term, whether training is permitted, whether synthetic transformation is permitted, and a withdrawal mechanism. Licensing evidence is linked to the media, but access to originals is restricted. Synthetic voices also record their provenance and terms of service — "AI generated" is not an assumption of zero rights risk.
A revocation test simulates the termination of one voice licence: the system should find unreleased dialogue, published platform packages, campaign derivatives and backups, stop new use and produce a disposition list.
Content and brand safety.
The red team attempts to generate discriminatory content, misleading medical claims, impersonation of real people, material creating risk for minors, gratuitous violence and brand-prohibited content. Rules and human review are applied by project, platform and territory. Refusals preserve the minimum necessary evidence, avoiding repeated propagation of sensitive content through logs.
The complexity of storytelling must not be caught by crude keyword matching. The automated system identifies risk and context first; boundary cases are judged by an authorized human who records the basis for the decision.
89.3 Release protection, incident response, and evidence retention#
The release attack surface.
Simulate an old version being uploaded, unapproved material being substituted, a cover art and episode mismatch, a missing subtitle track, a platform token leak and a scheduled release in the wrong time zone. The release service validates an immutable manifest, checksums, two-person approval and the destination platform.
Platform credentials are not held by generation or editing agents. Publishing uses short-lived permissions, and on completion records the platform receipt and the remote content hash.
Incident response.
On discovering a leak, first stop the spread, revoke tokens, freeze related tasks and preserve evidence — then assess impact and notify. Do not begin by deleting logs, and do not treat "we pulled the file" as the end of it. The runbook defines the security owner, legal, client communication and the platform contact.
The retrospective examines every control point along the attack chain: where it could have been blocked, why it was not, how long detection took and how long recovery took. Fixes go into tests, not just into a policy document.
When rights and security retention conflict.
Auditing requires retention; privacy and contracts may require deletion. Use a layered policy: delete the original identifiable media and sensitive fields, while retaining irreversible checksums, decision types and necessary legal proof. Retention periods are set by purpose and legal advice — not kept forever "just in case".
89.4 Security acceptance and a tabletop exercise#
Fault tree.
A model leaking future plot: context minimization failed. A malicious web page triggering an outbound transfer: content and instructions were not separated, and privileges were too broad. A licence revocation that cannot find every asset: rights never entered the dependency graph. An old version being published: the manifest and checksums were not enforced. Keys appearing in logs: secrets were passed as ordinary parameters.
SOP, checklist and deliverables.
Build the threat model. Separate assets into domains. Enforce least privilege. Isolate uploads. Test prompt injection. Rehearse licence revocation. Attack the release chain. Verify incident response. Remediate and add regressions. Review terms and dependencies on a schedule.
- External content is untrusted by default.
- A deceived model still cannot obtain dangerous permissions.
- A licence revocation can query every derivative.
- Release uses an immutable manifest and a two-person gate.
- Security incidents have evidence preservation and a notification path.
Exercise: run ten red team cases against the pilot project, delivering threat_model.yaml, permission_matrix.csv,
redteam_cases/, rights_revocation_report.md, incident_runbook.md and remediation_tests/.
The tabletop exercise.
Simulate a client discovering that one campaign asset used an unapproved voice. Within ten minutes the team must answer: which platforms carry it, which task derived it, who approved it, what the licence gap is, how to stop it, who must be notified, and whether a replacement version is ready. The facilitator keeps adding information — the platform cache is still reachable, another language version is also live — to test the dependency queries and the communication chain.
The exercise does not end by rating how quickly individuals reacted. It ends by finding gaps in the system. If the release list lives in someone's personal spreadsheet, if legal cannot find the evidence, or if operations does not know where the rollback entry point is, each becomes an improvement to asset relationships, rights snapshots and the runbook respectively, with a retest date scheduled. Security maturity comes from repeatable controls, not from one person's good memory.
A note on sources#
This chapter describes a rehearsal method rather than a specific security stack. What transfers is modelling threats by asset and path, keeping capability out of the model's reach so deception alone cannot complete an attack, and treating rights revocation as a queryable dependency problem.