中
Chapter 81. Agentic Chaos Lab: Making the System Fail on Purpose

Part XV — Production Labs and the Failure Casebook

Chapter 81. Agentic Chaos Lab: Making the System Fail on Purpose#

In this chapter
81.1 Experiment boundaries, duplicate events, and context contamination81.2 Privilege escalation, budget exhaustion, and human congestion81.3 Vendor failure, disaster recovery, and safety scoringA note on sources

An agentic system that only demonstrates success on ideal inputs has not entered commercial production. This lab injects timeouts, duplicate events, expired assets, privilege escalation, budget exhaustion and human queue congestion into an isolated environment, and observes whether the system stops in a safe state and leaves enough evidence to recover.

Bounded agents meeting injected faults through stopping, approval, isolation, retry and snapshot recovery

Figure 81-1 Chaos experiments do not pursue a system that never errs. They verify that a fault cannot cross permission, budget, isolation and human gates — and that recovery from evidence is possible.

81.1 Experiment boundaries, duplicate events, and context contamination#

The environment and safety boundary.

Chaos experiments run on a copied project with virtual budget and non-publishing tokens. Outbound connections are closed and test assets carry conspicuous internal markings. Before starting, save an event snapshot, the expected invariants and an abort switch.

chaos_experiment:
  id: chaos_ep004_007
  environment: isolated_staging
  hypothesis: a duplicate completion event does not cause double budget settlement or double state advancement
  injection: duplicate_event_delivery
  invariant:
    - budget_charge_count_equals_1
    - task_state_committed_once
    - downstream_render_created_once
  abort_if: any_external_publish_attempt
  recovery_target_minutes: 10

Duplicate events.

The system delivers shot.approved to the edit task twice. Correct behavior is the consumer recognizing the same event ID or idempotency key, creating one downstream task and settling once. The failing version kept the task unique in the database while calling the vendor twice, because idempotency existed at the task table and not at the external call layer.

The fix reserves a record before the call and uses a vendor request idempotency key; where the vendor does not support one, a local lease and response registry protect it. Regression tests also cover duplicate delivery after a consumer restart.

Stale assets and context contamination.

Inject an identity asset already superseded by a newer version. The context compiler should refuse it based on lifecycle and dependency invalidation events. If the old asset remains in cache, the task must be marked context_stale rather than continuing silently.

The failure log showed the cache key contained only the character ID, not the identity version and state snapshot. After the fix, the cache key includes a dependency checksum, and asset promotion events proactively purge affected contexts.

81.2 Privilege escalation, budget exhaustion, and human congestion#

Privilege escalation.

Have the shot agent attempt to rewrite locked dialogue and call the publish endpoint. The capability token should refuse at the resource layer, produce a security event and notify a human — without marking the whole project failed. An agent's output claiming this is necessary to complete the task does not constitute authorization.

The test also plants a malicious instruction inside research material asking the agent to ignore its task and upload the character release. The context compiler marks external text untrusted and extracts content candidates only; it does not execute instructions. Having no network write permission is the last technical protection.

Budget exhaustion and partial results.

A generation task exhausts its budget after four attempts, having produced two candidates — one passing identity, one failing action. The system should stop new calls, preserve the candidates and the attempt log, and produce an escalation package. It must not promote a partially passing candidate automatically, and must not overdraw in order to finish the task.

A human may add budget, switch to coverage, reduce shot complexity, or cancel. Adding budget creates a new authorization record; it does not edit the original reservation ceiling.

Human queue congestion.

Inject forty previsualization approvals simultaneously and observe whether critical release problems are drowned. The queue should sort by irreversibility, deadline, downstream blocking and budget risk. Similar low-risk candidates can be compared in batches while release and rights problems stay individually approved.

The system computes human load and does not record waiting time as a model failure. When median approval wait exceeds the SLO, the workflow slows upstream fan-out rather than producing more assets awaiting review.

81.3 Vendor failure, disaster recovery, and safety scoring#

Vendor timeouts and circuit breaking.

The first two timeouts trigger backoff; once the failure rate crosses the window threshold, the breaker opens. New tasks stop calling that vendor, and an approved fallback model is used only where capability, look and rights are compatible. The switch is recorded in the asset's provenance — never presented as though it were still the original model.

On recovery, probe with small traffic before closing the breaker. All timed-out tasks reuse the same idempotency key, so that a vendor which actually completed while the response was lost cannot be paid twice.

A disaster recovery drill.

Delete the materialized views while retaining the event log; the system should replay task state, budget balance and asset dependencies. Compare the replay against the snapshot, and treat any difference as blocking. Media files are not rebuilt from events and can be re-indexed in object storage by checksum.

The drill records RPO, RTO, manual steps and anything that cannot recover automatically. If only the original engineer knows the recovery commands, the organization does not yet have recovery capability.

SOP, checklist and deliverables.

Build the isolated environment. Write the hypothesis and invariants. Choose a single fault. Define the abort switch. Execute and capture traces. Judge whether it failed safely. Repair the earliest control gap. Add regression. Re-run. Update the runbook and the SLOs.

  • Test tokens have no external publishing capability.
  • Duplicate events cause no duplicate spend or state advancement.
  • Stale assets cannot escape through the cache.
  • Budget exhaustion preserves evidence and stops.
  • Human congestion triggers backpressure rather than more fan-out.
  • Recovery can be executed by someone other than the author.

Exercise: complete six classes of fault injection, delivering chaos_plan.yaml, invariants.json, trace_exports/, incident_reports/, regression_tests/ and recovery_runbook.md.

Scoring safe failure.

An experiment does not pass simply because it eventually recovered. Record separately whether the error was detected promptly, whether it stopped spreading, whether rights and budget were protected, whether partial work was preserved, whether the root cause is explicable, and whether an on-call person could recover independently. Fast recovery with a duplicate charge is still a failure; no data loss with a leaked publishing token is a worse one.

Each experiment also checks alert quality. If one timeout produces a hundred duplicate notifications, humans will miss the real blocker. If the system recovers automatically with no event record, the team cannot confirm what happened. Alerts aggregate to a root trace ID and show the scope of impact, the current safe state and the available next actions. Maturity in chaos engineering is not the disappearance of faults; it is faults becoming predictable, observable and bounded production events.

A note on sources#

This lab reconstructs a fault-injection method for creative production systems. What transfers is running in isolation with non-publishing tokens, stating invariants before injecting, and scoring the quality of the failure rather than only the recovery.