中
Chapter 113. Agentic Ops Practicum: Fail Safely, and Let Someone Else Recover

Part XIX — Practicums and the Capstone

Chapter 113. Agentic Ops Practicum: Fail Safely, and Let Someone Else Recover#

In this chapter
113.1 The ops boundary, inspection, and staged releases113.2 Permissions, idempotency, and budget backpressure113.3 Dead letters, human queues, and acceptance113.4 A night shift, queue drills, handover, and the retrospectiveA note on sources

The operations role maintains the agent registry, workflows, permissions, budget, events, evaluators and recovery. The goal is not a model that never errs. It is that errors stay contained, evidence is preserved, tasks can be recovered, and irreversible decisions still belong to the right people.

113.1 The ops boundary, inspection, and staged releases#

The boundary.

Operations can deploy workflows, roll back prompts, open and close circuit breakers, handle dead letters and restore state. It cannot approve story, rights or release. Inputs are versioned components and policy; outputs are a healthy system, traces, events, SLOs and incident evidence.

Production permissions are granted short-lived, and the test environment has neither real publishing nor real budget.

Daily inspection.

Check queue backlog, dead letters, circuit breakers, budget anomalies, stale context, evaluator drift, human waiting and the release queue. Watch trends, not only red lights.

API success rate and production approval rate are separate. A vendor at 99 percent success with a falling candidate approval rate still needs investigating.

Staged releases.

Prompts, adapters, rules and evaluators all have tests, staged rollout and rollback. The ops lead deploys a new prompt version to 10 percent of traffic and compares schema failures, rejection rate, cost and error types.

An escaped blocker triggers automatic rollback. Already-generated candidates are stamped with their version; history is not reinterpreted.

113.2 Permissions, idempotency, and budget backpressure#

Permission red teaming.

Have the shot agent try to modify the locked script, exceed the budget on a call, and publish. The resource layer must refuse and record a security event. The word "urgent" in a prompt does not widen a token.

Inject upload instructions into external research text: the agent cannot complete the attack because it holds no outbound permission.

Idempotency and duplication.

Deliver an approval event twice, and confirm it creates one downstream task, one budget settlement and one vendor call. Crash a worker after the call and before the record is written, and have the ops lead recover using the call intent and a query for the remote job.

Manually deleting duplicate records does not count as a pass; the business invariant has to be fixed.

Budget and backpressure.

When the generation queue exceeds human review capacity, the workflow slows down instead of fanning out further. When budget runs out, partial candidates are preserved and an escalation package is generated — no silent de-escalation.

The ops lead adjusts concurrency and priority so release blockers and low-risk previsualization do not compete for one queue.

113.3 Dead letters, human queues, and acceptance#

Dead letters and recovery.

Handle a task that entered the dead letter queue on a schema conflict: inspect the original input, fix upstream, create a new task and link the old record. Editing the dead letter payload and replaying it is not allowed.

Delete a materialized view and rebuild it from events and snapshots, comparing task, budget and asset state afterwards.

The human queue.

Approval is a formal node with an SLO, evidence and escalation. Candidates can be compared in batches, while rights and release are confirmed item by item. A congested queue triggers upstream backpressure and a notification to production.

Operations does not click approve on an approver's behalf.

Acceptance and deliverables.

Observability 15 percent; security and permissions 15 percent; staged releases 15 percent; idempotency and recovery 20 percent; budget backpressure 15 percent; dead letter recovery 10 percent; runbook handover 10 percent.

  • Irreversible actions have a technical gate.
  • Duplicate events do not spend twice.
  • Components can be staged and rolled back.
  • Budget exhaustion stops safely.
  • Someone who did not build it can recover from the runbook.

Deliverables: ops_dashboard_spec.md, canary_report.csv, security_events.jsonl, idempotency_tests/, dead_letter_report.md, recovery_drill.md and oncall_runbook.md.

113.4 A night shift, queue drills, handover, and the retrospective#

Simulated night shift: a duplicate event and a budget anomaly.

At 01:14, monitoring shows the same shot billed twice by the vendor. The on-call ops lead first pauses new fan-out for that task type and follows the correlation ID through the events: shot.approved was delivered twice, the task table created only one row, but a worker crashed and restarted before its external call intent had been committed, and called again.

The ops lead queries both remote jobs, keeps the earlier result, marks the other as a duplicate charge and notifies production. The fix persists the call intent before calling, uses a stable idempotency key, and recovers remote jobs on restart. They then inject "crash immediately after a successful call" in the test environment and confirm only one external job results. Traffic reopens at 5 percent.

At 02:06, a retired identity asset reaches a prompt from cache. The ops lead traces the asset promotion event and finds the cache invalidation consumer in the dead letter queue. Every affected candidate is quarantined, the cache key gains an identity version checksum, and the dead letter fix creates new tasks rather than editing old events.

Backpressure on the human queue.

By morning the visual approval backlog is 83 items. The ops lead reduces concurrency for low-risk generation, raises the priority of principal close-ups and release blockers, and notifies production of the capacity change. They do not let the agents auto-approve low-scoring candidates.

The runbook handover test.

A second ops lead replays the incident from the runbook: finding the trace from the alert, judging the safe state, tripping the breaker, recovering and verifying. Any step that requires the original on-call to explain it becomes documentation or an automated query.

The facilitator's note: operations does not take pride in "the service stayed up." Duplicate charges, contamination by a retired asset or a permission overreach are production incidents even with no downtime. A sound system is safe, explicable and recoverable, and it respects the boundary of human decisions.

The live retrospective.

The facilitator supplies a distributed trace running from asset approval to a duplicate charge. The ops lead marks the invariants between event, task, call intent, vendor job and settlement, and proposes the crash point and the recovery algorithm. "Add a database lock" is not an answer on its own; they must say what the lock protects, and how lease expiry and duplicate messages are handled.

The second question asks for a minimum capability token for the shot agent: read the current scene pack, write candidates, make three low-resolution model calls — and neither change locked dialogue, read licence originals, nor publish. The third supplies an approval backlog and asks for a priority and backpressure design, with automatic loosening of thresholds prohibited.

Finally, someone else performs a snapshot restore from the ops lead's runbook. If a step depends on an implicit environment variable, a personal key or the author's memory, the handover item fails. Production evidence includes traces, policy versions, drill logs, recovery comparisons and rollbacks. Operational capability is proven by business consistency under failure — not by green dashboards on a normal day.

After recovery, run an invariant audit confirming budget, task, asset and approval counts match the pre-failure state. A service that starts while the business ledgers disagree is still a failed recovery.

A note on sources#

This is a rehearsal design rather than a specific platform's runbook. What transfers is separating API health from production health, proving idempotency by crashing on purpose, applying backpressure instead of fanning out, and testing recovery through someone who did not build the system.