Production operations

Production Operations

A company operating model for accountable AI services: decision rights, shared rails, workflow ownership, proof gates, telemetry, cost, recovery, and retirement.

Use these artifacts after design begins—not only after launch. They define how a production-agent workflow is promoted, observed, contained, recovered, changed, reviewed, and retired.

Need Start with Output
Establish a company adoption operating model Operate and Scale and workflow portfolio review Named decision rights, shared-versus-workflow capability boundary, temporary delivery substitutions, and proof-to-operation gates
Decide whether to promote Release gates Evidence-backed hold, shadow, canary, bounded-production, or autonomy decision
Define runtime evidence Telemetry contract Trace, effect, version, identity, outcome, and cost events
Set service targets SLO scorecard Segment-specific objectives, budgets, capacity, and recovery policy
Operate data readiness Data quality and drift Source, quality, lineage, correction, drift, rebaseline, and retirement decisions
Contain and recover Incident runbook Scoped containment, reconciliation, regression, ownership, and re-enable decision
Change behavior safely Change management Compatible release bundle, evaluation report, canary, rollback, and post-change decision
Keep dependency views useful Map freshness and change impact Derived maps with source provenance, freshness, impact review, and no shadow authority
Detect dangerous divergence Behavior monitoring Independent intent/action signal routed to trusted containment controls
Govern capability provenance Capability supply chain Admitted, pinned, constrained, monitored, and revocable tools, MCP servers, skills, CLIs, and code packages
Measure adoption and transfer ownership Delivery and adoption plan and customer handoff Predeclared adoption contract, exercised harness ownership, and artifact lineage
Review production value and service health Production service review Expand, continue, constrain, pause, improve, or retire decision
Review an applied-AI workflow portfolio Workflow portfolio review Cohort-aware investment, continuation, operating ownership, productization, transfer, capacity, or exit decision
Route field learning Field-learning register Confidentiality-reviewed customer configuration, product backlog, reusable artifact, or retirement input
Improve or retire Operate and Scale Gated compatible release or verified decommission

Start adoption instrumentation, artifact-lineage capture, and customer harness pairing during the pilot. The recurring cadence and improve/retire sequence are in Operate and Scale.