aBmeSubscribe
CL-010·CL Track·Advanced·10–45 hrs saved

Stop Confusing "It Deploys" with "It's Ready" — Evidence-Based Production Launch for Cloud Workloads

An AI-assisted workflow to run a risk-based production-readiness review, surface launch blockers, classify residual risk, and produce an operational acceptance package before real users arrive.

3Phases
25Prompts
10–45Hours saved
10Deliverables

Executive Brief

Your Challenge

Your workload deploys, the pipeline is green, and the scanner reports no criticals — so the business wants to launch. But none of that proves the workload can support real users, be operated at 2 a.m., recover from a regional outage, or stay inside its budget. Production readiness is an operational-risk decision, not a project milestone, and treating launch as a deadline instead of a governance gate is how workloads enter production with no owner, untested alerts, and unverified backups.

Common Obstacles

The failures cluster in predictable ways. Reviews happen too late, when readiness gaps are expensive and politically hard to fix. Checklists get treated as approval, hiding the fact that a single critical blocker can outweigh a high overall score. Backup gets mistaken for recovery, metrics get mistaken for monitoring, and healthy infrastructure gets mistaken for working business journeys. Underneath all of it: project teams self-approve risk they have no authority to accept, and schedule pressure quietly redefines what "ready" means.

The ABME Approach

This workflow does it in the right order: classify workload criticality, define mandatory versus advisory controls, assemble an evidence register, then assess readiness domain by domain — business, architecture, security, identity, data, reliability, observability, recovery, operations, cost, and compliance. Every finding is classified as a launch blocker, conditional blocker, high-priority risk, improvement, or observation, with named owners and completion dates. The prompt pack drives the assessment; go/no-go criteria, a launch runbook, a rollback plan, a hypercare plan, and formal risk acceptance turn the review into an auditable production decision that a qualified human signs.

Insight Summary

A workload is production ready when the organization can operate it, secure it, observe it, recover it, fund it, and accept its residual risks — not merely when the application can be deployed.
phase-1

Criticality determines control depth. Applying Tier 0 rigor to a Tier 3 service wastes effort; applying Tier 3 rigor to a Tier 0 service is how launches become incidents.

phase-2

Treat missing evidence as unknown, never as passed. Absence of a control record is not proof the control exists — it is a gap in the launch decision.

phase-2

A configured backup is not a recovery capability, a collected metric is not monitoring, and a successful deployment is not a ready service. Each requires validation, not assumption.

phase-3

A single unresolved critical blocker may override a high readiness score — the scorecard supports judgment, it does not replace risk authority.

tactical

Project managers and engineers should not accept enterprise-level business risk; residual risk requires acceptance by the authorized decision-maker, with an expiration date.

The Journey

Three phases; each lists the tools you'll use there.

1

Scope Criticality and Evidence Standards

Classify the workload, name owners, and define what evidence a launch decision requires before assessing anything.
  • Gather the workload's architecture, ownership, and launch context
  • Classify workload criticality using the tier model
  • Separate mandatory controls from advisory improvements
  • Establish the evidence register and evidence standards
  • Confirm risk-acceptance and go/no-go authority
2

Assess Readiness Domain by Domain

Use the prompt pack to evaluate every readiness domain against evidence, classifying each finding.
  • Run the primary prompt with the assembled evidence
  • Assess each readiness domain and classify Ready through Not Ready
  • Separate launch blockers, conditional blockers, and high-priority risks
  • Populate the production-readiness scorecard
  • Flag missing evidence as unknown, not passed
3

Decide, Launch, and Stabilize

Turn the assessment into a go/no-go decision, a launch runbook, a rollback plan, risk acceptance, and hypercare.
  • Define objective go/no-go criteria
  • Produce the launch runbook and rollback plan
  • Record risk acceptance and time-bound exceptions
  • Confirm operational acceptance of the service
  • Run hypercare, validate post-launch, and complete the validation checklist

What's Inside the Execution Layer

Numbered deliverables grouped by phase. Membership unlocks every tool.

1. PHASE 1Checklistprotected

Prerequisites Checklist

Gather every piece of evidence the AI needs to assess production readiness before the first prompt runs.
Use this to
  • Assemble the full evidence base before assessing readiness
  • Surface missing evidence early, before it blocks a launch decision
  • Confirm ownership, cost, and recovery inputs are in hand

Gather as much of the following as possible:

2. PHASE 1Matrixprotected

Evidence Register

A single register that tracks every piece of readiness evidence, its source, owner, validity, and result.
Use this to
  • Record and attribute every piece of readiness evidence
  • Track evidence validity periods and review status
  • Distinguish validated evidence from missing or expired evidence
Evidence IDReadiness domainDescriptionSourceOwnerDateValidity periodResultRelated controlReviewerStatus
3. PHASE 2Matrixprotected

Production Readiness Scorecard

A weighted scorecard summarizing readiness by domain, with the explicit rule that a single critical blocker can override a high score.
Use this to
  • Summarize readiness status across all domains
  • Weight domains by importance for the workload
  • Support the go/no-go decision without letting a score hide a blocker
Reference rows from the blueprint — downloads ship as an empty skeleton
DomainWeightStatus
Business10%Ready
Architecture10%Ready with risk
Security15%Ready
Identity10%Ready
Network8%Ready
Data10%Ready
Reliability10%Conditionally ready
Observability10%Ready
Recovery10%Not ready
Operations7%Ready
RubricA score should support judgment, not replace it. A single critical blocker may override a high overall score.
4. PHASE 1Matrixprotected

Responsibility Matrix

A RACI-style matrix assigning accountability for each readiness capability across the business, service, and operations roles.
Use this to
  • Assign accountability for each readiness capability
  • Clarify who is accountable versus responsible for go/no-go and risk acceptance
  • Prevent project teams from self-approving enterprise risk
Reference rows from the blueprint — downloads ship as an empty skeleton
CapabilityBusiness OwnerService OwnerWorkload TeamPlatform TeamSecurityOperations
Business readinessAccountableConsultedInformedInformedInformedInformed
Technical readinessInformedAccountableResponsibleConsultedConsultedConsulted
Security readinessInformedConsultedResponsibleSupportsAccountableInformed
MonitoringInformedAccountableResponsibleSupportsConsultedResponsible
RecoveryConsultedAccountableResponsibleSupportsConsultedResponsible
Support modelInformedConsultedConsultedConsultedInformedAccountable
Cost readinessAccountable for budgetResponsibleProvides usage inputsProvides platform costsInformedInformed
Go/no-goAccountable for business acceptanceResponsible for service recommendationConsultedConsultedConsultedConsulted
Risk acceptanceAccountable at authorized levelResponsible for documentingConsultedConsultedConsultedConsulted
Operational acceptanceInformedResponsibleSupportsSupportsConsultedAccountable
5. PHASE 2Prompt Packprotected

Production Readiness Prompt Pack

One comprehensive primary prompt and twenty-four targeted follow-ups that drive an evidence-based production-readiness assessment and its supporting artifacts.
Use this to
  • Run a full evidence-based readiness assessment and launch recommendation
  • Generate scorecards, blockers, go/no-go criteria, runbooks, and rollback plans
  • Review individual domains — security, capacity, recovery, cost, and more

Primary Prompt

Start here with as much of the assembled evidence as you have.
You are a senior cloud production-readiness architect, site reliability engineer, cloud security architect, operations leader, application architect, incident-response advisor, disaster recovery specialist, FinOps practitioner, and technology risk advisor.

I will provide some or all of the following:
• Business objectives
• Workload description
• Criticality
• Launch date
• Architecture
• Data flows
• Dependency map
• Service inventory
• Ownership
• Test results
• User acceptance
• Security assessment
• Vulnerability findings
• Penetration-test results
• Identity design
• Network design
• Data classification
• Compliance requirements
• Performance baseline
• Load-test results
• Capacity model
• Service-level objectives
• Monitoring dashboards
• Alert definitions
• Logging design
• Incident-response plan
• Backup configuration
• Restore-test results
• Disaster recovery design
• Recovery-test results
• Deployment pipeline
• Release plan
• Rollback plan
• Change record
• Support model
• Runbooks
• Cost estimate
• Budget approval
• Risk register
• Exception register
• Vendor support
• Communication plan
• Hypercare plan

Your task is to perform a production-readiness assessment and produce an evidence-based launch recommendation.

Do not assume a successful deployment means the workload is production ready.
Do not approve production based solely on checklist completion.

First:

1. Summarize:
   • Business capability
   • Users and customers
   • Workload criticality
   • Launch scope
   • Launch date
   • Service-level expectations
   • Support expectations
   • Regulatory or contractual obligations

2. Separate:
   • Confirmed facts
   • Validated evidence
   • Reported observations
   • Inferences
   • Assumptions
   • Unknowns

3. Identify missing evidence that materially affects the production decision.

4. Identify:
   • Business owner
   • Service owner
   • Technical owner
   • Data owner
   • Security owner
   • Support owner
   • Budget owner
   • Risk-acceptance authority

5. Assess readiness across:
   • Business
   • Ownership
   • Architecture
   • Dependencies
   • Identity
   • Network
   • DNS
   • Certificates
   • Security
   • Vulnerabilities
   • Secrets
   • Encryption
   • Data
   • Database
   • Functional testing
   • Integration testing
   • Performance
   • Capacity
   • Reliability
   • Observability
   • Logging
   • Alerts
   • Incident response
   • Backup
   • Recovery
   • Disaster recovery
   • Deployment
   • Rollback
   • Operations
   • Support
   • Change management
   • Cost
   • Compliance
   • Documentation
   • User readiness
   • Hypercare

6. Classify each domain as:
   • Ready
   • Ready with accepted risk
   • Conditionally ready
   • Not ready
   • Unable to assess
   • Not applicable

7. Separate:
   • Launch blockers
   • Conditional blockers
   • High-priority risks
   • Improvements
   • Observations

8. For each finding provide:
   • Finding ID
   • Domain
   • Classification
   • Severity
   • Confidence
   • Evidence
   • Business impact
   • Technical impact
   • Security impact
   • Operational impact
   • Recommended action
   • Owner
   • Required completion date
   • Validation
   • Launch impact
   • Risk approver where applicable

9. Evaluate whether:
   • Critical business journeys pass
   • Critical dependencies are validated
   • Capacity meets baseline and peak demand
   • Scaling limits are understood
   • Monitoring covers service health and business outcomes
   • Alerts route to named responders
   • Backup is configured
   • Restore has been tested
   • Recovery objectives are approved
   • Recovery has been tested
   • Incident response is operational
   • Support has accepted the service
   • Rollback is practical
   • Cost is approved and attributable
   • Compliance obligations are satisfied
   • Residual risks are properly accepted

Requirements:
• A single critical blocker may override a high readiness score.
• Do not treat missing evidence as evidence of readiness.
• Do not treat configured backup as successful recovery.
• Do not treat scanner output as complete security assurance.
• Do not treat architecture documentation as proof of implementation.
• Do not recommend launch solely because of schedule pressure.
• Require named owners for all launch conditions.
• Require expiration dates for exceptions.
• Identify irreversible changes.
• Define rollback triggers and decision authority.
• Define objective go/no-go criteria.
• Define hypercare and exit criteria.
• Define post-launch validation.
• Identify where human testing, security review, legal review, compliance approval, recovery testing, capacity testing, or business acceptance is still required.
• State when evidence is insufficient.
• Require qualified human approval for production launch.

Then produce:

1. Executive summary.
2. Production recommendation.
3. Business and workload context.
4. Criticality assessment.
5. Ownership assessment.
6. Evidence register.
7. Production-readiness scorecard.
8. Launch blockers.
9. Conditional blockers.
10. Accepted risks.
11. Business readiness.
12. Architecture and dependency readiness.
13. Identity and network readiness.
14. Security readiness.
15. Data and database readiness.
16. Functional and integration readiness.
17. Performance and capacity readiness.
18. Reliability readiness.
19. Observability and logging readiness.
20. Incident-response readiness.
21. Backup and recovery readiness.
22. Deployment and rollback readiness.
23. Operational-support readiness.
24. Cost readiness.
25. Compliance readiness.
26. Documentation and user readiness.
27. Go/no-go criteria.
28. Launch runbook.
29. Hypercare plan.
30. Post-launch validation.
31. Risk register.
32. Exception register.
33. Responsibility matrix.
34. Open questions.
35. Required approvals.
36. Final recommendation.

Perform a Production Readiness Review

For a focused readiness classification pass.
Assess this workload for production readiness.
Classify each domain as:
• Ready
• Ready with accepted risk
• Conditionally ready
• Not ready
• Unable to assess
• Not applicable
Identify blockers, missing evidence, accepted risks, owners, and required completion dates.

Create a Production Readiness Checklist

To tailor a checklist to a specific workload.
Create a risk-based production-readiness checklist for this workload.
Tailor the checklist to:
• Criticality
• Architecture
• Data classification
• User population
• Recovery requirements
• Compliance
• Support model
• Launch method
Separate mandatory controls from advisory improvements.

Create a Production Readiness Scorecard

To generate a weighted scorecard.
Create a weighted production-readiness scorecard.
Include:
• Business
• Ownership
• Architecture
• Security
• Identity
• Network
• Data
• Reliability
• Performance
• Observability
• Recovery
• Operations
• Cost
• Compliance
Define how critical blockers override the score.

Identify Launch Blockers

To isolate what must be resolved before launch.
Review the supplied evidence and identify production launch blockers.
For each blocker include:
• Evidence
• Business impact
• Technical impact
• Required remediation
• Owner
• Validation
• Completion deadline
• Go/no-go impact
Do not classify missing evidence as passed.

Create Go/No-Go Criteria

To define objective launch gates.
Create objective production go/no-go criteria.
Include:
• Business acceptance
• Testing
• Security
• Data
• Capacity
• Monitoring
• Alerts
• Backup
• Recovery
• Support
• Rollback
• Cost
• Compliance
• Change approval
• Required decision-makers

Review Service Ownership

To confirm no ownership role is missing or temporary.
Review production ownership for this service.
Confirm:
• Business owner
• Service owner
• Technical owner
• Application owner
• Data owner
• Security owner
• Support owner
• Budget owner
• Vendor owner
• Risk approver
Identify missing, temporary, or conflicting ownership.

Review Production Architecture

For an architecture readiness pass.
Review this architecture for production readiness.
Assess:
• Failure domains
• Dependencies
• Single points of failure
• Scalability
• Service limits
• Regional design
• Security boundaries
• Data flows
• Recovery
• Supportability
• Operational complexity
• Cost
Separate launch blockers from future improvements.

Review Dependency Readiness

To assess each production dependency.
Review all production dependencies.
For each dependency assess:
• Owner
• Service level
• Authentication
• Connectivity
• Capacity
• Failure behavior
• Timeout
• Retry
• Monitoring
• Support
• Recovery
• Fallback
• Validation evidence

Review Security Readiness

For a security readiness pass.
Assess security readiness for production.
Review:
• Threat model
• Identity
• Privileged access
• Network exposure
• Secrets
• Encryption
• Vulnerabilities
• Logging
• Threat detection
• Incident response
• Security exceptions
• Data protection
Identify findings that should block launch.

Review Performance and Capacity

To validate capacity assumptions.
Assess performance and capacity readiness.
Review:
• Baseline load
• Peak load
• Growth
• Response time
• Throughput
• Error rate
• Concurrency
• Resource utilization
• Scaling
• Quotas
• Provider limits
• Dependency limits
• Cost
Identify missing tests and unsafe assumptions.

Review Observability Readiness

For an observability pass.
Assess observability readiness.
Review:
• Service-health metrics
• Application metrics
• Infrastructure metrics
• Dependency monitoring
• Business metrics
• Logs
• Traces
• Dashboards
• Synthetic tests
• Alert ownership
• Runbooks
• Escalation
• Alert testing
• Cost controls

Review Alert Quality

To audit the alert catalog.
Review the production alert catalog.
For each alert assess:
• Signal
• Threshold
• Duration
• Severity
• Business impact
• Owner
• Routing
• Runbook
• Escalation
• False-positive risk
• Maintenance behavior
• Test evidence
Identify noisy, missing, or unactionable alerts.

Review Backup and Recovery Readiness

To verify recovery, not just backup.
Assess backup and recovery readiness.
Review:
• Backup scope
• Frequency
• Retention
• Encryption
• Immutability
• Geographic protection
• Failure alerts
• Restore permissions
• Restore-test evidence
• RTO
• RPO
• Recovery runbook
• Recovery dependencies
• Failover
• Failback
Do not treat successful backup jobs as proof of recovery.

Create an Operational Handover Package

To transfer the service to operations.
Create an operational handover package.
Include:
• Service definition
• Ownership
• Architecture
• Dependencies
• Support model
• Monitoring
• Alerts
• Runbooks
• Backup
• Recovery
• Access
• Deployment
• Incident response
• Vendor contacts
• Known issues
• Risks
• Cost ownership
• Training
• Acceptance criteria

Create a Launch Runbook

To build the step-by-step launch sequence.
Create a detailed production launch runbook.
For each step include:
• Step number
• Planned time
• Owner
• Preconditions
• Action
• Expected result
• Validation
• Failure response
• Rollback impact
• Evidence
• Status
Include explicit go/no-go checkpoints.

Create a Rollback Plan

To define how the launch is reversed.
Create a production rollback plan.
Define:
• Trigger
• Decision authority
• Decision deadline
• Application rollback
• Infrastructure rollback
• Database rollback
• Data reconciliation
• DNS rollback
• Configuration rollback
• User communication
• Evidence
• Forward-fix criteria
• Irreversible actions

Review Database Change Readiness

When a database change is part of the launch.
Assess the production readiness of this database change.
Review:
• Compatibility
• Schema changes
• Data migration
• Locking
• Duration
• Backward compatibility
• Application dependencies
• Backup
• Rollback
• Validation
• Monitoring
• Performance
• Maintenance window

Create a Hypercare Plan

To plan the stabilization period after launch.
Create a production hypercare plan.
Define:
• Duration
• Staffing
• Support channel
• Monitoring
• Status cadence
• Defect handling
• Business reconciliation
• Vendor support
• Cost review
• Escalation
• Exit criteria
• Transition to normal operations

Create a Risk-Acceptance Record

To formally document an accepted residual risk.
Create a formal risk-acceptance record.
Include:
• Risk
• Evidence
• Business impact
• Technical impact
• Probability
• Compensating controls
• Owner
• Approver
• Acceptance period
• Expiration
• Remediation plan
• Target date
• Review cadence
Do not accept risk without appropriate authority.

Create an Exception Register

To track time-bound control exceptions.
Create a production-readiness exception register.
For each exception include:
• Control
• Justification
• Risk
• Compensating control
• Owner
• Approver
• Start date
• Expiration
• Remediation
• Validation
• Status

Review Cost Readiness

For a financial readiness pass.
Assess financial readiness for production.
Review:
• Cost estimate
• Budget
• Cost center
• Budget owner
• Tags
• Forecast
• Scaling cost
• Logging cost
• Backup cost
• Disaster recovery cost
• Support cost
• Licensing
• Marketplace charges
• Budget alerts
• Anomaly detection
• Unit economics

Create a Post-Launch Validation Plan

To define what is monitored after go-live.
Create a post-launch validation plan.
Monitor:
• Availability
• Performance
• Errors
• Traffic
• Capacity
• Scaling
• Business transactions
• Data quality
• Integrations
• Security alerts
• Backup
• Cost
• Support tickets
• User feedback
Define owners, thresholds, cadence, and escalation.

Perform a Post-Launch Review

After stabilization, to close the loop.
Perform a post-launch review.
Assess:
• Launch execution
• Incidents
• Defects
• Performance
• Capacity
• User impact
• Data reconciliation
• Security
• Monitoring
• Alert quality
• Support
• Cost
• Hypercare
• Outstanding risks
• Lessons learned
Recommend changes to the production-readiness standard.
6. PHASE 3Templateprotected

Go/No-Go Criteria

A structured set of objective go and no-go conditions that gate the production launch decision.
Use this to
  • Define objective launch gates before the review meeting
  • Separate go conditions from disqualifying no-go conditions
  • Ensure required decision-makers and approvals are captured

Go Criteria

Conditions that must all be satisfied for launch to proceed. Example go criteria: business acceptance complete, critical tests pass, no unresolved critical security findings, target environment stable, monitoring active, alerts tested, backup complete, recovery validated, support staffed, rollback available, change approved, budget approved, required decision-makers present.
[List the go conditions for this workload]

No-Go Criteria

Any single condition that disqualifies launch. Example no-go criteria: missing service owner, failed critical business test, data reconciliation failure, unsupported production dependency, critical vulnerability, missing recovery for critical workload, no monitoring for key service health, rollback unavailable, operations has not accepted support, required compliance approval missing.
[List the no-go conditions for this workload]

Required Decision-Makers

Name the individuals whose presence and approval are required for the go/no-go decision.
[Business owner, service owner, security, operations, ...]
7. PHASE 3Templateprotected

Launch Runbook

A step-by-step launch runbook capturing timing, owners, validation, and failure responses for each launch action.
Use this to
  • Sequence the launch with owners and timings
  • Define validation and failure response for each step
  • Embed explicit go/no-go checkpoints in the launch flow

Step Number

Sequential number for the launch step.
[1]

Planned Time

The scheduled time for this step.
[19:00]

Owner

The role accountable for executing this step.
[Change Manager]

Action

The action to perform.
[...]

Preconditions

What must be true before this step begins.
[...]

Expected Result

What success looks like for this step.
[...]

Validation

How the result is verified.
[...]

Failure Response

What to do if the step fails, including rollback impact.
[Delay / No-go / Roll back application]

Evidence

Evidence captured to record completion.
[...]

Completion Status

The recorded status of the step.
[Complete / Held / Reverted]
8. PHASE 3Templateprotected

Rollback Plan

A rollback plan that defines triggers, decision authority, reversal steps, and irreversible actions before launch.
Use this to
  • Define objective rollback triggers and decision authority
  • Sequence application, infrastructure, and data rollback steps
  • Identify irreversible actions and forward-fix criteria

Rollback Trigger

The objective condition(s) that initiate rollback.
[...]

Decision Authority

Who decides to roll back.
[...]

Maximum Decision Time

The deadline for making the rollback decision.
[...]

Application Rollback

How the application is reverted.
[...]

Infrastructure Rollback

How infrastructure changes are reverted.
[...]

Database Rollback

How database changes are reverted, including data reconciliation.
[...]

DNS Rollback

How DNS changes are reverted.
[...]

Configuration Rollback

How configuration changes are reverted.
[...]

User Communication

How users are informed of the rollback.
[...]

Evidence

Evidence captured during rollback.
[...]

Forward-Fix Criteria

When to fix forward instead of rolling back.
[...]

Irreversible Actions

Actions that cannot be undone and require explicit approval.
[...]
9. PHASE 3Templateprotected

Risk-Acceptance Record

A formal record that documents each accepted residual risk, its compensating controls, approver, and expiration.
Use this to
  • Document residual risks accepted at launch
  • Assign a named approver with appropriate authority
  • Set expiration, remediation, and review cadence for each risk

Risk ID

Unique identifier for the risk.
[...]

Description

What the risk is.
[...]

Business Impact

The business consequence if the risk materializes.
[...]

Technical Impact

The technical consequence if the risk materializes.
[...]

Probability

Likelihood of the risk materializing.
[...]

Compensating Control

Controls that reduce the risk.
[...]

Owner

Who owns the risk.
[...]

Risk Approver

The authorized individual accepting the risk. Project managers and engineers should not accept enterprise-level business risk unless authorized.
[...]

Acceptance Date

When the risk was accepted.
[...]

Expiration

When the acceptance expires.
[...]

Remediation Plan

How the risk will be resolved.
[...]

Target Date

When remediation is due.
[...]

Review Cadence

How often the risk is reviewed.
[...]
10. PHASE 3Checklistprotected

Validation Checklist

The comprehensive acceptance checklist a workload must pass across every readiness domain before production launch.
Use this to
  • Verify readiness across business, security, data, operations, and cost
  • Confirm operational acceptance and go/no-go authority before launch
  • Provide auditable evidence that each domain was checked

Confirm the following before production launch:

Business and Ownership

Architecture and Dependencies

Identity and Network

Security

Data and Database

Testing

Reliability and Observability

Operations and Release

Financial and Governance

Launch and Hypercare

🔒 The full execution layer — every checklist, matrix, and the prompt pack — is included with ABME membership.

Unlock Full Blueprint

Full Playbook

Overviewpublic

A cloud workload is not ready for production merely because:

  • The application deploys.
  • Infrastructure exists.
  • A test page loads.
  • A pipeline succeeds.
  • A security scanner reports no critical findings.
  • A business owner wants to launch.
  • A migration deadline has arrived.

Production readiness is the evidence-based determination that a workload can safely support real users, business transactions, operational support, security obligations, recovery requirements, and financial accountability.

A production-ready workload should be:

  • Functionally correct
  • Secure
  • Observable
  • Supportable
  • Recoverable
  • Scalable
  • Resilient
  • Compliant
  • Cost-accountable
  • Operationally owned
  • Governed through controlled change

The production-readiness process should answer: “Do we have enough evidence to place this workload into production, operate it responsibly, detect and respond to failure, recover it when necessary, and understand the residual risks?”

A production review should not become a checklist ceremony performed at the end of a project. Readiness should be built throughout the delivery lifecycle. Late reviews frequently uncover:

  • Missing ownership
  • Incomplete monitoring
  • Weak alert routing
  • Unvalidated backup
  • Unclear recovery objectives
  • Excessive permissions
  • Incomplete testing
  • Missing documentation
  • Uncontrolled deployment access
  • Unbudgeted cloud services
  • Unsupported dependencies
  • Undefined incident response
  • No rollback plan
  • No operational handover

This workflow uses AI to help structure a production-readiness assessment, evaluate evidence, identify blockers, classify residual risk, prepare launch criteria, and create an operational acceptance package.

The objective is not to eliminate all risk. The objective is to ensure risk is known, controlled, assigned, and accepted at the appropriate level before production use.

Business Problempublic

Organizations often treat production launch as a project-management milestone rather than an operational-risk decision. This can lead to workloads entering production with:

  • No service owner
  • No support model
  • No service-level objectives
  • Missing dashboards
  • Untested alerts
  • Unverified backups
  • Unrehearsed recovery
  • Incomplete security controls
  • Missing cost ownership
  • Undocumented dependencies
  • Unapproved exceptions
  • Weak rollback procedures
  • No hypercare plan
  • No user communication
  • No capacity evidence
  • No post-launch review

When these gaps are discovered after launch:

  • Incidents last longer.
  • Users lose confidence.
  • Support teams improvise.
  • Recovery attempts fail.
  • Security findings become emergencies.
  • Cloud costs exceed forecasts.
  • Project teams remain permanent support teams.
  • Audit evidence is incomplete.
  • Technical debt becomes operational debt.

A strong production-readiness process establishes objective launch gates and requires accountable decision-makers to accept residual risk.

Typical Use Casespublic

Use this workflow when:

  • Launching a new cloud application
  • Moving a migrated workload into production
  • Promoting a major application release
  • Launching a new managed service
  • Deploying a Kubernetes workload
  • Launching a public API
  • Launching an internal enterprise service
  • Moving from pilot to general availability
  • Launching a regulated workload
  • Replacing a critical legacy system
  • Preparing for a high-traffic event
  • Launching in a new region
  • Completing an operational handover
  • Reviewing a troubled preproduction environment
  • Creating a standardized production-readiness review
  • Establishing a production launch gate
  • Preparing executive launch approval
  • Preparing audit evidence
  • Creating a reusable service-acceptance process

Do NOT Use This Workflow Whenpublic

This workflow is not intended to:

  • Replace application testing
  • Replace penetration testing
  • Replace architecture review
  • Replace legal or compliance approval
  • Replace business acceptance
  • Guarantee zero incidents
  • Approve a launch solely because a checklist is complete
  • Treat documentation as proof of operational capability
  • Treat a backup configuration as proof of recoverability
  • Treat a successful deployment as proof of service readiness
  • Treat scanner output as complete security assurance
  • Accept unresolved critical risks without appropriate authority
  • Launch because a deadline exists
  • Allow project teams to self-approve all risk
  • Automatically deploy AI-generated changes
  • Store credentials or production secrets in prompts
  • Use generic acceptance criteria for every workload
  • Ignore workload criticality
  • Ignore customer and user impact
  • Ignore financial accountability

Production approval must remain a human governance decision supported by evidence.

Expected Outcomepublic

After completing this workflow, you should have:

  • Production-readiness scope
  • Business and technical ownership
  • Workload criticality classification
  • Production-readiness scorecard
  • Evidence register
  • Functional readiness assessment
  • Architecture readiness assessment
  • Security readiness assessment
  • Identity readiness assessment
  • Network readiness assessment
  • Data readiness assessment
  • Reliability readiness assessment
  • Capacity and performance assessment
  • Observability assessment
  • Backup and recovery assessment
  • Operational support assessment
  • Incident response assessment
  • Change and release assessment
  • Cost and financial assessment
  • Compliance assessment
  • Dependency assessment
  • Launch runbook
  • Rollback plan
  • Go/no-go criteria
  • Risk register
  • Exception register
  • Hypercare plan
  • Operational acceptance
  • Post-launch validation plan
  • Executive launch recommendation

🔒 The complete playbook — reference models, worked examples, and operational guidance — is included with ABME membership.

Unlock Full Blueprint

Production Readiness Objectivesprotected

A complete production-readiness review should answer:

  1. What business capability is being launched?
  2. Who owns the service?
  3. Who supports the service?
  4. What is the workload criticality?
  5. Who are the users and customers?
  6. What service levels are expected?
  7. What dependencies exist?
  8. What happens when a dependency fails?
  9. What security controls are required?
  10. What data is processed?
  11. What compliance obligations apply?
  12. What capacity has been validated?
  13. What performance is acceptable?
  14. How will the service be monitored?
  15. Which alerts require action?
  16. How will incidents be handled?
  17. What is backed up?
  18. How will recovery occur?
  19. What changes are allowed at launch?
  20. How will rollback work?
  21. What costs are expected?
  22. Who owns the budget?
  23. Which risks remain?
  24. Who may accept those risks?
  25. What conditions block launch?
  26. What evidence is required?
  27. What happens during hypercare?
  28. When does operations formally accept the service?
  29. How will launch success be measured?
  30. When will the post-launch review occur?

Production Readiness Principlesprotected

Recommended principles include:

  • Readiness is evidence-based.
  • Criticality determines control depth.
  • Production ownership must be explicit.
  • Security is required before launch.
  • Recovery must be tested.
  • Monitoring must be actionable.
  • Alerts need owners.
  • Capacity must be demonstrated.
  • Dependencies must be understood.
  • Rollback must be defined.
  • Exceptions must be time-bound.
  • Residual risks require named acceptance.
  • Cost ownership is part of readiness.
  • Operational handover is a formal event.
  • Launch success includes stabilization.
  • Lessons learned should improve the standard.

Workload Criticalityprotected

Tier 0 — Mission Critical

Failure may cause severe financial loss, safety impact, regulatory breach, major customer harm, or enterprise-wide disruption.
  • Highly available architecture
  • Strong recovery
  • Continuous monitoring
  • Formal incident command
  • Extensive testing
  • Senior risk approval

Tier 1 — Business Critical

Failure materially affects key business operations or customers.
  • Defined RTO and RPO
  • Redundancy
  • Strong monitoring
  • Tested recovery
  • Formal support
  • Controlled deployment

Tier 2 — Important

Failure affects productivity or a limited business process.
  • Standard backup
  • Business-hours support
  • Defined restoration
  • Basic redundancy
  • Standard monitoring

Tier 3 — Standard

Failure has limited and recoverable business impact. Controls may be proportionate and simpler.

Readiness Decision Categoriesprotected

Ready

All mandatory controls are implemented and validated.

Ready with Accepted Risk

No critical blockers remain, but documented residual risks have been accepted by authorized stakeholders.

Conditionally Ready

Launch is permitted only after specified conditions are completed before a defined deadline.

Not Ready

One or more blockers create unacceptable production risk.

Unable to Assess

Required evidence is missing or unreliable.

Evidence Standardsprotected

Evidence may include:

  • Test results
  • Architecture diagrams
  • Configuration exports
  • Security scans
  • Penetration-test reports
  • Performance-test results
  • Recovery-test evidence
  • Monitoring screenshots
  • Alert tests
  • Change records
  • Deployment logs
  • Cost estimates
  • Budget approvals
  • Compliance approvals
  • Support runbooks
  • Business sign-off
  • Incident-response exercises
  • Dependency maps

Evidence should be:

  • Current
  • Relevant
  • Reproducible
  • Attributable
  • Retained
  • Approved where required

Mandatory Versus Advisory Controlsprotected

Mandatory Controls

Failure blocks production unless an authorized exception exists.
  • Critical security vulnerability
  • Missing business owner
  • Missing data-protection control
  • No recovery capability for a critical workload
  • Unsupported production dependency
  • No rollback for a high-risk launch
  • No monitoring for critical service health

Advisory Controls

Recommended improvements that do not independently block launch. Advisory findings should still have owners and target dates.

Blocker Classificationprotected

Launch Blocker

Launch should not proceed.

Conditional Blocker

Launch may proceed only after a specific condition is completed.

High-Priority Risk

Requires accepted risk and near-term remediation.

Improvement

Should be planned but does not block launch.

Observation

Requires no immediate action.

Readiness Domainsprotected

Business Readiness

Confirm:

  • Business objective
  • Business owner
  • Product owner
  • User population
  • Customer impact
  • Launch date
  • Business calendar
  • Communication plan
  • Acceptance criteria
  • Training
  • Support expectations
  • Financial approval
  • Legal and compliance approval

Service Ownership

Every production service should have:

  • Business owner
  • Service owner
  • Technical owner
  • Application owner
  • Data owner
  • Security owner
  • Support owner
  • Budget owner
  • Vendor owner where applicable

Ownership should not be assigned only to a temporary project team.

Service Definition

Document:

  • Service name
  • Purpose
  • Customers
  • Users
  • Criticality
  • Business hours
  • Support hours
  • Service-level targets
  • Dependencies
  • Data classification
  • Deployment model
  • Recovery tier
  • Cost center
  • Escalation path

Architecture Readiness

Validate:

  • Approved target architecture
  • Component inventory
  • Data flows
  • Trust boundaries
  • Failure domains
  • Dependency behavior
  • Scalability
  • High availability
  • Service limits
  • Provider quotas
  • Regional design
  • Technology support
  • Lifecycle
  • Technical debt
  • Known constraints

Architecture Review Questions

  • Does the design meet business objectives?
  • Are all critical dependencies identified?
  • Are single points of failure known?
  • Are failure modes understood?
  • Is resilience proportional to criticality?
  • Are service limits understood?
  • Are provider-region constraints known?
  • Is the design supportable by available staff?
  • Are architecture exceptions documented?
  • Is the design unnecessarily complex?

Dependency Readiness

For each dependency document:

  • Dependency name
  • Owner
  • Service level
  • Authentication
  • Network path
  • Failure mode
  • Timeout
  • Retry behavior
  • Capacity
  • Monitoring
  • Support contact
  • Recovery expectation
  • Alternative or fallback
  • Validation evidence

Third-Party Dependency Review

Assess:

  • Vendor support
  • Contract
  • Service-level agreement
  • Security review
  • Privacy review
  • Data location
  • Integration limits
  • Rate limits
  • Maintenance windows
  • Incident notification
  • Exit strategy
  • Financial exposure

Identity Readiness

Validate:

  • User authentication
  • Single sign-on
  • Multi-factor authentication
  • Conditional access
  • Role mapping
  • Least privilege
  • Privileged access
  • Service identities
  • Workload identities
  • Guest access
  • Emergency access
  • Access reviews
  • Joiner, mover, and leaver process
  • Authentication logging

Privileged Access

Confirm:

  • Named privileged users
  • Separate administrative accounts
  • Just-in-time access
  • Approval
  • Session logging
  • MFA
  • Emergency access
  • Periodic review
  • Credential rotation
  • No shared administrator credentials

Workload Identity

For each service identity verify:

  • Owner
  • Purpose
  • Permission scope
  • Authentication method
  • Credential lifetime
  • Rotation
  • Logging
  • Dependency
  • Failure behavior
  • Decommissioning

Prefer short-lived managed or federated identities over static credentials.

Network Readiness

Validate:

  • Addressing
  • Routing
  • Segmentation
  • Firewalls
  • Security groups
  • Load balancers
  • Private endpoints
  • Internet ingress
  • Internet egress
  • DNS
  • Certificates
  • DDoS protection
  • Hybrid connectivity
  • Remote administration
  • Flow logging
  • Network monitoring
  • Capacity
  • Failover

Internet Exposure

For every public endpoint document:

  • Business purpose
  • Protocol
  • Port
  • Authentication
  • TLS
  • Web application firewall
  • DDoS protection
  • Rate limiting
  • Logging
  • Owner
  • Vulnerability testing
  • Certificate renewal
  • Abuse monitoring

DNS Readiness

Confirm:

  • Production records
  • Ownership
  • TTL
  • Private versus public resolution
  • Health checks
  • Failover behavior
  • Certificate names
  • Monitoring
  • Change process
  • Rollback records
  • Expiration and renewal

Certificate Readiness

Validate:

  • Certificate authority
  • Subject names
  • Expiration
  • Renewal automation
  • Private-key protection
  • Load-balancer configuration
  • Client trust
  • Revocation
  • Alerting
  • Emergency replacement

Security Readiness

Review:

  • Threat model
  • Security architecture
  • Vulnerability findings
  • Penetration testing
  • Dependency scanning
  • Configuration scanning
  • Secrets
  • Encryption
  • Logging
  • Threat detection
  • Identity
  • Network exposure
  • Data protection
  • Security exceptions
  • Incident response

Security Findings

Classify findings as:

  • Critical blocker
  • High risk
  • Medium risk
  • Low risk
  • Informational

A severity score alone should not determine launch readiness. Consider:

  • Exploitability
  • Exposure
  • Data impact
  • Business impact
  • Compensating controls
  • Detection
  • Recovery
  • Time to remediation

Vulnerability Readiness

Confirm:

  • Operating-system scanning
  • Container-image scanning
  • Dependency scanning
  • Source-code scanning
  • Infrastructure scanning
  • Runtime scanning where applicable
  • Remediation ownership
  • Exception process
  • Rescan evidence
  • Vulnerability age

Secret Readiness

Confirm:

  • No secrets in source control
  • No secrets in application images
  • No secrets in build logs
  • Approved secret store
  • Access control
  • Rotation
  • Expiration
  • Audit logging
  • Emergency rotation
  • Ownership

Encryption Readiness

Validate:

  • Data at rest
  • Data in transit
  • Internal service communication
  • Backup encryption
  • Key ownership
  • Key rotation
  • Key recovery
  • Certificate management
  • Customer-managed key requirements
  • Hardware security module requirements

Data Readiness

Assess:

  • Data classification
  • Data owner
  • Data location
  • Data residency
  • Data flow
  • Retention
  • Archival
  • Deletion
  • Legal hold
  • Privacy
  • Backup
  • Recovery
  • Masking
  • Production test-data restrictions
  • Data quality
  • Data reconciliation

Database Readiness

Validate:

  • Engine and version
  • Support status
  • High availability
  • Backup
  • Point-in-time recovery
  • Maintenance
  • Patching
  • Capacity
  • Performance
  • Connection limits
  • Authentication
  • Encryption
  • Monitoring
  • Query performance
  • Schema migration
  • Rollback
  • Licensing

Data Migration Readiness

For migrated data confirm:

  • Source of truth
  • Migration completed
  • Counts reconciled
  • Checksums or validation
  • Business totals reconciled
  • In-flight transactions handled
  • Data quality issues documented
  • Rollback implications understood
  • Business owner approved

Functional Readiness

Validate:

  • Core user journeys
  • Business rules
  • Error handling
  • Notifications
  • Reports
  • Scheduled jobs
  • Batch processes
  • Administrative functions
  • Integrations
  • Accessibility
  • Localization where required
  • Browser or client compatibility
  • Mobile behavior where applicable

Test Coverage

Testing may include:

  • Unit testing
  • Component testing
  • Integration testing
  • Contract testing
  • End-to-end testing
  • Regression testing
  • User acceptance testing
  • Performance testing
  • Security testing
  • Recovery testing
  • Operational testing
  • Rollback testing

Test Defect Review

For every unresolved defect document:

  • Defect ID
  • Severity
  • Business impact
  • Technical impact
  • Workaround
  • Owner
  • Target resolution
  • Launch impact
  • Risk approver
  • Acceptance status

Performance Readiness

Validate:

  • Response time
  • Throughput
  • Concurrency
  • Batch duration
  • Queue depth
  • Database performance
  • Network latency
  • Storage performance
  • Error rate
  • Scaling
  • Failover performance
  • Resource saturation
  • Performance under degraded dependencies

Capacity Model

Document:

  • Expected baseline
  • Expected peak
  • Seasonal peak
  • Growth assumption
  • Capacity margin
  • Scaling trigger
  • Scaling limit
  • Quota
  • Provider limit
  • Downstream capacity
  • Cost impact

Load Testing

Load testing should:

  • Reflect realistic behavior.
  • Include representative data.
  • Include critical integrations.
  • Measure system limits.
  • Test scaling.
  • Test recovery after load.
  • Identify bottlenecks.
  • Validate alerting.
  • Record cost.

Stress and Failure Testing

Where appropriate, test:

  • Resource exhaustion
  • Dependency latency
  • Dependency failure
  • Network interruption
  • Database failover
  • Node loss
  • Zone loss
  • Regional impairment
  • Queue backlog
  • Certificate failure
  • Credential expiration

Reliability Readiness

Validate:

  • Availability objectives
  • Redundancy
  • Health checks
  • Failover
  • Retry behavior
  • Timeouts
  • Circuit breakers
  • Queue handling
  • Graceful degradation
  • Rate limits
  • Idempotency
  • Data consistency
  • Maintenance behavior
  • Recovery procedures

Service-Level Objectives

Define measurable objectives such as:

  • Availability
  • Latency
  • Throughput
  • Error rate
  • Data freshness
  • Completion time
  • Recovery time
  • Support response

Service-level objectives should align with business expectations.

Error Budgets

For mature services, define:

  • SLO
  • Error budget
  • Measurement window
  • Consumption rate
  • Alert thresholds
  • Release restrictions
  • Escalation
  • Review cadence

Observability Readiness

Production monitoring should cover:

  • Availability
  • Latency
  • Errors
  • Traffic
  • Saturation
  • Dependencies
  • Business transactions
  • Security
  • Cost
  • Backup
  • Recovery
  • Deployment health

Monitoring Layers

Include:

  • Infrastructure monitoring
  • Platform monitoring
  • Application monitoring
  • Database monitoring
  • Network monitoring
  • User-experience monitoring
  • Synthetic monitoring
  • Business-process monitoring
  • Security monitoring
  • Cost monitoring

Dashboard Readiness

Dashboards should provide:

  • Service health
  • Key SLOs
  • Current incidents
  • Error rates
  • Latency
  • Traffic
  • Resource saturation
  • Dependency health
  • Deployment markers
  • Business metrics
  • Cost anomalies

Dashboards should support operational decisions, not merely display available metrics.

Alert Readiness

For every material alert define:

  • Alert condition
  • Threshold
  • Duration
  • Severity
  • Owner
  • Routing
  • Business impact
  • Runbook
  • Escalation
  • Suppression rules
  • Maintenance behavior
  • Validation evidence

Alert Testing

Test:

  • Alert generation
  • Notification
  • On-call receipt
  • Escalation
  • Runbook access
  • Acknowledgment
  • Resolution
  • Closure
  • Maintenance suppression

An untested alert should not be assumed operational.

Logging Readiness

Confirm:

  • Application logs
  • Infrastructure logs
  • Identity logs
  • Security logs
  • Administrative logs
  • Network logs
  • Database logs
  • Audit logs
  • Retention
  • Access
  • Redaction
  • Time synchronization
  • Correlation identifiers
  • Searchability
  • Cost controls

Sensitive Data in Logs

Prevent or control:

  • Passwords
  • Tokens
  • Private keys
  • Session identifiers
  • Personal information
  • Payment information
  • Health information
  • Confidential business data
  • Connection strings

Incident Response Readiness

Confirm:

  • On-call schedule
  • Incident severity model
  • Escalation
  • Incident commander
  • Technical responders
  • Business contacts
  • Security contacts
  • Vendor contacts
  • Communications
  • Evidence retention
  • Post-incident review
  • Runbooks

Incident Scenarios

Prepare for:

  • Application outage
  • Database outage
  • Regional impairment
  • Identity failure
  • Network failure
  • Data corruption
  • Security compromise
  • Credential compromise
  • Dependency outage
  • Performance degradation
  • Capacity exhaustion
  • Cost spike
  • Failed deployment

Runbook Readiness

Runbooks should exist for:

  • Service restart
  • Scaling
  • Dependency failure
  • Database failover
  • Backup restore
  • Certificate renewal
  • Credential rotation
  • Alert response
  • Deployment rollback
  • Access failure
  • Queue backlog
  • Regional failover
  • Security containment
  • Vendor escalation

Backup Readiness

Validate:

  • Backup scope
  • Frequency
  • Retention
  • Encryption
  • Immutability
  • Geographic protection
  • Monitoring
  • Ownership
  • Failure alerting
  • Restore permissions
  • Backup cost

Recovery Readiness

Validate:

  • RTO
  • RPO
  • Recovery architecture
  • Recovery dependencies
  • Recovery runbook
  • Data recovery
  • Identity recovery
  • DNS recovery
  • Network recovery
  • Application recovery
  • Communication
  • Failover
  • Failback
  • Recovery testing

Restore Testing

Evidence should include:

  • Restore date
  • Data restored
  • Environment
  • Duration
  • Integrity result
  • Application validation
  • RTO result
  • RPO result
  • Defects
  • Owner
  • Follow-up actions

Disaster Recovery Readiness

For workloads requiring disaster recovery, confirm:

  • Recovery environment
  • Infrastructure reproducibility
  • Replication
  • Capacity
  • DNS
  • Identity
  • Network
  • Secrets
  • Keys
  • Monitoring
  • Security
  • Runbook
  • Decision authority
  • Exercise results
  • Failback

Operational Support Readiness

Confirm:

  • Support hours
  • On-call coverage
  • Support tiers
  • Escalation
  • Service desk knowledge
  • Support group
  • Ticket routing
  • Vendor support
  • Runbooks
  • Known-error records
  • Operational training
  • Support acceptance

Operational Handover

The handover package should include:

  • Service definition
  • Architecture
  • Dependency map
  • Support model
  • Monitoring
  • Alerts
  • Runbooks
  • Backup
  • Recovery
  • Access
  • Deployment
  • Incident process
  • Vendor contacts
  • Known issues
  • Risks
  • Cost ownership
  • Documentation repository

Service Acceptance

Operations should formally accept the service after confirming:

  • Ownership
  • Access
  • Monitoring
  • Alerts
  • Runbooks
  • Backup
  • Recovery
  • Training
  • Support contacts
  • Escalation
  • Known risks
  • Documentation
  • Staffing

Change Readiness

Confirm:

  • Change policy
  • Release process
  • Deployment window
  • Branch protection
  • Approval
  • Separation of duties
  • Rollback
  • Database-change procedure
  • Feature-flag process
  • Emergency change
  • Evidence retention

Deployment Readiness

Validate:

  • Artifact versioning
  • Build reproducibility
  • Artifact integrity
  • Security scans
  • Environment configuration
  • Secret retrieval
  • Infrastructure plan
  • Database migrations
  • Deployment sequence
  • Health checks
  • Rollback
  • Post-deployment validation
  • Ownership

Release Strategy

Possible approaches include:

  • Standard rolling deployment
  • Blue-green
  • Canary
  • Feature flags
  • Phased user rollout
  • Regional rollout
  • Shadow traffic

Select based on:

  • Risk
  • Architecture
  • Data compatibility
  • User impact
  • Rollback capability
  • Operational maturity

Database Change Readiness

Database changes should define:

  • Compatibility
  • Migration sequence
  • Locking risk
  • Long-running operations
  • Backward compatibility
  • Rollback
  • Data backup
  • Validation
  • Application release dependency
  • Maintenance window

Rollback Readiness

Define:

  • Rollback trigger
  • Decision authority
  • Maximum decision time
  • Application rollback
  • Infrastructure rollback
  • Database rollback
  • Data reconciliation
  • DNS rollback
  • Configuration rollback
  • User communication
  • Evidence
  • Forward-fix criteria

Irreversible Changes

Identify:

  • Data transformations
  • Destructive schema changes
  • Key rotation
  • Source-system retirement
  • License transfer
  • Contract change
  • Endpoint removal
  • Data deletion
  • Permanent migration steps

Require explicit approval.

Financial Readiness

Validate:

  • Cost estimate
  • Budget
  • Cost center
  • Budget owner
  • Tagging
  • Forecast
  • Scaling cost
  • Logging cost
  • Backup cost
  • Disaster recovery cost
  • Support cost
  • Marketplace charges
  • License cost
  • Anomaly alerts
  • Cost review cadence

Cost Guardrails

Potential controls include:

  • Budget alerts
  • Forecast alerts
  • Anomaly detection
  • Scaling limits
  • Sandbox expiration
  • Log-volume alerts
  • Storage-growth alerts
  • High-cost resource approval
  • Commitment review
  • Unit-cost monitoring

Unit Economics

Where applicable define:

  • Cost per customer
  • Cost per user
  • Cost per transaction
  • Cost per API call
  • Cost per device
  • Cost per report
  • Cost per workload
  • Cost per revenue dollar

Compliance Readiness

Assess applicable obligations such as:

  • Data privacy
  • Data residency
  • Retention
  • Audit logging
  • Access review
  • Encryption
  • Vulnerability management
  • Incident notification
  • Vendor management
  • Business continuity
  • Evidence retention

Only apply frameworks relevant to the workload.

Documentation Readiness

Required documentation may include:

  • Service overview
  • Architecture diagram
  • Data-flow diagram
  • Dependency map
  • Ownership
  • Support model
  • Deployment procedure
  • Rollback procedure
  • Monitoring
  • Alerts
  • Runbooks
  • Backup
  • Recovery
  • Security design
  • Cost model
  • Known issues
  • Risk register
  • Contact list

User Readiness

Confirm:

  • User communication
  • Training
  • Support instructions
  • Access
  • Client requirements
  • New URLs
  • Maintenance notice
  • Known limitations
  • Feedback channel
  • Accessibility
  • Launch timing

Hypercare

A hypercare plan may include:

  • Duration
  • Extended staffing
  • Dedicated support channel
  • Increased monitoring
  • Daily status
  • Rapid defect triage
  • Vendor support
  • Business reconciliation
  • Cost review
  • Exit criteria
  • Transition to normal operations

Hypercare Exit Criteria

Examples:

  • No unresolved critical defects
  • Incident rate within expected range
  • Performance stable
  • Data reconciled
  • Support volume declining
  • Monitoring effective
  • Operations accepts steady-state support
  • Business owner approves
  • Cost within expected range

Post-Launch Validation

Validate after launch:

  • Availability
  • Performance
  • Error rates
  • User access
  • Business transactions
  • Data quality
  • Integrations
  • Security alerts
  • Backup
  • Cost
  • Support tickets
  • User feedback
  • Capacity
  • Scaling

Go/No-Go Criteriaprotected

Example go criteria:

  • Business acceptance complete
  • Critical tests pass
  • No unresolved critical security findings
  • Target environment stable
  • Monitoring active
  • Alerts tested
  • Backup complete
  • Recovery validated
  • Support staffed
  • Rollback available
  • Change approved
  • Budget approved
  • Required decision-makers present

Example no-go criteria:

  • Missing service owner
  • Failed critical business test
  • Data reconciliation failure
  • Unsupported production dependency
  • Critical vulnerability
  • Missing recovery for critical workload
  • No monitoring for key service health
  • Rollback unavailable
  • Operations has not accepted support
  • Required compliance approval missing

Risk Acceptanceprotected

Each accepted risk should include:

  • Risk ID
  • Description
  • Business impact
  • Technical impact
  • Probability
  • Compensating control
  • Owner
  • Risk approver
  • Acceptance date
  • Expiration
  • Remediation plan
  • Target date
  • Review cadence

Project managers and engineers should not accept enterprise-level business risk unless authorized.

Exception Managementprotected

Exceptions should be:

  • Specific
  • Justified
  • Approved
  • Time-bound
  • Monitored
  • Assigned
  • Reviewed
  • Closed

Avoid generic exceptions such as “launch deadline.”

Production Readiness Review Meetingprotected

Participants may include:

  • Business owner
  • Service owner
  • Product owner
  • Application team
  • Platform team
  • Security
  • Operations
  • Network
  • Identity
  • Data
  • Finance
  • Compliance
  • Change management
  • Vendor representative where required

Review agenda:

  1. Business objective
  2. Workload criticality
  3. Scope
  4. Architecture
  5. Test results
  6. Security
  7. Data
  8. Reliability
  9. Monitoring
  10. Recovery
  11. Operations
  12. Cost
  13. Open risks
  14. Exceptions
  15. Go/no-go decision
  16. Conditions and owners

Decision record should document:

  • Decision
  • Date
  • Scope
  • Participants
  • Evidence reviewed
  • Blockers
  • Accepted risks
  • Conditions
  • Approvers
  • Launch window
  • Next review

Production Readiness Metricsprotected

Useful metrics include:

  • Workloads reviewed
  • Ready on first review
  • Average readiness lead time
  • Blockers per workload
  • Critical security blockers
  • Missing ownership findings
  • Missing-monitoring findings
  • Failed recovery tests
  • Unresolved high risks at launch
  • Exceptions granted
  • Exception age
  • Post-launch incidents
  • Rollbacks
  • Hypercare duration
  • Budget variance
  • Performance variance
  • Support-ticket volume
  • Time to operational acceptance
  • Post-launch defect rate
  • Evidence completeness

Production Readiness Maturity Modelprotected

Level 1 — Ad Hoc

  • Launch decisions are informal.
  • Evidence is inconsistent.
  • Operations learns about services late.
  • Risks are poorly documented.

Level 2 — Checklist Driven

  • Standard checklist exists.
  • Reviews occur near launch.
  • Evidence quality varies.
  • Exceptions are informal.

Level 3 — Governed

  • Risk-based reviews
  • Mandatory evidence
  • Named ownership
  • Formal risk acceptance
  • Operational acceptance
  • Repeatable launch gates

Level 4 — Integrated

  • Readiness built into delivery lifecycle
  • Automated evidence
  • Continuous security and policy validation
  • Standard platform capabilities
  • Measured launch outcomes

Level 5 — Optimized

  • Risk-based automation
  • Predictive readiness
  • Production telemetry feeds future reviews
  • Exception trends drive platform improvements
  • Developer experience and reliability are jointly optimized

Governanceprotected

Define standards for:

  • Workload criticality
  • Mandatory production controls
  • Required evidence
  • Review timing
  • Review participants
  • Go/no-go authority
  • Risk acceptance
  • Exceptions
  • Security approval
  • Compliance approval
  • Operational acceptance
  • Budget approval
  • Launch communication
  • Hypercare
  • Post-launch review
  • Evidence retention
  • Standard updates

Example Workload Inputprotected

Workload: Public Customer Portal

Criticality: Tier 1 — Business Critical

Architecture:

  • Public web application
  • Managed application platform
  • Managed relational database
  • Object storage
  • Content delivery network
  • Web application firewall
  • Federated customer identity
  • Central logging
  • Multi-zone deployment
  • Secondary-region backup

Launch Constraints:

  • Public launch date already announced
  • Expected launch-day traffic is uncertain
  • Penetration testing found two medium findings
  • Database restore succeeded in test
  • Regional failover has not been exercised
  • Operations has not yet signed service acceptance
  • Cost estimate excludes increased launch-day logging

Example Executive Recommendationprotected

Recommendation: Conditionally Ready

The workload should not receive final production approval until the following conditions are completed:

  1. Operations formally accepts the support model.
  2. Launch-day capacity assumptions are validated through a representative load test.
  3. Logging-volume cost is included in the approved budget.
  4. Regional recovery limitations are documented and accepted.
  5. The two unresolved security findings receive formal disposition.

No confirmed critical security or functional blocker was identified. The announced launch date should not override the unresolved operational-ownership condition.

Example Production Readiness Findingsprotected

PRD-001 — Operations Has Not Accepted Service Ownership

Classification: Launch Blocker · Severity: Critical · Confidence: High

Evidence: No signed operational acceptance exists, and the on-call group has not completed service training.

Business Impact: Production incidents may not receive timely response.

Recommended Action: Complete operational handover, verify access, test alert routing, and record formal acceptance.

PRD-002 — Launch-Day Capacity Is Unvalidated

Classification: Conditional Blocker · Severity: High · Confidence: High

Evidence: Performance tests covered approximately thirty percent of forecast peak traffic, and launch-day demand is uncertain.

Recommended Action: Run a representative load test, validate autoscaling, confirm quotas, and establish launch-day scaling guardrails.

PRD-003 — Regional Failover Has Not Been Tested

Classification: High-Priority Risk · Severity: High · Confidence: High

Evidence: The secondary-region design exists, but DNS, data recovery, identity, and application startup have not been exercised together.

Recommended Action: Complete an integrated recovery exercise or formally accept the current recovery limitation before launch.

PRD-004 — Launch Logging Cost Missing from Forecast

Classification: Conditional Blocker · Severity: Medium · Confidence: High

Evidence: Debug and access-log volumes will be increased during launch, but the approved cost estimate reflects normal retention and ingestion.

Recommended Action: Estimate temporary and steady-state logging cost, confirm budget ownership, and define a date to return to normal logging levels.

Example Go/No-Go Criteriaprotected

Go:

  • Operations has accepted the service.
  • Critical user journeys pass.
  • Load testing validates expected peak demand.
  • No critical security findings remain.
  • Medium findings have approved dispositions.
  • Monitoring and alerts are active and tested.
  • Backup and restore evidence is current.
  • Rollback remains practical.
  • Budget approval includes launch conditions.
  • Business owner approves launch.

No-Go:

  • No active support ownership
  • Failed payment or login journey
  • Unresolved critical vulnerability
  • Database reconciliation failure
  • Capacity below expected demand
  • Missing alert routing
  • Rollback unavailable
  • Required compliance approval absent

Example Launch Runbook Extractprotected

StepTimeOwnerActionValidationFailure Response
119:00Change ManagerOpen launch bridgeRequired participants presentDelay launch
219:10Release ManagerConfirm approved artifactVersion and signature matchNo-go
319:20Database OwnerApply backward-compatible schema changeHealth and validation checks passStop and assess
419:40DevOps EngineerDeploy production applicationHealth checks passRoll back application
520:00Test LeadExecute critical user journeysAll critical tests passGo/no-go review
620:20OperationsConfirm alerts and dashboardsSignals visible and routedNo-go
720:30Business OwnerApprove public launchApproval recordedDelay
820:40Network EngineerEnable public trafficSynthetic and real traffic healthyRevert routing
921:00Incident LeadBegin hypercare monitoringMetrics within thresholdsEscalate

Example Risk Registerprotected

RiskProbabilityImpactMitigationOwner
Launch traffic exceeds forecastMediumHighLoad testing, quotas, scaling, CDNService Owner
Customer identity provider degradesLowCriticalMonitoring, rate limits, vendor escalationIdentity Owner
Logging cost spikesHighMediumVolume alerts, temporary retention, budgetFinOps
Regional recovery failsMediumHighRecovery exercise and accepted limitationDR Owner
Medium security finding becomes exploitableLowHighCompensating control and target remediationSecurity Owner
Support team lacks application knowledgeMediumHighTraining, runbooks, hypercareOperations
Database migration requires forward recoveryLowCriticalBackward compatibility and tested backupDBA

Automation Opportunitiesprotected

  • Readiness intake
  • Criticality scoring
  • Ownership validation
  • Evidence collection
  • Test-result aggregation
  • Security finding aggregation
  • Backup validation
  • Monitoring checks
  • Alert checks
  • Cost checks
  • Scorecard generation
  • Blocker detection
  • Exception tracking
  • Approval routing
  • Launch runbook generation
  • Hypercare reporting
  • Post-launch review
  • A mature automated workflow could import workload metadata, determine required controls by criticality, collect pipeline/testing/security/compliance evidence, validate monitoring, alerts, backup, recovery, ownership, support, cost, and budget, generate a readiness scorecard, highlight blockers, route conditions and risks to owners, require human approval, record the go/no-go decision, monitor hypercare, compare launch outcomes to readiness predictions, and improve standards continuously.

Pro Tipsprotected

  • Define readiness requirements at project start.
  • Tailor controls to workload criticality.
  • Name the service owner early.
  • Separate mandatory controls from improvements.
  • Require evidence for every material conclusion.
  • Treat missing evidence as unknown—not passed.
  • Test critical business journeys.
  • Test dependency failure.
  • Validate provider quotas.
  • Include launch-day and seasonal capacity.
  • Monitor business outcomes as well as infrastructure.
  • Test alert routing before launch.
  • Test backup restoration.
  • Test integrated recovery.
  • Define objective rollback triggers.
  • Identify irreversible changes.
  • Require operations to accept the service.
  • Assign risk acceptance to the correct authority.
  • Time-limit exceptions.
  • Include cost and budget ownership.
  • Define hypercare exit criteria.
  • Schedule the post-launch review before launch.
  • Use production outcomes to improve the standard.
  • Require qualified human approval for every production decision.

Common Mistakesprotected

  • Reviewing production readiness too late — readiness gaps discovered immediately before launch are expensive and politically difficult to resolve.
  • Treating a checklist as approval — a checklist records evidence; it does not replace technical judgment or risk authority.
  • Using an overall score to hide a critical blocker — a single unresolved critical issue may outweigh a high average score.
  • Launching without a service owner — temporary project ownership does not create sustainable production accountability.
  • Launching without operational acceptance — a service should not enter production when no team is prepared to support it.
  • Assuming monitoring exists because metrics are collected — metrics must be connected to dashboards, alerts, ownership, and response procedures.
  • Creating alerts without testing them — an alert may fail to route, escalate, or provide enough context.
  • Testing infrastructure but not business journeys — healthy infrastructure does not prove that customers can complete critical transactions.
  • Ignoring dependency failure — external services, identity, DNS, and databases often determine actual availability.
  • Using average load for capacity planning — average demand does not validate peak behavior.
  • Ignoring provider quotas — autoscaling cannot exceed provider or account limits.
  • Treating backup as recovery — recovery requires restoration, validation, dependencies, and operational procedures.
  • Failing to test rollback — rollback procedures often fail because data and schema behavior were not considered.
  • Accepting permanent exceptions — every exception should have an expiration and remediation owner.
  • Allowing schedule pressure to define risk — an announced date does not reduce security, recovery, or support obligations.
  • Excluding cost from readiness — production workloads require accountable budgets, cost allocation, and anomaly detection.
  • Ignoring logging cost during launch — temporary increased logging may create significant financial impact.
  • Failing to define hypercare exit — project teams may remain indefinitely responsible for production support.
  • Failing to perform a post-launch review — production telemetry and incidents should improve future readiness decisions.
  • Using AI-generated readiness assessments without evidence — AI may infer controls that do not exist, misunderstand workload criticality, or overlook provider-specific risks. Every material conclusion requires evidence and qualified human validation.

Security Considerationsprotected

  • Production-readiness materials may contain sensitive information such as architecture, network ranges, DNS names, security controls, vulnerabilities, identity roles, administrative paths, data classifications, recovery procedures, incident contacts, vendor dependencies, cost details, customer information, launch dates, and known weaknesses.
  • Before sharing information with an AI system: remove credentials.
  • Remove tokens.
  • Remove private keys.
  • Remove active connection strings.
  • Remove production secrets.
  • Sanitize customer and employee data.
  • Sanitize vulnerability details where required.
  • Sanitize network and identity details where required.
  • Follow source-code handling requirements.
  • Follow incident and vulnerability disclosure procedures.
  • Confirm the AI platform is approved.
  • Confirm retention and model-training settings.
  • Restrict distribution of generated assessments.
  • Do not use AI-generated approval language as a substitute for formal human authorization.

Brian Diamond

Founder, BrianOnAI

Twenty-five years designing, operating, and governing enterprise infrastructure — from MSP operations across dozens of client environments to enterprise infrastructure leadership. This blueprint codifies the operating model he's implemented in production, not theory.

⚠ Normalization Warnings — 12 for review

  • RESTRUCTURE: The document contains ~50+ domain readiness H1 sections (Business Readiness through Post-Launch Validation). To avoid a flat 50+ section render, all were grouped under one body/group 'Readiness Domains'. Confirm grouping strategy — an alternative would split into thematic sub-groups (e.g., Security & Identity, Data, Reliability & Observability, Operations & Change, Cost & Compliance).
  • CLASSIFICATION TO CONFIRM: 'Prerequisites' classified as a checklist TOOL 'Prerequisites Checklist' (gather-before-start items). Alternative: body/prose.
  • CLASSIFICATION TO CONFIRM: 'Evidence Register', 'Production Readiness Scorecard', and 'Responsibility Matrix' classified as matrix TOOLS (columnar, filled-in). Scorecard and Responsibility Matrix have doc-provided example_rows verbatim; Evidence Register has no rows in the doc (skeleton). Confirm.
  • CLASSIFICATION TO CONFIRM: 'Workload Criticality', 'Readiness Decision Categories', 'Mandatory Versus Advisory Controls', 'Blocker Classification', and 'Production Readiness Maturity Model' classified as body/reference (consulted tiered models). Confirm none should be tools.
  • RESTRUCTURE: 'Primary Prompt' and 'Follow-Up Prompts' (two source H1s, 25 prompts total) combined into one prompt_pack tool. Prompt text is verbatim; 'when' guidance lines are editorial additions.
  • CLASSIFICATION TO CONFIRM: 'Go/No-Go Criteria' appears twice — once as a reference-style prose body section (kept as body/prose with the doc's example criteria) and again derived into a 'Go/No-Go Criteria' template TOOL. The doc's 'Example Go/No-Go Criteria' remains a body/example. Confirm no duplication concern.
  • TOOL DERIVATION: 'Launch Runbook', 'Rollback Plan', 'Risk-Acceptance Record', and 'Go/No-Go Criteria' were derived into template TOOLS from the doc's fill-in field lists (Launch Runbook / Rollback Readiness / Risk Acceptance / Go-No-Go Criteria prose sections). Their body/prose counterparts were kept as reference within Readiness Domains. Confirm this reference-plus-tool split is desired rather than one or the other.
  • MATRIX: 'Evidence Register' rubric left empty (rubric:'') because the doc specifies fields but no scoring/usage scheme. skeleton_rows:0 per empty-skeleton rule.
  • PLAYBOOK: 'Implementation Roadmap' Phases 1–5 mapped to playbook.roadmap using phase names as horizons (the doc uses named phases, not time horizons). quick_wins intentionally empty — the doc has no Quick Wins section.
  • EXAMPLES: 'Example Workload Input', 'Example Executive Recommendation', 'Example Production Readiness Findings', 'Example Go/No-Go Criteria', 'Example Launch Runbook Extract', and 'Example Risk Register' all classified body/example. Tables preserved as HTML tables.
  • STATS: prompts=25 (1 primary + 24 follow-ups). deliverables=10 counts the distinct tools. Confirm counting convention.
  • OVERLAY: headline, subhead, teaser, exec_brief, and insights are written per voice rules and are not from the doc verbatim — flagged for overlay diff review.

SEO Block

  • Title tag: Cloud Production Readiness Review | ABME (40 chars)
  • Meta: Run an evidence-based production-readiness review for any cloud workload — launch blockers, go/no-go criteria, risk acceptance, and operational handover included. (162 chars)
  • Schema: HowTo · noindex: false
  • Related: cl-001, cl-002, cl-003, cl-004, cl-005, cl-006, cl-007, cl-008, cl-009, ai-003, ai-004, ai-007, ai-008, ai-009, ai-010
  • Keywords: cloud production readiness, production readiness review, go no-go criteria, launch runbook, operational acceptance, cloud workload launch gate, production readiness scorecard, risk acceptance record, hypercare plan, rollback plan, disaster recovery readiness, cloud launch governance
Copied