Stop Confusing "It Deploys" with "It's Ready" — Evidence-Based Production Launch for Cloud Workloads
An AI-assisted workflow to run a risk-based production-readiness review, surface launch blockers, classify residual risk, and produce an operational acceptance package before real users arrive.
Executive Brief
Your Challenge
Your workload deploys, the pipeline is green, and the scanner reports no criticals — so the business wants to launch. But none of that proves the workload can support real users, be operated at 2 a.m., recover from a regional outage, or stay inside its budget. Production readiness is an operational-risk decision, not a project milestone, and treating launch as a deadline instead of a governance gate is how workloads enter production with no owner, untested alerts, and unverified backups.
Common Obstacles
The failures cluster in predictable ways. Reviews happen too late, when readiness gaps are expensive and politically hard to fix. Checklists get treated as approval, hiding the fact that a single critical blocker can outweigh a high overall score. Backup gets mistaken for recovery, metrics get mistaken for monitoring, and healthy infrastructure gets mistaken for working business journeys. Underneath all of it: project teams self-approve risk they have no authority to accept, and schedule pressure quietly redefines what "ready" means.
The ABME Approach
This workflow does it in the right order: classify workload criticality, define mandatory versus advisory controls, assemble an evidence register, then assess readiness domain by domain — business, architecture, security, identity, data, reliability, observability, recovery, operations, cost, and compliance. Every finding is classified as a launch blocker, conditional blocker, high-priority risk, improvement, or observation, with named owners and completion dates. The prompt pack drives the assessment; go/no-go criteria, a launch runbook, a rollback plan, a hypercare plan, and formal risk acceptance turn the review into an auditable production decision that a qualified human signs.
Insight Summary
A workload is production ready when the organization can operate it, secure it, observe it, recover it, fund it, and accept its residual risks — not merely when the application can be deployed.
Criticality determines control depth. Applying Tier 0 rigor to a Tier 3 service wastes effort; applying Tier 3 rigor to a Tier 0 service is how launches become incidents.
Treat missing evidence as unknown, never as passed. Absence of a control record is not proof the control exists — it is a gap in the launch decision.
A configured backup is not a recovery capability, a collected metric is not monitoring, and a successful deployment is not a ready service. Each requires validation, not assumption.
A single unresolved critical blocker may override a high readiness score — the scorecard supports judgment, it does not replace risk authority.
Project managers and engineers should not accept enterprise-level business risk; residual risk requires acceptance by the authorized decision-maker, with an expiration date.
The Journey
Three phases; each lists the tools you'll use there.
Scope Criticality and Evidence Standards
- Gather the workload's architecture, ownership, and launch context
- Classify workload criticality using the tier model
- Separate mandatory controls from advisory improvements
- Establish the evidence register and evidence standards
- Confirm risk-acceptance and go/no-go authority
Assess Readiness Domain by Domain
- Run the primary prompt with the assembled evidence
- Assess each readiness domain and classify Ready through Not Ready
- Separate launch blockers, conditional blockers, and high-priority risks
- Populate the production-readiness scorecard
- Flag missing evidence as unknown, not passed
Decide, Launch, and Stabilize
- Define objective go/no-go criteria
- Produce the launch runbook and rollback plan
- Record risk acceptance and time-bound exceptions
- Confirm operational acceptance of the service
- Run hypercare, validate post-launch, and complete the validation checklist
What's Inside the Execution Layer
Numbered deliverables grouped by phase. Membership unlocks every tool.
Prerequisites Checklist
- Assemble the full evidence base before assessing readiness
- Surface missing evidence early, before it blocks a launch decision
- Confirm ownership, cost, and recovery inputs are in hand
Gather as much of the following as possible:
Evidence Register
- Record and attribute every piece of readiness evidence
- Track evidence validity periods and review status
- Distinguish validated evidence from missing or expired evidence
| Evidence ID | Readiness domain | Description | Source | Owner | Date | Validity period | Result | Related control | Reviewer | Status |
|---|
Production Readiness Scorecard
- Summarize readiness status across all domains
- Weight domains by importance for the workload
- Support the go/no-go decision without letting a score hide a blocker
| Domain | Weight | Status |
|---|---|---|
| Business | 10% | Ready |
| Architecture | 10% | Ready with risk |
| Security | 15% | Ready |
| Identity | 10% | Ready |
| Network | 8% | Ready |
| Data | 10% | Ready |
| Reliability | 10% | Conditionally ready |
| Observability | 10% | Ready |
| Recovery | 10% | Not ready |
| Operations | 7% | Ready |
Responsibility Matrix
- Assign accountability for each readiness capability
- Clarify who is accountable versus responsible for go/no-go and risk acceptance
- Prevent project teams from self-approving enterprise risk
| Capability | Business Owner | Service Owner | Workload Team | Platform Team | Security | Operations |
|---|---|---|---|---|---|---|
| Business readiness | Accountable | Consulted | Informed | Informed | Informed | Informed |
| Technical readiness | Informed | Accountable | Responsible | Consulted | Consulted | Consulted |
| Security readiness | Informed | Consulted | Responsible | Supports | Accountable | Informed |
| Monitoring | Informed | Accountable | Responsible | Supports | Consulted | Responsible |
| Recovery | Consulted | Accountable | Responsible | Supports | Consulted | Responsible |
| Support model | Informed | Consulted | Consulted | Consulted | Informed | Accountable |
| Cost readiness | Accountable for budget | Responsible | Provides usage inputs | Provides platform costs | Informed | Informed |
| Go/no-go | Accountable for business acceptance | Responsible for service recommendation | Consulted | Consulted | Consulted | Consulted |
| Risk acceptance | Accountable at authorized level | Responsible for documenting | Consulted | Consulted | Consulted | Consulted |
| Operational acceptance | Informed | Responsible | Supports | Supports | Consulted | Accountable |
Production Readiness Prompt Pack
- Run a full evidence-based readiness assessment and launch recommendation
- Generate scorecards, blockers, go/no-go criteria, runbooks, and rollback plans
- Review individual domains — security, capacity, recovery, cost, and more
Primary Prompt
Start here with as much of the assembled evidence as you have.You are a senior cloud production-readiness architect, site reliability engineer, cloud security architect, operations leader, application architect, incident-response advisor, disaster recovery specialist, FinOps practitioner, and technology risk advisor. I will provide some or all of the following: • Business objectives • Workload description • Criticality • Launch date • Architecture • Data flows • Dependency map • Service inventory • Ownership • Test results • User acceptance • Security assessment • Vulnerability findings • Penetration-test results • Identity design • Network design • Data classification • Compliance requirements • Performance baseline • Load-test results • Capacity model • Service-level objectives • Monitoring dashboards • Alert definitions • Logging design • Incident-response plan • Backup configuration • Restore-test results • Disaster recovery design • Recovery-test results • Deployment pipeline • Release plan • Rollback plan • Change record • Support model • Runbooks • Cost estimate • Budget approval • Risk register • Exception register • Vendor support • Communication plan • Hypercare plan Your task is to perform a production-readiness assessment and produce an evidence-based launch recommendation. Do not assume a successful deployment means the workload is production ready. Do not approve production based solely on checklist completion. First: 1. Summarize: • Business capability • Users and customers • Workload criticality • Launch scope • Launch date • Service-level expectations • Support expectations • Regulatory or contractual obligations 2. Separate: • Confirmed facts • Validated evidence • Reported observations • Inferences • Assumptions • Unknowns 3. Identify missing evidence that materially affects the production decision. 4. Identify: • Business owner • Service owner • Technical owner • Data owner • Security owner • Support owner • Budget owner • Risk-acceptance authority 5. Assess readiness across: • Business • Ownership • Architecture • Dependencies • Identity • Network • DNS • Certificates • Security • Vulnerabilities • Secrets • Encryption • Data • Database • Functional testing • Integration testing • Performance • Capacity • Reliability • Observability • Logging • Alerts • Incident response • Backup • Recovery • Disaster recovery • Deployment • Rollback • Operations • Support • Change management • Cost • Compliance • Documentation • User readiness • Hypercare 6. Classify each domain as: • Ready • Ready with accepted risk • Conditionally ready • Not ready • Unable to assess • Not applicable 7. Separate: • Launch blockers • Conditional blockers • High-priority risks • Improvements • Observations 8. For each finding provide: • Finding ID • Domain • Classification • Severity • Confidence • Evidence • Business impact • Technical impact • Security impact • Operational impact • Recommended action • Owner • Required completion date • Validation • Launch impact • Risk approver where applicable 9. Evaluate whether: • Critical business journeys pass • Critical dependencies are validated • Capacity meets baseline and peak demand • Scaling limits are understood • Monitoring covers service health and business outcomes • Alerts route to named responders • Backup is configured • Restore has been tested • Recovery objectives are approved • Recovery has been tested • Incident response is operational • Support has accepted the service • Rollback is practical • Cost is approved and attributable • Compliance obligations are satisfied • Residual risks are properly accepted Requirements: • A single critical blocker may override a high readiness score. • Do not treat missing evidence as evidence of readiness. • Do not treat configured backup as successful recovery. • Do not treat scanner output as complete security assurance. • Do not treat architecture documentation as proof of implementation. • Do not recommend launch solely because of schedule pressure. • Require named owners for all launch conditions. • Require expiration dates for exceptions. • Identify irreversible changes. • Define rollback triggers and decision authority. • Define objective go/no-go criteria. • Define hypercare and exit criteria. • Define post-launch validation. • Identify where human testing, security review, legal review, compliance approval, recovery testing, capacity testing, or business acceptance is still required. • State when evidence is insufficient. • Require qualified human approval for production launch. Then produce: 1. Executive summary. 2. Production recommendation. 3. Business and workload context. 4. Criticality assessment. 5. Ownership assessment. 6. Evidence register. 7. Production-readiness scorecard. 8. Launch blockers. 9. Conditional blockers. 10. Accepted risks. 11. Business readiness. 12. Architecture and dependency readiness. 13. Identity and network readiness. 14. Security readiness. 15. Data and database readiness. 16. Functional and integration readiness. 17. Performance and capacity readiness. 18. Reliability readiness. 19. Observability and logging readiness. 20. Incident-response readiness. 21. Backup and recovery readiness. 22. Deployment and rollback readiness. 23. Operational-support readiness. 24. Cost readiness. 25. Compliance readiness. 26. Documentation and user readiness. 27. Go/no-go criteria. 28. Launch runbook. 29. Hypercare plan. 30. Post-launch validation. 31. Risk register. 32. Exception register. 33. Responsibility matrix. 34. Open questions. 35. Required approvals. 36. Final recommendation.
Perform a Production Readiness Review
For a focused readiness classification pass.Assess this workload for production readiness. Classify each domain as: • Ready • Ready with accepted risk • Conditionally ready • Not ready • Unable to assess • Not applicable Identify blockers, missing evidence, accepted risks, owners, and required completion dates.
Create a Production Readiness Checklist
To tailor a checklist to a specific workload.Create a risk-based production-readiness checklist for this workload. Tailor the checklist to: • Criticality • Architecture • Data classification • User population • Recovery requirements • Compliance • Support model • Launch method Separate mandatory controls from advisory improvements.
Create a Production Readiness Scorecard
To generate a weighted scorecard.Create a weighted production-readiness scorecard. Include: • Business • Ownership • Architecture • Security • Identity • Network • Data • Reliability • Performance • Observability • Recovery • Operations • Cost • Compliance Define how critical blockers override the score.
Identify Launch Blockers
To isolate what must be resolved before launch.Review the supplied evidence and identify production launch blockers. For each blocker include: • Evidence • Business impact • Technical impact • Required remediation • Owner • Validation • Completion deadline • Go/no-go impact Do not classify missing evidence as passed.
Create Go/No-Go Criteria
To define objective launch gates.Create objective production go/no-go criteria. Include: • Business acceptance • Testing • Security • Data • Capacity • Monitoring • Alerts • Backup • Recovery • Support • Rollback • Cost • Compliance • Change approval • Required decision-makers
Review Service Ownership
To confirm no ownership role is missing or temporary.Review production ownership for this service. Confirm: • Business owner • Service owner • Technical owner • Application owner • Data owner • Security owner • Support owner • Budget owner • Vendor owner • Risk approver Identify missing, temporary, or conflicting ownership.
Review Production Architecture
For an architecture readiness pass.Review this architecture for production readiness. Assess: • Failure domains • Dependencies • Single points of failure • Scalability • Service limits • Regional design • Security boundaries • Data flows • Recovery • Supportability • Operational complexity • Cost Separate launch blockers from future improvements.
Review Dependency Readiness
To assess each production dependency.Review all production dependencies. For each dependency assess: • Owner • Service level • Authentication • Connectivity • Capacity • Failure behavior • Timeout • Retry • Monitoring • Support • Recovery • Fallback • Validation evidence
Review Security Readiness
For a security readiness pass.Assess security readiness for production. Review: • Threat model • Identity • Privileged access • Network exposure • Secrets • Encryption • Vulnerabilities • Logging • Threat detection • Incident response • Security exceptions • Data protection Identify findings that should block launch.
Review Performance and Capacity
To validate capacity assumptions.Assess performance and capacity readiness. Review: • Baseline load • Peak load • Growth • Response time • Throughput • Error rate • Concurrency • Resource utilization • Scaling • Quotas • Provider limits • Dependency limits • Cost Identify missing tests and unsafe assumptions.
Review Observability Readiness
For an observability pass.Assess observability readiness. Review: • Service-health metrics • Application metrics • Infrastructure metrics • Dependency monitoring • Business metrics • Logs • Traces • Dashboards • Synthetic tests • Alert ownership • Runbooks • Escalation • Alert testing • Cost controls
Review Alert Quality
To audit the alert catalog.Review the production alert catalog. For each alert assess: • Signal • Threshold • Duration • Severity • Business impact • Owner • Routing • Runbook • Escalation • False-positive risk • Maintenance behavior • Test evidence Identify noisy, missing, or unactionable alerts.
Review Backup and Recovery Readiness
To verify recovery, not just backup.Assess backup and recovery readiness. Review: • Backup scope • Frequency • Retention • Encryption • Immutability • Geographic protection • Failure alerts • Restore permissions • Restore-test evidence • RTO • RPO • Recovery runbook • Recovery dependencies • Failover • Failback Do not treat successful backup jobs as proof of recovery.
Create an Operational Handover Package
To transfer the service to operations.Create an operational handover package. Include: • Service definition • Ownership • Architecture • Dependencies • Support model • Monitoring • Alerts • Runbooks • Backup • Recovery • Access • Deployment • Incident response • Vendor contacts • Known issues • Risks • Cost ownership • Training • Acceptance criteria
Create a Launch Runbook
To build the step-by-step launch sequence.Create a detailed production launch runbook. For each step include: • Step number • Planned time • Owner • Preconditions • Action • Expected result • Validation • Failure response • Rollback impact • Evidence • Status Include explicit go/no-go checkpoints.
Create a Rollback Plan
To define how the launch is reversed.Create a production rollback plan. Define: • Trigger • Decision authority • Decision deadline • Application rollback • Infrastructure rollback • Database rollback • Data reconciliation • DNS rollback • Configuration rollback • User communication • Evidence • Forward-fix criteria • Irreversible actions
Review Database Change Readiness
When a database change is part of the launch.Assess the production readiness of this database change. Review: • Compatibility • Schema changes • Data migration • Locking • Duration • Backward compatibility • Application dependencies • Backup • Rollback • Validation • Monitoring • Performance • Maintenance window
Create a Hypercare Plan
To plan the stabilization period after launch.Create a production hypercare plan. Define: • Duration • Staffing • Support channel • Monitoring • Status cadence • Defect handling • Business reconciliation • Vendor support • Cost review • Escalation • Exit criteria • Transition to normal operations
Create a Risk-Acceptance Record
To formally document an accepted residual risk.Create a formal risk-acceptance record. Include: • Risk • Evidence • Business impact • Technical impact • Probability • Compensating controls • Owner • Approver • Acceptance period • Expiration • Remediation plan • Target date • Review cadence Do not accept risk without appropriate authority.
Create an Exception Register
To track time-bound control exceptions.Create a production-readiness exception register. For each exception include: • Control • Justification • Risk • Compensating control • Owner • Approver • Start date • Expiration • Remediation • Validation • Status
Review Cost Readiness
For a financial readiness pass.Assess financial readiness for production. Review: • Cost estimate • Budget • Cost center • Budget owner • Tags • Forecast • Scaling cost • Logging cost • Backup cost • Disaster recovery cost • Support cost • Licensing • Marketplace charges • Budget alerts • Anomaly detection • Unit economics
Create a Post-Launch Validation Plan
To define what is monitored after go-live.Create a post-launch validation plan. Monitor: • Availability • Performance • Errors • Traffic • Capacity • Scaling • Business transactions • Data quality • Integrations • Security alerts • Backup • Cost • Support tickets • User feedback Define owners, thresholds, cadence, and escalation.
Perform a Post-Launch Review
After stabilization, to close the loop.Perform a post-launch review. Assess: • Launch execution • Incidents • Defects • Performance • Capacity • User impact • Data reconciliation • Security • Monitoring • Alert quality • Support • Cost • Hypercare • Outstanding risks • Lessons learned Recommend changes to the production-readiness standard.
Go/No-Go Criteria
- Define objective launch gates before the review meeting
- Separate go conditions from disqualifying no-go conditions
- Ensure required decision-makers and approvals are captured
Go Criteria
No-Go Criteria
Required Decision-Makers
Launch Runbook
- Sequence the launch with owners and timings
- Define validation and failure response for each step
- Embed explicit go/no-go checkpoints in the launch flow
Step Number
Planned Time
Owner
Action
Preconditions
Expected Result
Validation
Failure Response
Evidence
Completion Status
Rollback Plan
- Define objective rollback triggers and decision authority
- Sequence application, infrastructure, and data rollback steps
- Identify irreversible actions and forward-fix criteria
Rollback Trigger
Decision Authority
Maximum Decision Time
Application Rollback
Infrastructure Rollback
Database Rollback
DNS Rollback
Configuration Rollback
User Communication
Evidence
Forward-Fix Criteria
Irreversible Actions
Risk-Acceptance Record
- Document residual risks accepted at launch
- Assign a named approver with appropriate authority
- Set expiration, remediation, and review cadence for each risk
Risk ID
Description
Business Impact
Technical Impact
Probability
Compensating Control
Owner
Risk Approver
Acceptance Date
Expiration
Remediation Plan
Target Date
Review Cadence
Validation Checklist
- Verify readiness across business, security, data, operations, and cost
- Confirm operational acceptance and go/no-go authority before launch
- Provide auditable evidence that each domain was checked
Confirm the following before production launch:
Business and Ownership
Architecture and Dependencies
Identity and Network
Security
Data and Database
Testing
Reliability and Observability
Operations and Release
Financial and Governance
Launch and Hypercare
🔒 The full execution layer — every checklist, matrix, and the prompt pack — is included with ABME membership.
Unlock Full BlueprintFull Playbook
Overviewpublic
A cloud workload is not ready for production merely because:
- The application deploys.
- Infrastructure exists.
- A test page loads.
- A pipeline succeeds.
- A security scanner reports no critical findings.
- A business owner wants to launch.
- A migration deadline has arrived.
Production readiness is the evidence-based determination that a workload can safely support real users, business transactions, operational support, security obligations, recovery requirements, and financial accountability.
A production-ready workload should be:
- Functionally correct
- Secure
- Observable
- Supportable
- Recoverable
- Scalable
- Resilient
- Compliant
- Cost-accountable
- Operationally owned
- Governed through controlled change
The production-readiness process should answer: “Do we have enough evidence to place this workload into production, operate it responsibly, detect and respond to failure, recover it when necessary, and understand the residual risks?”
A production review should not become a checklist ceremony performed at the end of a project. Readiness should be built throughout the delivery lifecycle. Late reviews frequently uncover:
- Missing ownership
- Incomplete monitoring
- Weak alert routing
- Unvalidated backup
- Unclear recovery objectives
- Excessive permissions
- Incomplete testing
- Missing documentation
- Uncontrolled deployment access
- Unbudgeted cloud services
- Unsupported dependencies
- Undefined incident response
- No rollback plan
- No operational handover
This workflow uses AI to help structure a production-readiness assessment, evaluate evidence, identify blockers, classify residual risk, prepare launch criteria, and create an operational acceptance package.
The objective is not to eliminate all risk. The objective is to ensure risk is known, controlled, assigned, and accepted at the appropriate level before production use.
Business Problempublic
Organizations often treat production launch as a project-management milestone rather than an operational-risk decision. This can lead to workloads entering production with:
- No service owner
- No support model
- No service-level objectives
- Missing dashboards
- Untested alerts
- Unverified backups
- Unrehearsed recovery
- Incomplete security controls
- Missing cost ownership
- Undocumented dependencies
- Unapproved exceptions
- Weak rollback procedures
- No hypercare plan
- No user communication
- No capacity evidence
- No post-launch review
When these gaps are discovered after launch:
- Incidents last longer.
- Users lose confidence.
- Support teams improvise.
- Recovery attempts fail.
- Security findings become emergencies.
- Cloud costs exceed forecasts.
- Project teams remain permanent support teams.
- Audit evidence is incomplete.
- Technical debt becomes operational debt.
A strong production-readiness process establishes objective launch gates and requires accountable decision-makers to accept residual risk.
Typical Use Casespublic
Use this workflow when:
- Launching a new cloud application
- Moving a migrated workload into production
- Promoting a major application release
- Launching a new managed service
- Deploying a Kubernetes workload
- Launching a public API
- Launching an internal enterprise service
- Moving from pilot to general availability
- Launching a regulated workload
- Replacing a critical legacy system
- Preparing for a high-traffic event
- Launching in a new region
- Completing an operational handover
- Reviewing a troubled preproduction environment
- Creating a standardized production-readiness review
- Establishing a production launch gate
- Preparing executive launch approval
- Preparing audit evidence
- Creating a reusable service-acceptance process
Do NOT Use This Workflow Whenpublic
This workflow is not intended to:
- Replace application testing
- Replace penetration testing
- Replace architecture review
- Replace legal or compliance approval
- Replace business acceptance
- Guarantee zero incidents
- Approve a launch solely because a checklist is complete
- Treat documentation as proof of operational capability
- Treat a backup configuration as proof of recoverability
- Treat a successful deployment as proof of service readiness
- Treat scanner output as complete security assurance
- Accept unresolved critical risks without appropriate authority
- Launch because a deadline exists
- Allow project teams to self-approve all risk
- Automatically deploy AI-generated changes
- Store credentials or production secrets in prompts
- Use generic acceptance criteria for every workload
- Ignore workload criticality
- Ignore customer and user impact
- Ignore financial accountability
Production approval must remain a human governance decision supported by evidence.
Expected Outcomepublic
After completing this workflow, you should have:
- Production-readiness scope
- Business and technical ownership
- Workload criticality classification
- Production-readiness scorecard
- Evidence register
- Functional readiness assessment
- Architecture readiness assessment
- Security readiness assessment
- Identity readiness assessment
- Network readiness assessment
- Data readiness assessment
- Reliability readiness assessment
- Capacity and performance assessment
- Observability assessment
- Backup and recovery assessment
- Operational support assessment
- Incident response assessment
- Change and release assessment
- Cost and financial assessment
- Compliance assessment
- Dependency assessment
- Launch runbook
- Rollback plan
- Go/no-go criteria
- Risk register
- Exception register
- Hypercare plan
- Operational acceptance
- Post-launch validation plan
- Executive launch recommendation
🔒 The complete playbook — reference models, worked examples, and operational guidance — is included with ABME membership.
Unlock Full BlueprintProduction Readiness Objectivesprotected
A complete production-readiness review should answer:
- What business capability is being launched?
- Who owns the service?
- Who supports the service?
- What is the workload criticality?
- Who are the users and customers?
- What service levels are expected?
- What dependencies exist?
- What happens when a dependency fails?
- What security controls are required?
- What data is processed?
- What compliance obligations apply?
- What capacity has been validated?
- What performance is acceptable?
- How will the service be monitored?
- Which alerts require action?
- How will incidents be handled?
- What is backed up?
- How will recovery occur?
- What changes are allowed at launch?
- How will rollback work?
- What costs are expected?
- Who owns the budget?
- Which risks remain?
- Who may accept those risks?
- What conditions block launch?
- What evidence is required?
- What happens during hypercare?
- When does operations formally accept the service?
- How will launch success be measured?
- When will the post-launch review occur?
Production Readiness Principlesprotected
Recommended principles include:
- Readiness is evidence-based.
- Criticality determines control depth.
- Production ownership must be explicit.
- Security is required before launch.
- Recovery must be tested.
- Monitoring must be actionable.
- Alerts need owners.
- Capacity must be demonstrated.
- Dependencies must be understood.
- Rollback must be defined.
- Exceptions must be time-bound.
- Residual risks require named acceptance.
- Cost ownership is part of readiness.
- Operational handover is a formal event.
- Launch success includes stabilization.
- Lessons learned should improve the standard.
Workload Criticalityprotected
Tier 0 — Mission Critical
- Highly available architecture
- Strong recovery
- Continuous monitoring
- Formal incident command
- Extensive testing
- Senior risk approval
Tier 1 — Business Critical
- Defined RTO and RPO
- Redundancy
- Strong monitoring
- Tested recovery
- Formal support
- Controlled deployment
Tier 2 — Important
- Standard backup
- Business-hours support
- Defined restoration
- Basic redundancy
- Standard monitoring
Tier 3 — Standard
Readiness Decision Categoriesprotected
Ready
Ready with Accepted Risk
Conditionally Ready
Not Ready
Unable to Assess
Evidence Standardsprotected
Evidence may include:
- Test results
- Architecture diagrams
- Configuration exports
- Security scans
- Penetration-test reports
- Performance-test results
- Recovery-test evidence
- Monitoring screenshots
- Alert tests
- Change records
- Deployment logs
- Cost estimates
- Budget approvals
- Compliance approvals
- Support runbooks
- Business sign-off
- Incident-response exercises
- Dependency maps
Evidence should be:
- Current
- Relevant
- Reproducible
- Attributable
- Retained
- Approved where required
Mandatory Versus Advisory Controlsprotected
Mandatory Controls
- Critical security vulnerability
- Missing business owner
- Missing data-protection control
- No recovery capability for a critical workload
- Unsupported production dependency
- No rollback for a high-risk launch
- No monitoring for critical service health
Advisory Controls
Blocker Classificationprotected
Launch Blocker
Conditional Blocker
High-Priority Risk
Improvement
Observation
Readiness Domainsprotected
Business Readiness
Confirm:
- Business objective
- Business owner
- Product owner
- User population
- Customer impact
- Launch date
- Business calendar
- Communication plan
- Acceptance criteria
- Training
- Support expectations
- Financial approval
- Legal and compliance approval
Service Ownership
Every production service should have:
- Business owner
- Service owner
- Technical owner
- Application owner
- Data owner
- Security owner
- Support owner
- Budget owner
- Vendor owner where applicable
Ownership should not be assigned only to a temporary project team.
Service Definition
Document:
- Service name
- Purpose
- Customers
- Users
- Criticality
- Business hours
- Support hours
- Service-level targets
- Dependencies
- Data classification
- Deployment model
- Recovery tier
- Cost center
- Escalation path
Architecture Readiness
Validate:
- Approved target architecture
- Component inventory
- Data flows
- Trust boundaries
- Failure domains
- Dependency behavior
- Scalability
- High availability
- Service limits
- Provider quotas
- Regional design
- Technology support
- Lifecycle
- Technical debt
- Known constraints
Architecture Review Questions
- Does the design meet business objectives?
- Are all critical dependencies identified?
- Are single points of failure known?
- Are failure modes understood?
- Is resilience proportional to criticality?
- Are service limits understood?
- Are provider-region constraints known?
- Is the design supportable by available staff?
- Are architecture exceptions documented?
- Is the design unnecessarily complex?
Dependency Readiness
For each dependency document:
- Dependency name
- Owner
- Service level
- Authentication
- Network path
- Failure mode
- Timeout
- Retry behavior
- Capacity
- Monitoring
- Support contact
- Recovery expectation
- Alternative or fallback
- Validation evidence
Third-Party Dependency Review
Assess:
- Vendor support
- Contract
- Service-level agreement
- Security review
- Privacy review
- Data location
- Integration limits
- Rate limits
- Maintenance windows
- Incident notification
- Exit strategy
- Financial exposure
Identity Readiness
Validate:
- User authentication
- Single sign-on
- Multi-factor authentication
- Conditional access
- Role mapping
- Least privilege
- Privileged access
- Service identities
- Workload identities
- Guest access
- Emergency access
- Access reviews
- Joiner, mover, and leaver process
- Authentication logging
Privileged Access
Confirm:
- Named privileged users
- Separate administrative accounts
- Just-in-time access
- Approval
- Session logging
- MFA
- Emergency access
- Periodic review
- Credential rotation
- No shared administrator credentials
Workload Identity
For each service identity verify:
- Owner
- Purpose
- Permission scope
- Authentication method
- Credential lifetime
- Rotation
- Logging
- Dependency
- Failure behavior
- Decommissioning
Prefer short-lived managed or federated identities over static credentials.
Network Readiness
Validate:
- Addressing
- Routing
- Segmentation
- Firewalls
- Security groups
- Load balancers
- Private endpoints
- Internet ingress
- Internet egress
- DNS
- Certificates
- DDoS protection
- Hybrid connectivity
- Remote administration
- Flow logging
- Network monitoring
- Capacity
- Failover
Internet Exposure
For every public endpoint document:
- Business purpose
- Protocol
- Port
- Authentication
- TLS
- Web application firewall
- DDoS protection
- Rate limiting
- Logging
- Owner
- Vulnerability testing
- Certificate renewal
- Abuse monitoring
DNS Readiness
Confirm:
- Production records
- Ownership
- TTL
- Private versus public resolution
- Health checks
- Failover behavior
- Certificate names
- Monitoring
- Change process
- Rollback records
- Expiration and renewal
Certificate Readiness
Validate:
- Certificate authority
- Subject names
- Expiration
- Renewal automation
- Private-key protection
- Load-balancer configuration
- Client trust
- Revocation
- Alerting
- Emergency replacement
Security Readiness
Review:
- Threat model
- Security architecture
- Vulnerability findings
- Penetration testing
- Dependency scanning
- Configuration scanning
- Secrets
- Encryption
- Logging
- Threat detection
- Identity
- Network exposure
- Data protection
- Security exceptions
- Incident response
Security Findings
Classify findings as:
- Critical blocker
- High risk
- Medium risk
- Low risk
- Informational
A severity score alone should not determine launch readiness. Consider:
- Exploitability
- Exposure
- Data impact
- Business impact
- Compensating controls
- Detection
- Recovery
- Time to remediation
Vulnerability Readiness
Confirm:
- Operating-system scanning
- Container-image scanning
- Dependency scanning
- Source-code scanning
- Infrastructure scanning
- Runtime scanning where applicable
- Remediation ownership
- Exception process
- Rescan evidence
- Vulnerability age
Secret Readiness
Confirm:
- No secrets in source control
- No secrets in application images
- No secrets in build logs
- Approved secret store
- Access control
- Rotation
- Expiration
- Audit logging
- Emergency rotation
- Ownership
Encryption Readiness
Validate:
- Data at rest
- Data in transit
- Internal service communication
- Backup encryption
- Key ownership
- Key rotation
- Key recovery
- Certificate management
- Customer-managed key requirements
- Hardware security module requirements
Data Readiness
Assess:
- Data classification
- Data owner
- Data location
- Data residency
- Data flow
- Retention
- Archival
- Deletion
- Legal hold
- Privacy
- Backup
- Recovery
- Masking
- Production test-data restrictions
- Data quality
- Data reconciliation
Database Readiness
Validate:
- Engine and version
- Support status
- High availability
- Backup
- Point-in-time recovery
- Maintenance
- Patching
- Capacity
- Performance
- Connection limits
- Authentication
- Encryption
- Monitoring
- Query performance
- Schema migration
- Rollback
- Licensing
Data Migration Readiness
For migrated data confirm:
- Source of truth
- Migration completed
- Counts reconciled
- Checksums or validation
- Business totals reconciled
- In-flight transactions handled
- Data quality issues documented
- Rollback implications understood
- Business owner approved
Functional Readiness
Validate:
- Core user journeys
- Business rules
- Error handling
- Notifications
- Reports
- Scheduled jobs
- Batch processes
- Administrative functions
- Integrations
- Accessibility
- Localization where required
- Browser or client compatibility
- Mobile behavior where applicable
Test Coverage
Testing may include:
- Unit testing
- Component testing
- Integration testing
- Contract testing
- End-to-end testing
- Regression testing
- User acceptance testing
- Performance testing
- Security testing
- Recovery testing
- Operational testing
- Rollback testing
Test Defect Review
For every unresolved defect document:
- Defect ID
- Severity
- Business impact
- Technical impact
- Workaround
- Owner
- Target resolution
- Launch impact
- Risk approver
- Acceptance status
Performance Readiness
Validate:
- Response time
- Throughput
- Concurrency
- Batch duration
- Queue depth
- Database performance
- Network latency
- Storage performance
- Error rate
- Scaling
- Failover performance
- Resource saturation
- Performance under degraded dependencies
Capacity Model
Document:
- Expected baseline
- Expected peak
- Seasonal peak
- Growth assumption
- Capacity margin
- Scaling trigger
- Scaling limit
- Quota
- Provider limit
- Downstream capacity
- Cost impact
Load Testing
Load testing should:
- Reflect realistic behavior.
- Include representative data.
- Include critical integrations.
- Measure system limits.
- Test scaling.
- Test recovery after load.
- Identify bottlenecks.
- Validate alerting.
- Record cost.
Stress and Failure Testing
Where appropriate, test:
- Resource exhaustion
- Dependency latency
- Dependency failure
- Network interruption
- Database failover
- Node loss
- Zone loss
- Regional impairment
- Queue backlog
- Certificate failure
- Credential expiration
Reliability Readiness
Validate:
- Availability objectives
- Redundancy
- Health checks
- Failover
- Retry behavior
- Timeouts
- Circuit breakers
- Queue handling
- Graceful degradation
- Rate limits
- Idempotency
- Data consistency
- Maintenance behavior
- Recovery procedures
Service-Level Objectives
Define measurable objectives such as:
- Availability
- Latency
- Throughput
- Error rate
- Data freshness
- Completion time
- Recovery time
- Support response
Service-level objectives should align with business expectations.
Error Budgets
For mature services, define:
- SLO
- Error budget
- Measurement window
- Consumption rate
- Alert thresholds
- Release restrictions
- Escalation
- Review cadence
Observability Readiness
Production monitoring should cover:
- Availability
- Latency
- Errors
- Traffic
- Saturation
- Dependencies
- Business transactions
- Security
- Cost
- Backup
- Recovery
- Deployment health
Monitoring Layers
Include:
- Infrastructure monitoring
- Platform monitoring
- Application monitoring
- Database monitoring
- Network monitoring
- User-experience monitoring
- Synthetic monitoring
- Business-process monitoring
- Security monitoring
- Cost monitoring
Dashboard Readiness
Dashboards should provide:
- Service health
- Key SLOs
- Current incidents
- Error rates
- Latency
- Traffic
- Resource saturation
- Dependency health
- Deployment markers
- Business metrics
- Cost anomalies
Dashboards should support operational decisions, not merely display available metrics.
Alert Readiness
For every material alert define:
- Alert condition
- Threshold
- Duration
- Severity
- Owner
- Routing
- Business impact
- Runbook
- Escalation
- Suppression rules
- Maintenance behavior
- Validation evidence
Alert Testing
Test:
- Alert generation
- Notification
- On-call receipt
- Escalation
- Runbook access
- Acknowledgment
- Resolution
- Closure
- Maintenance suppression
An untested alert should not be assumed operational.
Logging Readiness
Confirm:
- Application logs
- Infrastructure logs
- Identity logs
- Security logs
- Administrative logs
- Network logs
- Database logs
- Audit logs
- Retention
- Access
- Redaction
- Time synchronization
- Correlation identifiers
- Searchability
- Cost controls
Sensitive Data in Logs
Prevent or control:
- Passwords
- Tokens
- Private keys
- Session identifiers
- Personal information
- Payment information
- Health information
- Confidential business data
- Connection strings
Incident Response Readiness
Confirm:
- On-call schedule
- Incident severity model
- Escalation
- Incident commander
- Technical responders
- Business contacts
- Security contacts
- Vendor contacts
- Communications
- Evidence retention
- Post-incident review
- Runbooks
Incident Scenarios
Prepare for:
- Application outage
- Database outage
- Regional impairment
- Identity failure
- Network failure
- Data corruption
- Security compromise
- Credential compromise
- Dependency outage
- Performance degradation
- Capacity exhaustion
- Cost spike
- Failed deployment
Runbook Readiness
Runbooks should exist for:
- Service restart
- Scaling
- Dependency failure
- Database failover
- Backup restore
- Certificate renewal
- Credential rotation
- Alert response
- Deployment rollback
- Access failure
- Queue backlog
- Regional failover
- Security containment
- Vendor escalation
Backup Readiness
Validate:
- Backup scope
- Frequency
- Retention
- Encryption
- Immutability
- Geographic protection
- Monitoring
- Ownership
- Failure alerting
- Restore permissions
- Backup cost
Recovery Readiness
Validate:
- RTO
- RPO
- Recovery architecture
- Recovery dependencies
- Recovery runbook
- Data recovery
- Identity recovery
- DNS recovery
- Network recovery
- Application recovery
- Communication
- Failover
- Failback
- Recovery testing
Restore Testing
Evidence should include:
- Restore date
- Data restored
- Environment
- Duration
- Integrity result
- Application validation
- RTO result
- RPO result
- Defects
- Owner
- Follow-up actions
Disaster Recovery Readiness
For workloads requiring disaster recovery, confirm:
- Recovery environment
- Infrastructure reproducibility
- Replication
- Capacity
- DNS
- Identity
- Network
- Secrets
- Keys
- Monitoring
- Security
- Runbook
- Decision authority
- Exercise results
- Failback
Operational Support Readiness
Confirm:
- Support hours
- On-call coverage
- Support tiers
- Escalation
- Service desk knowledge
- Support group
- Ticket routing
- Vendor support
- Runbooks
- Known-error records
- Operational training
- Support acceptance
Operational Handover
The handover package should include:
- Service definition
- Architecture
- Dependency map
- Support model
- Monitoring
- Alerts
- Runbooks
- Backup
- Recovery
- Access
- Deployment
- Incident process
- Vendor contacts
- Known issues
- Risks
- Cost ownership
- Documentation repository
Service Acceptance
Operations should formally accept the service after confirming:
- Ownership
- Access
- Monitoring
- Alerts
- Runbooks
- Backup
- Recovery
- Training
- Support contacts
- Escalation
- Known risks
- Documentation
- Staffing
Change Readiness
Confirm:
- Change policy
- Release process
- Deployment window
- Branch protection
- Approval
- Separation of duties
- Rollback
- Database-change procedure
- Feature-flag process
- Emergency change
- Evidence retention
Deployment Readiness
Validate:
- Artifact versioning
- Build reproducibility
- Artifact integrity
- Security scans
- Environment configuration
- Secret retrieval
- Infrastructure plan
- Database migrations
- Deployment sequence
- Health checks
- Rollback
- Post-deployment validation
- Ownership
Release Strategy
Possible approaches include:
- Standard rolling deployment
- Blue-green
- Canary
- Feature flags
- Phased user rollout
- Regional rollout
- Shadow traffic
Select based on:
- Risk
- Architecture
- Data compatibility
- User impact
- Rollback capability
- Operational maturity
Database Change Readiness
Database changes should define:
- Compatibility
- Migration sequence
- Locking risk
- Long-running operations
- Backward compatibility
- Rollback
- Data backup
- Validation
- Application release dependency
- Maintenance window
Rollback Readiness
Define:
- Rollback trigger
- Decision authority
- Maximum decision time
- Application rollback
- Infrastructure rollback
- Database rollback
- Data reconciliation
- DNS rollback
- Configuration rollback
- User communication
- Evidence
- Forward-fix criteria
Irreversible Changes
Identify:
- Data transformations
- Destructive schema changes
- Key rotation
- Source-system retirement
- License transfer
- Contract change
- Endpoint removal
- Data deletion
- Permanent migration steps
Require explicit approval.
Financial Readiness
Validate:
- Cost estimate
- Budget
- Cost center
- Budget owner
- Tagging
- Forecast
- Scaling cost
- Logging cost
- Backup cost
- Disaster recovery cost
- Support cost
- Marketplace charges
- License cost
- Anomaly alerts
- Cost review cadence
Cost Guardrails
Potential controls include:
- Budget alerts
- Forecast alerts
- Anomaly detection
- Scaling limits
- Sandbox expiration
- Log-volume alerts
- Storage-growth alerts
- High-cost resource approval
- Commitment review
- Unit-cost monitoring
Unit Economics
Where applicable define:
- Cost per customer
- Cost per user
- Cost per transaction
- Cost per API call
- Cost per device
- Cost per report
- Cost per workload
- Cost per revenue dollar
Compliance Readiness
Assess applicable obligations such as:
- Data privacy
- Data residency
- Retention
- Audit logging
- Access review
- Encryption
- Vulnerability management
- Incident notification
- Vendor management
- Business continuity
- Evidence retention
Only apply frameworks relevant to the workload.
Documentation Readiness
Required documentation may include:
- Service overview
- Architecture diagram
- Data-flow diagram
- Dependency map
- Ownership
- Support model
- Deployment procedure
- Rollback procedure
- Monitoring
- Alerts
- Runbooks
- Backup
- Recovery
- Security design
- Cost model
- Known issues
- Risk register
- Contact list
User Readiness
Confirm:
- User communication
- Training
- Support instructions
- Access
- Client requirements
- New URLs
- Maintenance notice
- Known limitations
- Feedback channel
- Accessibility
- Launch timing
Hypercare
A hypercare plan may include:
- Duration
- Extended staffing
- Dedicated support channel
- Increased monitoring
- Daily status
- Rapid defect triage
- Vendor support
- Business reconciliation
- Cost review
- Exit criteria
- Transition to normal operations
Hypercare Exit Criteria
Examples:
- No unresolved critical defects
- Incident rate within expected range
- Performance stable
- Data reconciled
- Support volume declining
- Monitoring effective
- Operations accepts steady-state support
- Business owner approves
- Cost within expected range
Post-Launch Validation
Validate after launch:
- Availability
- Performance
- Error rates
- User access
- Business transactions
- Data quality
- Integrations
- Security alerts
- Backup
- Cost
- Support tickets
- User feedback
- Capacity
- Scaling
Go/No-Go Criteriaprotected
Example go criteria:
- Business acceptance complete
- Critical tests pass
- No unresolved critical security findings
- Target environment stable
- Monitoring active
- Alerts tested
- Backup complete
- Recovery validated
- Support staffed
- Rollback available
- Change approved
- Budget approved
- Required decision-makers present
Example no-go criteria:
- Missing service owner
- Failed critical business test
- Data reconciliation failure
- Unsupported production dependency
- Critical vulnerability
- Missing recovery for critical workload
- No monitoring for key service health
- Rollback unavailable
- Operations has not accepted support
- Required compliance approval missing
Risk Acceptanceprotected
Each accepted risk should include:
- Risk ID
- Description
- Business impact
- Technical impact
- Probability
- Compensating control
- Owner
- Risk approver
- Acceptance date
- Expiration
- Remediation plan
- Target date
- Review cadence
Project managers and engineers should not accept enterprise-level business risk unless authorized.
Exception Managementprotected
Exceptions should be:
- Specific
- Justified
- Approved
- Time-bound
- Monitored
- Assigned
- Reviewed
- Closed
Avoid generic exceptions such as “launch deadline.”
Production Readiness Review Meetingprotected
Participants may include:
- Business owner
- Service owner
- Product owner
- Application team
- Platform team
- Security
- Operations
- Network
- Identity
- Data
- Finance
- Compliance
- Change management
- Vendor representative where required
Review agenda:
- Business objective
- Workload criticality
- Scope
- Architecture
- Test results
- Security
- Data
- Reliability
- Monitoring
- Recovery
- Operations
- Cost
- Open risks
- Exceptions
- Go/no-go decision
- Conditions and owners
Decision record should document:
- Decision
- Date
- Scope
- Participants
- Evidence reviewed
- Blockers
- Accepted risks
- Conditions
- Approvers
- Launch window
- Next review
Production Readiness Metricsprotected
Useful metrics include:
- Workloads reviewed
- Ready on first review
- Average readiness lead time
- Blockers per workload
- Critical security blockers
- Missing ownership findings
- Missing-monitoring findings
- Failed recovery tests
- Unresolved high risks at launch
- Exceptions granted
- Exception age
- Post-launch incidents
- Rollbacks
- Hypercare duration
- Budget variance
- Performance variance
- Support-ticket volume
- Time to operational acceptance
- Post-launch defect rate
- Evidence completeness
Production Readiness Maturity Modelprotected
Level 1 — Ad Hoc
- Launch decisions are informal.
- Evidence is inconsistent.
- Operations learns about services late.
- Risks are poorly documented.
Level 2 — Checklist Driven
- Standard checklist exists.
- Reviews occur near launch.
- Evidence quality varies.
- Exceptions are informal.
Level 3 — Governed
- Risk-based reviews
- Mandatory evidence
- Named ownership
- Formal risk acceptance
- Operational acceptance
- Repeatable launch gates
Level 4 — Integrated
- Readiness built into delivery lifecycle
- Automated evidence
- Continuous security and policy validation
- Standard platform capabilities
- Measured launch outcomes
Level 5 — Optimized
- Risk-based automation
- Predictive readiness
- Production telemetry feeds future reviews
- Exception trends drive platform improvements
- Developer experience and reliability are jointly optimized
Governanceprotected
Define standards for:
- Workload criticality
- Mandatory production controls
- Required evidence
- Review timing
- Review participants
- Go/no-go authority
- Risk acceptance
- Exceptions
- Security approval
- Compliance approval
- Operational acceptance
- Budget approval
- Launch communication
- Hypercare
- Post-launch review
- Evidence retention
- Standard updates
Example Workload Inputprotected
Workload: Public Customer Portal
Criticality: Tier 1 — Business Critical
Architecture:
- Public web application
- Managed application platform
- Managed relational database
- Object storage
- Content delivery network
- Web application firewall
- Federated customer identity
- Central logging
- Multi-zone deployment
- Secondary-region backup
Launch Constraints:
- Public launch date already announced
- Expected launch-day traffic is uncertain
- Penetration testing found two medium findings
- Database restore succeeded in test
- Regional failover has not been exercised
- Operations has not yet signed service acceptance
- Cost estimate excludes increased launch-day logging
Example Executive Recommendationprotected
Recommendation: Conditionally Ready
The workload should not receive final production approval until the following conditions are completed:
- Operations formally accepts the support model.
- Launch-day capacity assumptions are validated through a representative load test.
- Logging-volume cost is included in the approved budget.
- Regional recovery limitations are documented and accepted.
- The two unresolved security findings receive formal disposition.
No confirmed critical security or functional blocker was identified. The announced launch date should not override the unresolved operational-ownership condition.
Example Production Readiness Findingsprotected
PRD-001 — Operations Has Not Accepted Service Ownership
Classification: Launch Blocker · Severity: Critical · Confidence: High
Evidence: No signed operational acceptance exists, and the on-call group has not completed service training.
Business Impact: Production incidents may not receive timely response.
Recommended Action: Complete operational handover, verify access, test alert routing, and record formal acceptance.
PRD-002 — Launch-Day Capacity Is Unvalidated
Classification: Conditional Blocker · Severity: High · Confidence: High
Evidence: Performance tests covered approximately thirty percent of forecast peak traffic, and launch-day demand is uncertain.
Recommended Action: Run a representative load test, validate autoscaling, confirm quotas, and establish launch-day scaling guardrails.
PRD-003 — Regional Failover Has Not Been Tested
Classification: High-Priority Risk · Severity: High · Confidence: High
Evidence: The secondary-region design exists, but DNS, data recovery, identity, and application startup have not been exercised together.
Recommended Action: Complete an integrated recovery exercise or formally accept the current recovery limitation before launch.
PRD-004 — Launch Logging Cost Missing from Forecast
Classification: Conditional Blocker · Severity: Medium · Confidence: High
Evidence: Debug and access-log volumes will be increased during launch, but the approved cost estimate reflects normal retention and ingestion.
Recommended Action: Estimate temporary and steady-state logging cost, confirm budget ownership, and define a date to return to normal logging levels.
Example Go/No-Go Criteriaprotected
Go:
- Operations has accepted the service.
- Critical user journeys pass.
- Load testing validates expected peak demand.
- No critical security findings remain.
- Medium findings have approved dispositions.
- Monitoring and alerts are active and tested.
- Backup and restore evidence is current.
- Rollback remains practical.
- Budget approval includes launch conditions.
- Business owner approves launch.
No-Go:
- No active support ownership
- Failed payment or login journey
- Unresolved critical vulnerability
- Database reconciliation failure
- Capacity below expected demand
- Missing alert routing
- Rollback unavailable
- Required compliance approval absent
Example Launch Runbook Extractprotected
| Step | Time | Owner | Action | Validation | Failure Response |
|---|---|---|---|---|---|
| 1 | 19:00 | Change Manager | Open launch bridge | Required participants present | Delay launch |
| 2 | 19:10 | Release Manager | Confirm approved artifact | Version and signature match | No-go |
| 3 | 19:20 | Database Owner | Apply backward-compatible schema change | Health and validation checks pass | Stop and assess |
| 4 | 19:40 | DevOps Engineer | Deploy production application | Health checks pass | Roll back application |
| 5 | 20:00 | Test Lead | Execute critical user journeys | All critical tests pass | Go/no-go review |
| 6 | 20:20 | Operations | Confirm alerts and dashboards | Signals visible and routed | No-go |
| 7 | 20:30 | Business Owner | Approve public launch | Approval recorded | Delay |
| 8 | 20:40 | Network Engineer | Enable public traffic | Synthetic and real traffic healthy | Revert routing |
| 9 | 21:00 | Incident Lead | Begin hypercare monitoring | Metrics within thresholds | Escalate |
Example Risk Registerprotected
| Risk | Probability | Impact | Mitigation | Owner |
|---|---|---|---|---|
| Launch traffic exceeds forecast | Medium | High | Load testing, quotas, scaling, CDN | Service Owner |
| Customer identity provider degrades | Low | Critical | Monitoring, rate limits, vendor escalation | Identity Owner |
| Logging cost spikes | High | Medium | Volume alerts, temporary retention, budget | FinOps |
| Regional recovery fails | Medium | High | Recovery exercise and accepted limitation | DR Owner |
| Medium security finding becomes exploitable | Low | High | Compensating control and target remediation | Security Owner |
| Support team lacks application knowledge | Medium | High | Training, runbooks, hypercare | Operations |
| Database migration requires forward recovery | Low | Critical | Backward compatibility and tested backup | DBA |
Automation Opportunitiesprotected
- Readiness intake
- Criticality scoring
- Ownership validation
- Evidence collection
- Test-result aggregation
- Security finding aggregation
- Backup validation
- Monitoring checks
- Alert checks
- Cost checks
- Scorecard generation
- Blocker detection
- Exception tracking
- Approval routing
- Launch runbook generation
- Hypercare reporting
- Post-launch review
- A mature automated workflow could import workload metadata, determine required controls by criticality, collect pipeline/testing/security/compliance evidence, validate monitoring, alerts, backup, recovery, ownership, support, cost, and budget, generate a readiness scorecard, highlight blockers, route conditions and risks to owners, require human approval, record the go/no-go decision, monitor hypercare, compare launch outcomes to readiness predictions, and improve standards continuously.
Pro Tipsprotected
- Define readiness requirements at project start.
- Tailor controls to workload criticality.
- Name the service owner early.
- Separate mandatory controls from improvements.
- Require evidence for every material conclusion.
- Treat missing evidence as unknown—not passed.
- Test critical business journeys.
- Test dependency failure.
- Validate provider quotas.
- Include launch-day and seasonal capacity.
- Monitor business outcomes as well as infrastructure.
- Test alert routing before launch.
- Test backup restoration.
- Test integrated recovery.
- Define objective rollback triggers.
- Identify irreversible changes.
- Require operations to accept the service.
- Assign risk acceptance to the correct authority.
- Time-limit exceptions.
- Include cost and budget ownership.
- Define hypercare exit criteria.
- Schedule the post-launch review before launch.
- Use production outcomes to improve the standard.
- Require qualified human approval for every production decision.
Common Mistakesprotected
- Reviewing production readiness too late — readiness gaps discovered immediately before launch are expensive and politically difficult to resolve.
- Treating a checklist as approval — a checklist records evidence; it does not replace technical judgment or risk authority.
- Using an overall score to hide a critical blocker — a single unresolved critical issue may outweigh a high average score.
- Launching without a service owner — temporary project ownership does not create sustainable production accountability.
- Launching without operational acceptance — a service should not enter production when no team is prepared to support it.
- Assuming monitoring exists because metrics are collected — metrics must be connected to dashboards, alerts, ownership, and response procedures.
- Creating alerts without testing them — an alert may fail to route, escalate, or provide enough context.
- Testing infrastructure but not business journeys — healthy infrastructure does not prove that customers can complete critical transactions.
- Ignoring dependency failure — external services, identity, DNS, and databases often determine actual availability.
- Using average load for capacity planning — average demand does not validate peak behavior.
- Ignoring provider quotas — autoscaling cannot exceed provider or account limits.
- Treating backup as recovery — recovery requires restoration, validation, dependencies, and operational procedures.
- Failing to test rollback — rollback procedures often fail because data and schema behavior were not considered.
- Accepting permanent exceptions — every exception should have an expiration and remediation owner.
- Allowing schedule pressure to define risk — an announced date does not reduce security, recovery, or support obligations.
- Excluding cost from readiness — production workloads require accountable budgets, cost allocation, and anomaly detection.
- Ignoring logging cost during launch — temporary increased logging may create significant financial impact.
- Failing to define hypercare exit — project teams may remain indefinitely responsible for production support.
- Failing to perform a post-launch review — production telemetry and incidents should improve future readiness decisions.
- Using AI-generated readiness assessments without evidence — AI may infer controls that do not exist, misunderstand workload criticality, or overlook provider-specific risks. Every material conclusion requires evidence and qualified human validation.
Security Considerationsprotected
- Production-readiness materials may contain sensitive information such as architecture, network ranges, DNS names, security controls, vulnerabilities, identity roles, administrative paths, data classifications, recovery procedures, incident contacts, vendor dependencies, cost details, customer information, launch dates, and known weaknesses.
- Before sharing information with an AI system: remove credentials.
- Remove tokens.
- Remove private keys.
- Remove active connection strings.
- Remove production secrets.
- Sanitize customer and employee data.
- Sanitize vulnerability details where required.
- Sanitize network and identity details where required.
- Follow source-code handling requirements.
- Follow incident and vulnerability disclosure procedures.
- Confirm the AI platform is approved.
- Confirm retention and model-training settings.
- Restrict distribution of generated assessments.
- Do not use AI-generated approval language as a substitute for formal human authorization.
Related Blueprints
⚠ Normalization Warnings — 12 for review
- RESTRUCTURE: The document contains ~50+ domain readiness H1 sections (Business Readiness through Post-Launch Validation). To avoid a flat 50+ section render, all were grouped under one body/group 'Readiness Domains'. Confirm grouping strategy — an alternative would split into thematic sub-groups (e.g., Security & Identity, Data, Reliability & Observability, Operations & Change, Cost & Compliance).
- CLASSIFICATION TO CONFIRM: 'Prerequisites' classified as a checklist TOOL 'Prerequisites Checklist' (gather-before-start items). Alternative: body/prose.
- CLASSIFICATION TO CONFIRM: 'Evidence Register', 'Production Readiness Scorecard', and 'Responsibility Matrix' classified as matrix TOOLS (columnar, filled-in). Scorecard and Responsibility Matrix have doc-provided example_rows verbatim; Evidence Register has no rows in the doc (skeleton). Confirm.
- CLASSIFICATION TO CONFIRM: 'Workload Criticality', 'Readiness Decision Categories', 'Mandatory Versus Advisory Controls', 'Blocker Classification', and 'Production Readiness Maturity Model' classified as body/reference (consulted tiered models). Confirm none should be tools.
- RESTRUCTURE: 'Primary Prompt' and 'Follow-Up Prompts' (two source H1s, 25 prompts total) combined into one prompt_pack tool. Prompt text is verbatim; 'when' guidance lines are editorial additions.
- CLASSIFICATION TO CONFIRM: 'Go/No-Go Criteria' appears twice — once as a reference-style prose body section (kept as body/prose with the doc's example criteria) and again derived into a 'Go/No-Go Criteria' template TOOL. The doc's 'Example Go/No-Go Criteria' remains a body/example. Confirm no duplication concern.
- TOOL DERIVATION: 'Launch Runbook', 'Rollback Plan', 'Risk-Acceptance Record', and 'Go/No-Go Criteria' were derived into template TOOLS from the doc's fill-in field lists (Launch Runbook / Rollback Readiness / Risk Acceptance / Go-No-Go Criteria prose sections). Their body/prose counterparts were kept as reference within Readiness Domains. Confirm this reference-plus-tool split is desired rather than one or the other.
- MATRIX: 'Evidence Register' rubric left empty (rubric:'') because the doc specifies fields but no scoring/usage scheme. skeleton_rows:0 per empty-skeleton rule.
- PLAYBOOK: 'Implementation Roadmap' Phases 1–5 mapped to playbook.roadmap using phase names as horizons (the doc uses named phases, not time horizons). quick_wins intentionally empty — the doc has no Quick Wins section.
- EXAMPLES: 'Example Workload Input', 'Example Executive Recommendation', 'Example Production Readiness Findings', 'Example Go/No-Go Criteria', 'Example Launch Runbook Extract', and 'Example Risk Register' all classified body/example. Tables preserved as HTML tables.
- STATS: prompts=25 (1 primary + 24 follow-ups). deliverables=10 counts the distinct tools. Confirm counting convention.
- OVERLAY: headline, subhead, teaser, exec_brief, and insights are written per voice rules and are not from the doc verbatim — flagged for overlay diff review.
SEO Block
- Title tag: Cloud Production Readiness Review | ABME (40 chars)
- Meta: Run an evidence-based production-readiness review for any cloud workload — launch blockers, go/no-go criteria, risk acceptance, and operational handover included. (162 chars)
- Schema: HowTo · noindex: false
- Related: cl-001, cl-002, cl-003, cl-004, cl-005, cl-006, cl-007, cl-008, cl-009, ai-003, ai-004, ai-007, ai-008, ai-009, ai-010
- Keywords: cloud production readiness, production readiness review, go no-go criteria, launch runbook, operational acceptance, cloud workload launch gate, production readiness scorecard, risk acceptance record, hypercare plan, rollback plan, disaster recovery readiness, cloud launch governance
