Major Cloud Infrastructure Outage & Disaster Recovery Runbook
Manage multi-region failover, customer status communication, and incident post-mortems during a multi-hour cloud outage.
The Prompt
Copy and customize
Role: You are a VP of Infrastructure and Site Reliability Engineering directing outage recovery.
Context: A core AWS/GCP availability zone outage takes down your SaaS platform, causing customer login failures and triggering thousands of panic support tickets.
Task: Execute a Disaster Recovery Runbook and Multi-Region Failover Sequence.
Input Available:
- Primary vs Secondary cloud region architecture
- Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets
Output Format:
- Severity 1 Incident Triage Checklist (Incident Commander, Tech Lead, Comms Lead, Ops Lead)
- DNS Failover & Database Promotion Sequence (Switching read-replicas to master in secondary region)
- Live Status Page Updating Cadence (Updates every 20 minutes on status.company.com)
- Customer Support Escalation Macro Library (Clear, technical, honest answers)
- Service Level Agreement (SLA) Credit Calculation Model for impacted enterprise contracts
- Blameless Post-Mortem Template (Timeline, Root cause [5 Whys], Preventive action items)
- Notion SRE Incident Archive & Post-Mortem Tracker
Guardrails & Quality Control:
- Update status pages regularly even if there is no new technical progress to maintain user trust
- Conduct post-mortems with blameless language focused on systemic design improvements
This prompt is exclusive to Pro members
To view and copy this prompt, purchase any official TailorFlow Notion product (such as Business OS) to receive an exclusive unlock code.
Notion Ultimate Bundle
Get instant lifetime access to all current and future Notion templates. Transform your work and personal life with the ultimate Notion Bundle.

How to Use
Run this prompt in four steps
- 1Trigger the SRE on-call pager and spin up the Incident Command Slack channel.
- 2Follow the step-by-step database failover runbook.
- 3Publish the comprehensive blameless post-mortem in Notion within 48 hours of resolution.
When to Use
When to use this prompt
Use during service outages exceeding 15 minutes in duration.
Limitations · Worth Knowing
This prompt has limitations you must understand.
Failover mechanisms must be tested quarterly; untested recovery scripts frequently fail in real outages.