How would you answer an interview scenario involving AWS availability and disaster recovery in AWS?
For an interview scenario involving AWS availability and disaster recovery, I would first clarify the business goal, scale, constraints, and the failure or quality attribute the interviewer wants to explore. AWS resilience uses multiple Availability Zones, regional services, backups, replication, Route 53, and tested recovery strategies. In this scenario, given an RTO of 30 minutes and RPO of 5 minutes, explain how you would design recovery for ECS, RDS, S3, messaging, DNS, and secrets across AWS regions. For production, define RTO/RPO, select backup-and-restore, pilot-light, warm-standby, or multi-region patterns as justified, automate backups, test restoration, and document DNS and data failover procedures. I would then explain the main alternatives and tradeoffs, identify likely failure modes, and describe how I would validate the solution through testing, observability, security controls, and recovery or rollback planning.