DevOps · Monitoring
How should metrics and alerting be designed for production?
For production, choose meaningful SLIs, set thresholds with duration and context, reduce duplicate alerts, include runbook links, test alerts, and review noisy or unactionable rules regularly. Add automated tests and observability around the critical behavior, document ownership and failure handling, and review the design when traffic, dependencies, or security requirements change.