Operations
Incident response
Severity levels, first response, and post-incident reviews.
When error rates spike or a region degrades, follow a short playbook. Speed matters more than perfect diagnosis in the first fifteen minutes.
Severity levels
| Severity | Definition | Page |
|---|---|---|
| SEV-1 | Global outage or data loss risk | Immediate |
| SEV-2 | Single region degraded | 15 minutes |
| SEV-3 | Elevated errors, workaround exists | Business hours |
First response
- Acknowledge the alert in the incident channel.
- Check status.relay-edge.example for platform incidents.
- If customer-side, roll back the last deploy with relay rollback.
- Capture request_ids before rotating logs.
Danger
Do not redeploy speculative fixes during SEV-1. Stabilize first, then change code.
- Incident commander assigned
- Customer status page updated
- Post-incident review scheduled
Post-incident review template
After mitigation, write a timeline with detection, impact, and permanent fix. Share it within three business days.
