Demo Technical
Operations

Incident response

Severity levels, first response, and post-incident reviews.

When error rates spike or a region degrades, follow a short playbook. Speed matters more than perfect diagnosis in the first fifteen minutes.

Severity levels

SeverityDefinitionPage
SEV-1Global outage or data loss riskImmediate
SEV-2Single region degraded15 minutes
SEV-3Elevated errors, workaround existsBusiness hours

First response

  1. Acknowledge the alert in the incident channel.
  2. Check status.relay-edge.example for platform incidents.
  3. If customer-side, roll back the last deploy with relay rollback.
  4. Capture request_ids before rotating logs.
Danger
Do not redeploy speculative fixes during SEV-1. Stabilize first, then change code.
  • Incident commander assigned
  • Customer status page updated
  • Post-incident review scheduled
Post-incident review template
After mitigation, write a timeline with detection, impact, and permanent fix. Share it within three business days.