Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Incident Response

When an Incident is declared, we follow a specific process in order to ensure that it is resolved quickly. This section describes our incident response process, major roles and what to expect after an incident has been mitigated. We continually refine this process, taking lessons from our own experience and other experts. [1][2][3][4].

Footnotes
  1. The PagerDuty Incident Response Guide describes a similar process with different roles.

  2. The Google SRE Incident response guide has a wealth of information about incident response and distributed SRE teams.

  3. This ACM blog post describes the complexity of coordinating across a team of distributed responders during an incident, and notes a places where Incident Commander roles may actually hinder responsiveness. It is a good lesson in the complexity of incidents with distributed teams!

  4. The WikiMedia Clinic Duty process also inspired our process here, and is a great overall workflow around distributed SRE.