When an Incident is declared, we follow a specific process in order to ensure that it is resolved quickly. This section describes our incident response process, major roles and what to expect after an incident has been mitigated. We continually refine this process, taking lessons from our own experience and other experts. [1][2][3][4].
- Things to know before an incident
- Incident response process
- After the incident
- Paying out past Incident Report debt
The PagerDuty Incident Response Guide describes a similar process with different roles.
The Google SRE Incident response guide has a wealth of information about incident response and distributed SRE teams.
This ACM blog post describes the complexity of coordinating across a team of distributed responders during an incident, and notes a places where Incident Commander roles may actually hinder responsiveness. It is a good lesson in the complexity of incidents with distributed teams!
The WikiMedia Clinic Duty process also inspired our process here, and is a great overall workflow around distributed SRE.