A small incident record that helps the next responder
Capture decisions and evidence without filling a giant form.
Record the first known symptom
Write the affected task, first observed time and how the issue was detected. Separate the time an incident started from the time someone noticed it. If the start time is uncertain, say so rather than inventing precision.
Build a factual timeline
For each significant event, record time, observation, action and result. Identify which deployment or configuration was involved without copying secrets. A hypothesis belongs in the record as a hypothesis until evidence supports it.
Verify recovery through the user task
A process restarting is not the same as a customer completing the affected task. Check the actual flow and record the evidence used to declare recovery. Monitor for recurrence over an appropriate period and identify who owns follow-up.
Turn the record into one practical improvement
Choose an action that addresses the demonstrated failure: a clearer alert, a safer rollback, a missing test or a corrected runbook. Include an owner and completion condition. Avoid a long list of vague promises that cannot be checked later.
Keep working through the question
Unkillable Window · Published September 17, 2026. Prepared with AI assistance; editorial approach and corrections are described on our method page. Examples are illustrative, not case-study results.