First 30 Minutes Incident Response Checklist
What ops teams do when an incident hits: stabilize, measure, gather facts. Zero-fluff, time-boxed, repeatable.
Phase 0: The first 30 minutes (time boxed)
The first 30 minutes are about stabilization and measurement, not root cause. Do not try to solve everything at once. Focus on keeping the situation under control and collecting enough data to make the next hours productive.
0–5 minutes: Gather the facts
Who, what, when, where, how serious. No opinions yet.
- Identify the incident. What happened? Check status pages, alerts, and logs.
- Confirm scope. Which systems, users, or customers are affected?
- Identify owners. Who is responsible for the affected systems? Which person oncall or squad needs to know?
- Collect the baseline. What was normal before? Capture current performance (latency, error rate, uptime).
- Document. Create a short note: date, time, observed behavior, affected scope, owners.
5–15 minutes: Stabilize
Limit the blast radius, protect the rest of the system, and reduce user impact.
- Isolate or route. Disable the failing component, route traffic away, or add a failsafe.
- Minimize exposure. Temporarily freeze new signups, pause non-critical features, or show a maintenance page.
- Communicate internally. Notify the team and oncall. If applicable, trigger the incident runbook channel.
- Preserve evidence. Take screenshots, capture logs, or start a stack trace before you make changes.
- Measure impact. Collect concrete metrics: total users affected, lost revenue (if any), duration so far.
15–30 minutes: Measure and prepare
Quantify the situation and align on next steps.
- Complete the impact summary. How many users affected? Which features down? Is the problem growing?
- Set the recovery plan. Identify two to three potential solutions. Don't pick one yet—just list options.
- Prepare the communication. Draft an internal note and, if needed, a customer communication.
- Escalate. If you need security, legal, or PR involvement, trigger the appropriate process now.
- Document the decision. Record the chosen path and who is responsible for executing it.
Next steps after the first 30 minutes
By the end of the 30-minute window, you should have:
- A factual summary of what happened, when, and who it affected
- A stabilization plan that contains the system is safe and impact is under control
- Prepared options for the next stages (investigation, escalation, or workaround)
- A named owner for each next step and a clear time box for the next phase
The full Ops Starter Kit Vol 2 includes a 3-day incident response runbook, decision trees, playbooks, and cultural guidance to turn this initial stabilization into repeatable fixes. Get the complete kit at https://docs.ops-notes.com/ops-starter-kit-vol-2.
Put this to work. The Ops Starter Kit bundles the highest-leverage templates — incident response, runbooks, onboarding, checklists — into one download ($14).
Free start: grab the First 30 Minutes incident-response checklist, or get the free template library by email.