Here is the test, and it takes ten seconds to administer: if your pager went off at 2am, could the person who picks it up run the first thirty minutes without messaging you?
Not "could they eventually figure it out." Could they run it — declare, comms, triage, escalate — while you are unreachable, with no context except what is written down?
If your honest answer is "they'd call me," you don't have an incident process. You have a person. And that person, statistically, is you, asleep, on a plane, or in a meeting with your phone on silent.
We wrote this after watching the same failure repeat across small teams: the incident process lives in one head, works brilliantly, and then evaporates the first night it's actually needed.
What actually breaks at 2am
When the runbook author is asleep, six things break in a predictable order:
- Nobody declares. The first responder spends 20 minutes investigating privately, hoping it's nothing. Meanwhile the clock on your recovery time objective is running.
- Nobody roles up. "Incident Commander" feels like a title for bigger companies. So three people all kind of lead, and nobody talks to stakeholders.
- The status update never goes out. The first #incidents message is also the last one for 45 minutes. Customers find out from Twitter instead of your status page.
- Severity is guessed. Is a slow checkout P1 or P2? Without written thresholds, the night shift picks "medium" because that feels safer.
- Escalation stops at the first no-answer. Page the on-call, no reply after 5 minutes… and then? An unwritten escalation path is a dead escalation path.
- The timeline is lost. Nobody records what happened when, so the post-mortem becomes archaeology, and the fix becomes a guess.
None of these are skill problems. Every one of them is a writing problem: the decisions were made once — by you, calmly, in daylight — and never written down where the 2am version of your team can find them.
The 20-minute fix
You do not need enterprise process. You need one page and one rule.
The page is a first-30-minutes card with five answers on it:
- Who declares, and where (the exact command, channel, or button — not "raise an incident").
- The severity table with thresholds, e.g. "P1 = customers cannot check out; P2 = checkout works but is slow." Binary, not vibes.
- Who owns comms in minute 5, minute 30 — and the template sentence they post.
- The escalation ladder with real names and real wait times — "no reply in 5 minutes → next person on the list."
- Where the timeline lives (one channel, one thread, append-only).
The rule is: the card must work for someone who has never seen your system. If it references "the usual suspects" or "the Jenkins thing," it fails. Write it for the newest hire, because at 2am your senior engineer is also effectively a stranger to the situation — tired, rushed, and missing context.
Then drill it once
A card that has never been used is a document, not a process. Run a 30-minute tabletop: pick a Tuesday afternoon, break something harmless (or simulate it), and have someone who didn't write the card run it end to end.
You will find the gaps in minutes: the escalation list has a name of someone who left, the "status page" step assumes access nobody on night shift has, the severity thresholds have a tie (both "slow" and "down" match). Every one of those is a two-line edit — and every one is a failure you just avoided at 2am.
That's the entire trick. Small teams don't lose incidents to complexity; they lose them to unwritten decisions. Write them down, drill them once a quarter, and the newest hire runs a calmer first thirty minutes than most enterprise teams.
The 30-minute card, pre-written
If you'd rather not start from a blank page, we packaged the checklist we use — every prompt above, plus the timeline thread format and the comms templates — as a free download:
- The First 30 Minutes (free checklist): [hive80lab.gumroad.com/l/first-30-minutes](https://hive80lab.gumroad.com/l/first-30-minutes)
If you want the full kit behind it — severity matrix, on-call ladder builder, post-mortem template, and the tabletop script to drill it:
- Ops Starter Kit ($14): [hive80lab.gumroad.com/l/ops-starter-kit](https://hive80lab.gumroad.com/l/ops-starter-kit)
And if your next gap is what happens after hour one — stakeholder comms, exec updates, customer messaging — the advanced volume covers exactly that:
- Ops Starter Kit Vol. 2 — Advanced Incident Response & Comms (≈US$19.50): [hive80lab.gumroad.com/l/ops-starter-kit-vol-2](https://hive80lab.gumroad.com/l/ops-starter-kit-vol-2)
Start with the free card. Fill in the five answers tonight — it really is a 20-minute job — and the next 2am page gets answered by a process instead of a person.
Written by the Hive80 Lab crew. We build small, opinionated incident-response tools for teams of 2–20.