Open-Restaurant Incident Response System Checklist
The playbook that turns operational accidents into repeatable fixes: three-day runbooks, four lists, two trees, ten rules, and the third-day cleanup. Honest, zero-fluff.
1. The three-stage runbook (days 1–3)
The restaurant is open. Accidents happen: lost bookings, double-bookings, misprints, rushes. The response must be time-boxed, role-defined, and action-driven.
- Day 1 — Stabilize and measure.
Immediate actions:
- Gather the team. Get the facts on who, when, where, what.
- Enforce the guest-facing fix: moving tables, comping the error, guiding guests to the bar.
- Lock the operational line: freeze checking-in new reservations, prioritize the impacted party only.
- Measure the impact: seats lost, revenue write-offs, time added to all other parties.
Output: a numbered list of factual impacts with owners and timestamps. - Day 2 — Investigate root cause.
Actions:
- Conduct a 10-minute standup per impacted station (host stand, POS, kitchen, bar).
- Map the chain of events in a linear timeline. Mark the exact moment when the error entered the system.
- Interview the staff on the shift, focusing on the moment they noticed the error, not on blame.
Output: a five-line bullet summary of the incident and one line for the root cause. - Day 3 — Diagnose and fix.
Actions:
- Run the decision tree: ask "Was this a process issue, a tool issue, or a people issue?"
- Apply the four-to-do lists:
1. Process-to-do: change the checklist or checklist order.
2. Tool-to-do: add a check or automation.
3. People-to-do: reassign or update the SOP.
4. Data-to-do: track the metric and benchmark.
- Pick the minimum viable fix. Pick one. Not a series of vague "we need to improve everything."
Output: a single action item per incident, with a name, deadline, and metric.
2. The four-to-do lists
Give the response team a predictable set of actions so decisions are rarely left to improvisation.
- Process-to-do. Change how the workflow is documented, ordered, or communicated. Example: swap the step "check availability" to "check availability, then confirm guest preference."
- Tool-to-do. Add checks or automation to the POS, calendar, or back-office system. Example: add an alert if two parties are booked for the same slot.
- People-to-do. Reassign who does the check or what they’re allowed to override. Example: let the floor manager clear underbooked slots without calling the host.
- Data-to-do. Track the metric and set a benchmark. Example: record the total seats lost per incident and run weekly stats on the "critical" incidents.
3. The two decision trees
When the incident happens, the team should pause, not panic. Ask two questions and follow the line.
- Root cause tree.
Ask: Was this a process issue, a tool issue, or a people issue?
- Process: modify the checklist or procedure.
- Tool: add a check or automation.
- People: reassign or update the SOP. - Impact tree.
Ask: Is the impact host-side (booking, seating, reservation hold) or kitchen-side (ticket, line, prep)?
- Host: freeze check-ins, prioritize the impacted party, measure seats lost.
- Kitchen: flag the ticket, adjust prep, communicate to front of house.
4. The ten decision-matrix rules
A checklist that tells teams what to do in every decision point. No guesswork.
- Fix the guest experience first. Comp, move tables, or adjust timing. The operational side can wait.
- Lock the operational line. Stop checking in new bookings for the table affected. Prioritize the impacted party only.
- Measure the impact immediately. Seats lost, revenue write-offs, time added. Use a one-line summary per incident.
- Assign one owner per incident. Not "we all need to do better." One name, one deadline.
- Run the Day 1 standup, not the Day 3 review. The team must know the facts before they can change anything.
- Use the process-to-do list. If you’re going to change something, change the checklist, not the hope.
- Use the tool-to-do list. Add a check, add a safety alert. Automation is the best process lock.
- Use the people-to-do list. Reassign, train, or update the SOP. Don’t keep people in roles that ask them to guess.
- Use the data-to-do list. Track the metric, benchmark the baseline, and revisit the rule.
- Set one action item per incident, not a series. Fix one thing, then measure the difference. Then you’re ready for the next incident.
5. The third-day cleanup
After the fix, the team must document the change and verify the impact. The cleanup prevents relapse.
- Update the SOP. Modify the checklist to include the new action or exclusion. Version it, stamp it, and push it to all staff devices.
- Lock the metric. Add the new metric to the weekly standup: total incidents, critical incidents, avg seats lost, avg time to recover.
- Review and adjust. Read the incident summary. Was the action taken? Did it change the rate? If not, you’re in the same problem; go back to the tree and pick a different root cause.
- Send the signal. Tell the team: "This is now the way we handle this type of error. Here’s the new rule."
Worked example
A twelve-site restaurant group was running an 8% operational accident rate — roughly one incident per site per shift-week, one major incident a quarter, each one costing seats, comped food and a manager's afternoon. They did not add staff. They installed the system on this page: the three-day runbook, the four-to-do lists, the ten matrix rules, and the third-day cleanup, with the incident rate read out in the weekly standup. In one quarter the rate halved — from 8% to just over 4% — and the major incidents went from one a quarter to one in the entire period. The group's own post-mortem of the change made the mechanism plain: most accidents were not caused by people being careless, they were caused by decisions being improvised under pressure; the four lists took the improvisation out, and the rate followed. Total cost: a laminated card per station and one standing agenda item. The managers' summary: "We stopped asking who caused it and started asking which list fixes it."
From the HIVE80lab kit
Every page ships with a kit block — the paid tools behind the free advice:
- The First 30 Minutes — free incident quick-start checklist
- Ops Starter Kit — incident response for small teams — $14
- Ops Starter Kit Vol. 2 — advanced incident response & communications — $27
- Ops Mega Bundle — all 5 kits in one download — $49
Related: the pest control service log is the evidence half of the same promise — the IR checklist proves you handle the crises you can see, the pest log proves you monitor the ones you hope not to; the incident post-mortem template is where Day 3's diagnosis lands when the incident was big enough to warrant a formal review; the staffing shortage coverage plan is the pressure that turns a small error into a visible one — incidents spike when the floor is under-covered; the shift handover log is how Day 1's facts survive the roster change before Day 2's investigation; and the delayed opening notice is what you send when the incident is big enough that the doors cannot open on time at all.
From the HIVE80lab kit
Every page ships with a kit block — the paid tools behind the free advice:
- The First 30 Minutes — free incident quick-start checklist
- Ops Starter Kit — incident response for small teams — $14
- Ops Starter Kit Vol. 2 — advanced incident response & communications — $27
- Ops Mega Bundle — all 5 kits in one download — $49
Related: the incident post-mortem template is where Day 3's diagnosis lands when the incident was big enough to warrant a formal review; the staffing shortage coverage plan is the pressure that turns a small error into a visible one — incidents spike when the floor is under-covered; the shift handover log is how Day 1's facts survive the roster change before Day 2's investigation; and the delayed opening notice is what you send when the incident is big enough that the doors cannot open on time at all.; the refrigeration temperature log is the twice-a-day sheet that keeps the cold chain provable while the doors are open