Internet outage runbook: when the whole office goes dark
Tuesday, 11:05. The router's green lights turn amber, three people announce it in three Slack threads, and someone is already rebooting a switch nobody asked them to touch. An internet outage is the rare incident where the diagnosis is the easy part and everything else — customers, payments, phones, the queue of people asking "is it just me?" — is the hard part. The DNS outage runbook covers the case where the pipe works but names don't resolve; this page is the bigger, dumber failure: the pipe itself is gone, and the team needs to keep working and keep selling until it comes back. The one rule that runs the whole page: declare it once, work the ladder, and never let the customer be the one to tell you the internet is down.
1. First ten minutes: prove it is the ISP and not you
- One wired device, one known-good cable, one speed test. If the wired desktop on the good cable also has no internet, the office wifi is not the suspect. This single test prevents the classic hour lost re-flashing a router that was never the problem.
- Phone hotspot tells you which side of the modem you are on. If the office phone on cellular loads sites fine, your local network may still be suspect — but at minimum the world outside is working, which means this is "our pipe," not "the internet."
- Router lights and modem lights, read in order. Amber or red on the modem's WAN indicator with everything else green is an ISP-side failure signature. Note the exact pattern and the time — it is evidence for the support ticket you are about to open, and the first thirty minutes discipline applies to outages too.
- Check the ISP's own status page from a phone. If their status page is green while three people's traceroutes die at the same hop, screenshot it. A green status page during a real outage is exactly the kind of fact that turns a polite support call into an escalation.
- One reboot, of one device, by one person, then stop. A single power-cycle of the router/modem is a legitimate early step. A cycle of uncoordinated reboots from four people is how a ten-minute ISP blip becomes an afternoon of "which config did we lose?"
2. Declare the outage and staff it
- One channel, one owner, one message. Post to the team channel: what is down, who owns the incident, where updates will land. The incident communication templates work for internal outages just as well as customer-facing ones.
- Name the fallback owner for the moment the primary's laptop battery dies. An outage that kills the office wifi also kills the video bridge the incident owner was running the incident from. The on-call handoff rules apply here: state, then contact, then authority.
- Track the clock from the last verified-good moment. When did the internet last provably work? That timestamp starts your SLA math and ends the "was it down when I left for lunch?" debates.
- Kill the automation that depends on the dead pipe. Scheduled syncs, deploy pipelines, and cloud backup jobs will fail loudly or, worse, retry silently into the void. The cron monitoring rules earn their keep here — a paused backup job is a note in the log; a failed one is a mystery next week.
3. Stand up the failover lane
- The failover lane is a phone, a cable, and a rule. Any recent phone with a working data plan is a backup WAN. The rule: it is for the tasks that keep the business alive, not for YouTube "because nobody can do anything anyway."
- Name what gets the hotspot first. Order of precedence, written before the outage: the payment/POS path, the support inbox, the on-call bridge. Everything else waits. A hotspot shared without priorities is a hotspot shared by whoever shouted first.
- Know the data cap math before you need it. Ten gigabytes is about an hour of video calls or two days of email and dashboards. Tether one laptop as the office's shared gateway if you must, but video calls are the first thing that dies, not the last.
- A 4G/5G backup router is the grown-up version of this section. A backup router with a data SIM, tested quarterly, turns a two-hour outage into a shrug. If the office has none, the runbook's job is to survive today — and the after-action step is to price one.
- People who can work from home should go work from home. The remote work rules apply from day one: VPN up, no "just this once" shadow storage, personal hotspot hotspots stay personal devices. The office being dark is not a permission slip for security debt.
4. Keep the critical path alive offline
- Split the work into "cloud-required" and "works-anyway." Local file edits, design work on local assets, writing, and planning all survive an outage. Anything whose save button is a cloud icon does not — name it, so nobody loses an hour of work to a save that silently failed.
- Switch the cloud-app work to phone-tether, only for the alive lane. Sending invoices, answering the support inbox, running payroll: these ride the failover lane. Everything cloud-flavored but not urgent waits, in writing.
- Write down what the outage silently breaks. Door-badge provisioning, label printers that pull fonts from the cloud, the VoIP desk phones, the smartboard that signs in every morning. The list you capture today is the outage-impact map you reuse next time.
- Payments have a paper fallback that everyone knows. If the point-of-sale is cloud-bound, the manual-receipt procedure exists and someone on shift can run it — or the honest answer is "we cannot take payments right now," and the decision to close follows from that fact, not from hope.
5. Talk to customers before they ask
- Publish the first status note inside fifteen minutes, even if it says nothing new. "We are aware of an outage affecting our office systems and are working with our provider" is a complete first message. The status page template has the cadence and the phrasing; the one-hour update rhythm is what keeps the inbox quiet.
- Use phone and SMS if the web presence is unreachable. A status page hosted on the same dead pipe helps nobody. The personal phones of two named people are the emergency outbound channel until the pipe returns.
- Say what still works, not only what broke. "Order history is unavailable, support email is working, replies may be slower than usual" is a useful message. "We are experiencing technical difficulties" is wallpaper.
- Give the next update a time, not a promise. "Next update at 1:30pm" beats "shortly." An outage with a clock on it stays calm; an outage with only adjectives escalates.
6. Work the ISP ladder
- Open the ticket early, even if the status page says fine. The ticket timestamp is what you point at later for SLA credits, and tier-one scripts resolve "did you try rebooting" in minutes if you have already done it and can say so.
- Ask for the incident number and an ETA in writing. "We have an open ticket, incident #48211, provider ETA under review, last update 12:40" is a status update. "They are working on it" is a mood.
- Escalate with facts, not volume. Businesses affected, hours down, the green status page screenshot, the traceroute dying at the same hop. The vendor escalation ladder is the same climb with an ISP on top of it.
- Log every checkpoint time. Down at 11:05, ticket at 11:20, escalation at 12:15, restored at 14:50, verified 15:05. This log is the input to the after-action section and the evidence for any service-credit claim.
- Know what your contract promises before you need it. The vendor contract review should have surfaced the SLA, the credit formula, and the escalation contact. If it did not, this outage is the prompt to fetch that page of the contract.
7. Decide the shop-stays-open threshold
- Write the threshold down while nothing is down. "If the outage passes two hours and payments are still down, we close the floor and send people home" is a decision made in daylight. The same decision made at hour three is an argument made in the dark.
- Hourly workers get a real answer, not "wait and see." The compensation principle rhymes: people's time is a cost you either spend deliberately or waste resentfully. Two hours of paid "stand by" is cheaper than a day of ambiguity.
- Customers get the closure decision too. If the office closes, the status note says so, with the reopening condition. "Closed until we can take payments; next update 4pm" ends the trickle of "are you open?" calls.
- Sending people home is not failure. A five-person team idle on a dead pipe costs more in morale than the outage itself. The power-failure checklist makes the same point with worse weather: the runbook's job is the decision, not the heroics.
8. After reconnection: verify, log, harden
- Verify before you declare victory. Two pages on two networks, one upload, one video call, one payment test. The vendor outage rules apply: restored is a claim, verified is a fact.
- Un-pause what you paused, in the same breath. Backups, cron jobs, deploys, monitoring holds. A resumed schedule nobody resumed is next month's restore-test mystery.
- Replay the runbook against what actually happened. What did section 3 get wrong about your office? Which cloud app surprised you? Update the runbook the same day — an after-action note written a week later is fiction.
- Price the failover upgrade while the memory is fresh. The backup-SIM router, the second ISP on a different last-mile medium, the tethered-laptop gateway. One quote obtained this week is worth ten remembered intentions.
- Drill it quarterly, fifteen minutes. Cut the WAN (or just pretend to), walk the first three sections on paper, check the hotspot still has data and charge. The drill schedule exists so this rehearsal happens on a calendar, not only in outages.
- Feed the log into the annual review. Downtime hours, customer impact, what the failover lane actually carried. The annual review and the monitoring checklist both get a row out of it.
The whole discipline fits one sentence a five-person team can keep: prove it is the ISP, declare it once, keep the alive lane alive, tell people before they ask, climb the ladder with a ticket number, and decide daylight-bright whether the shop stays open. The internet coming back is the provider's job; the business surviving the wait is yours — and this page is how it does, on paper, before the next amber light.