Incident Severity Matrix Template: P1–P4 on One Page
One page. What P1–P4 actually mean on your team, the sixty-second severity call, and the upgrade/downgrade rule that keeps the numbers honest.
Without a severity matrix, every incident is sized by whoever is most awake and most scared — which means everything is “urgent” at 2am and nothing is urgent on a quiet Tuesday. The pager becomes a mood ring: real P1s get treated like tickets, and typos get escalated like breaches. The fix is not a longer process. It is one table, agreed while everyone is calm, that any human can apply in sixty seconds with no meeting and no guilt.
The severity matrix, one page
Four rows. If an incident needs more than four severities, you are running a bureaucracy, not an on-call. A row per level — what it means, an example your team recognizes, how fast the first human moves, how often everyone hears an update, and who is told:
| Sev | Definition | Example | First response | Updates | Who is told |
|---|---|---|---|---|---|
| P1 | Business stopped: customers cannot pay, log in, or reach their data; data loss or security breach in progress | Checkout down for everyone; production database unreachable; attacker indicator on a production box | Page on-call now; commander named within 5 minutes; work stops on everything else | Status page within 15 minutes; stakeholder update every 30 minutes until stable | On-call, commander, leadership, customers (status page) |
| P2 | Degraded but alive: a workaround exists or impact is partial | Checkout works but takes 20 seconds; one region down; nightly batch failed and can be re-run | On-call engages within 1 business hour | Every 2 hours or on any state change | On-call, system owner, lead of the affected team |
| P3 | Broken but invisible to customers | Internal dashboard erroring; one non-critical job failing; cosmetic bug on a marketing page | Next business day, as a ticket | In the ticket, no broadcasts | The team that owns it |
| P4 | Not broken: polish, docs, nice-to-have | Missing doc, stale copy, refactor nobody is blocked on | Backlog, whenever | Never | Nobody |
Three columns carry the weight. “First response” converts the label into a clock — a severity without a clock is an opinion. “Updates” is the anti-silence column: the number one thing that turns a bad Saturday into a churned account is not the outage, it is the six hours nobody heard anything. “Who is told” ends the “should I wake the founder?” debate — the row answers it, per level, in advance.
The sixty-second severity call
Whoever detects the incident asks three questions, in order, out loud:
- Is money or data moving — or stuck? Payments failing, orders not recording, data being written wrong or read by someone who should not: that is a P1 regardless of how few users are affected, because the blast radius is growing while you deliberate.
- Is there a workaround? No workaround and customers blocked: P1. Workaround exists: P2. Nobody notices: P3 or P4.
- Who has to know in the next hour? The row in the “Who is told” column — page them, then start the update clock.
Call it in sixty seconds, say the level and one reason out loud (“P1 — checkout is down, no workaround”), and let the incident commander take it from there. Uncertainty resolves one way: when torn between two levels, pick the higher one. Over-calling is cheap — one downgrade line in the log. Under-calling is the six silent hours that cost you the customer.
The upgrade/downgrade rule
Downgrades are free. Upgrades need no permission either. What is not allowed is calling it P2 and then not doing P2’s clock. Anyone can downgrade, any time, with one logged line: why, when, and who. Anyone can upgrade, any time, no permission asked — better wrong and early than right and late. The one hard rule: the duties are welded to the label. If you are calling it a P1, the status page and the thirty-minute updates are running; if you are calling it a P2, the two-hour updates are running. A label that skips its own comms clock is a lie, and the post-mortem treats it as the finding.
Comms lanes per severity
The severity label is also a routing table — it decides who hears what, without anyone deciding in the moment:
- P1: status page within 15 minutes (even with no ETA — “we know, we are on it” counts), stakeholder channel update every 30 minutes, direct note to affected customers the same day. Customers hear first, not last.
- P2: stakeholder channel on state changes; affected customers told when the workaround is not obvious.
- P3 / P4: the ticket. Nothing else. No channel-wide pings for a broken internal dashboard.
For what goes in each update, the incident communication templates carries the templates; the matrix only decides when they fire.
Three metrics that keep the matrix honest
- Share of P1s downgraded within the first hour. Healthy is most of them — early over-calls are the system working, not a failure. If nothing is ever downgraded, your P1 is either too broad or people are afraid to touch the label.
- Time from detect to label. Target under sixty seconds, measured per incident. If this creeps up, the matrix is being debated, not applied.
- Labels that skipped their own clock. A P1 with no status page, a P2 with no updates — count them at the post-mortem. The goal is zero; the point of the label is the clock, not the vocabulary.
Worked example: from six silent hours to a ninety-second call
A twelve-person B2B SaaS company. The year before: the payment webhook queue stopped draining at 21:40 on a Saturday. No matrix, so the on-call engineer — alone, unsure, and reading Stripe docs — logged it as “P3, weird webhook errors, will look Sunday.” Reconciliation failed Monday morning, one enterprise customer’s invoices were wrong for a week, and the account churned in the renewal month. Total silent hours: about six, all of them expensive.
They wrote the matrix above in one afternoon and pinned it in the on-call channel. Five months later, the same failure at 21:10 on a Saturday: the on-call asked the three questions — money stuck, no workaround — called P1 at 21:12, commander named by 21:17, status page up at 22:02 with a plain-language “we know, we are on it,” rollback of the previous webhook change executed at 23:05, root cause a partner-side certificate expiry discovered and documented that night. One incident was downgraded P1→P2 at 23:30 with a logged reason. Final metrics from the post-mortem: detect-to-label 90 seconds, first customer update 52 minutes, zero silent hours. Same team, same failure class, one page of difference.
Where this fits
The matrix is the vocabulary the rest of the incident system speaks: the first 30 minutes is what the label starts, the incident commander checklist is the human the label hands authority to, and the escalation policy decides who gets paged at each level. After the dust settles, the post-mortem template is where label misses get reviewed without blame. And the money behind all of it lives on one card: the incident response budget template — five line items, the 40/30/20/10 split, and what to fund before the first page. The declaration itself is a decision row: the RACI matrix template gives it one A (who declares) and one R (who leads) before 2 a.m. ever asks. If you would rather have the whole incident loop — severities, commander lanes, comms, the post-mortem — built and rehearsed for your estate by an outside pair, that is the Small-Team Ops Audit; and the Custom Incident Runbook turns your post-mortems into procedures your team can run without you. The Ops Starter Kit covers the incident half; Vol. 2 covers the on-call rotation that carries the pager.
Severity is what keeps response metrics honest: the incident metrics report computes medians per severity level — the levels this matrix defines — never a blended mean.