HIVE80lab — Ops notes

Playbook Selection Rules

Purpose: Define clear rules for selecting the correct playbook. Avoid decision paralysis and ensure consistent response.

Fundamental Rules

Rule 1: Start with First 30 Minutes

Every incident must start with first-30-minutes-incident-response playbook.

Rule 2: Severity Defines Tier, Not Playbook

P1 = Tier 1, P2 = Tier 2, P3 = Tier 3, P4 = Tier 4

Rule 3: Category Trumps Secondary Factors

Primary category overrides secondary factors.

Rule 4: Multiple Playbooks May Apply

Complex incidents may require multiple playbooks.

Rule 5: Always Document Reasoning

Each playbook selection must be documented with reasoning.

Category-Based Selection Rules

Availability Incidents

Primary Rule: Always use Availability playbook first.

SeverityPlaybookFocus
P1Tier 1 Availability PlaybookInvestigation and communication
P2Tier 2 Availability PlaybookResolution and customer comms
P3Tier 3 Availability PlaybookExecutive involvement, major comms
P4Tier 4 Availability PlaybookRCA, lessons learned

Sub-playbooks (when needed):

Security Incidents

Primary Rule: Always use Security playbook first.

SeverityPlaybookFocus
P1Tier 1 Security PlaybookContainment, preserve evidence, escalate
P2Tier 2 Security PlaybookInvestigation, customer notification
P3Tier 3 Security PlaybookLegal, compliance, press comms
P4Tier 4 Security PlaybookRCA, training, prevention

Sub-playbooks (when needed):

Performance Incidents

Primary Rule: Always use Performance playbook first.

SeverityPlaybookFocus
P1Tier 1 Performance PlaybookInvestigation, root cause
P2Tier 2 Performance PlaybookResolution, customer comms
P3Tier 3 Performance PlaybookEngineering changes, optimization
P4Tier 4 Performance PlaybookRCA, monitoring improvements

Sub-playbooks (when needed):

Functional Incidents

Primary Rule: Always use Functional playbook first.

SeverityPlaybookFocus
P1Tier 1 Functional PlaybookRoot cause, fix development
P2Tier 2 Functional PlaybookRollback, customer comms
P3Tier 3 Functional PlaybookEngineering changes, stakeholder comms
P4Tier 4 Functional PlaybookRCA, feature improvements

Sub-playbooks (when needed):

Data Incidents

Primary Rule: Always use Data playbook first.

SeverityPlaybookFocus
P1Tier 1 Data PlaybookContainment, preserve evidence
P2Tier 2 Data PlaybookRestore from backup, notify customers
P3Tier 3 Data PlaybookComplex restoration, compliance comms
P4Tier 4 Data PlaybookRCA, backup improvements

Sub-playbooks (when needed):

Third-Party Incidents

Primary Rule: Always use Third-Party playbook first.

SeverityPlaybookFocus
P1Tier 1 Third-Party PlaybookVendor escalation, workaround
P2Tier 2 Third-Party PlaybookCustomer comms, SLA tracking
P3Tier 3 Third-Party PlaybookVendor penalties, legal escalation
P4Tier 4 Third-Party PlaybookRCA, alternative providers

Sub-playbooks (when needed):

Tier-Specific Selection Rules

Tier 1 (15 min)

Tier 1 playbooks focus on investigation and communication.

Rules: 1. Never escalate without reason: Document why escalation is needed 2. Communicate early: Update Slack, incident timeline 3. Link to timeline: Track every step 4. Acknowledge immediately: Use emoji reactions, confirm receipt

Tier 1 Playbooks:

Tier 2 (30 min)

Tier 2 playbooks focus on resolution and customer comms.

Rules: 1. Allocate resources: Devs? QA? Design? Customer success? 2. Update timeline: Document every decision 3. Customer comms: If business impact > 30 min, notify customers 4. Escalate if needed: If unresolved in 30 min, escalate to Tier 3

Tier 2 Playbooks:

Tier 3 (1 hour)

Tier 3 playbooks focus on major incidents, executive involvement, stakeholder comms.

Rules: 1. Finalize plan: What will we do? What resources? What timeline? 2. Customer comms: Notify customers of impact, solution, compensation 3. Executive involvement: CTO / Director notified 4. Post-mortem started: Document incident, root cause, prevention

Tier 3 Playbooks:

Tier 4 (4 hours)

Tier 4 playbooks focus on RCA, lessons learned, prevention.

Rules: 1. Complete RCA: Document root cause, contributing factors 2. Prevention plan: What changes will prevent recurrence? 3. Track recurrence: Similar incidents in last 30 days? 4. Team learning: Share findings in weekly ops sync

Tier 4 Playbooks:

Special Cases

Mixed Impact Incidents

Example: Payment gateway down (Availability + Business + Data)

Rule: Use Availability playbook for Tier 1, Data playbook for Tier 2, Third-Party playbook for customer comms.

Selection:

Recurring Incidents

Rule: Use simulation-after-action-report playbook to prevent recurrence.

Steps: 1. Review post-mortem from previous incident 2. Identify root cause (use incident-taxonomy-root-cause-categorization) 3. Create prevention plan 4. Schedule mock drill (use incident-response-drill-schedule-template) 5. Track recurrence

Unexpected Third-Party Impact

Rule: Use customer-vendor-incident-comms playbook immediately.

Steps: 1. Notify customer (provide transparency) 2. Escalate to vendor (use vendor-incident-coordination) 3. Track SLA (record vendor response time) 4. Compensate if SLA violated

Documentation Fields

Usage: Use for every incident selection. Follow rules for consistency. Document decisions in incident timeline.

Product links: /l/ops-starter-kit-vol-2 | /l/ops-starter-kit | /l/automation-starter-pack

From the HIVE80lab kit

Part of the five-pillar incident-response set: see the pillars overview and the blameless post-incident review template.