HIVE80lab — Ops notes

On-Call Rotation Cadence

Purpose: Define on-call rotation schedules, escalation contacts, and standby processes. Move from "call the person who works last" to structured rotation.

Rotation Structure

TierRoleOn-CallEscalationBackupStandby Period
Tier 1SRE / Infrastructure@sre-lead-01PagerDuty P1@sre-lead-024 weeks
Tier 2Senior Engineering@senior-lead-01PagerDuty P2@senior-lead-024 weeks
Tier 3Engineering Director@cto-leadPagerDuty P3@vp-engineering1 month
Tier 4Operations Manager@ops-managerPagerDuty P4@team-lead2 weeks

Rotation Schedule

SRE / Infrastructure Rotation (4 weeks)

WeekOn-CallBackupHandoff DayHandoff Time
1@sre-lead-01@sre-lead-02Saturday17:00
2@sre-lead-02@sre-lead-01Saturday17:00
3@sre-lead-01@sre-lead-02Saturday17:00
4@sre-lead-02@sre-lead-01Saturday17:00

Senior Engineering Rotation (4 weeks)

WeekOn-CallBackupHandoff DayHandoff Time
1@senior-lead-01@senior-lead-02Saturday17:00
2@senior-lead-02@senior-lead-01Saturday17:00
3@senior-lead-01@senior-lead-02Saturday17:00
4@senior-lead-02@senior-lead-01Saturday17:00

C-Level Rotation (1 month)

MonthOn-CallBackupHandoff Day
September@cto-lead@vp-engineeringLast Saturday
October@vp-engineering@cto-leadLast Saturday

On-Call Responsibilities

Tier 1: SRE / Infrastructure On-Call

Tier 2: Senior Engineering On-Call

Tier 3: C-Level On-Call

Tier 4: Operations Manager On-Call

Standby Process

Backup On-Call

Standby Period (4 weeks)

Handoff Process

1. Primary pulls incident timeline for last incident 2. Primary posts handoff message in channel 3. Backup confirms receipt within 30 min 4. Backup links incident timeline in channel 5. Primary archives incident from active channel 6. Documentation shared: Handoff notes, incident timeline

Escalation Contacts

PagerDuty Integrations

Tier 1 Escalation:

Tier 2 Escalation:

Tier 3 Escalation:

Incident Rotation Workflow

Step 1: New Incident

1. Alert fires → primary on-call notified 2. Primary acks → backup notified of escalation path 3. Primary investigates → updates incident timeline 4. Primary communicates → posts updates to Slack

Step 2: Incident Escalation

1. No response after X min → backup activated 2. Backup takes over → links incident timeline 3. Backup communicates → posts updates to Slack 4. Primary archived → documentation preserved

Step 3: Incident Resolution

1. Primary resolves → marks incident resolved 2. Timeline updated → final status 3. Post-mortem written → linked in timeline 4. Team notified → lessons learned shared

Documentation Fields

Rotations and On-Call Duty

MonthTier 1 On-CallTier 2 On-CallTier 3 On-CallTier 4 On-Call
Sep 2026@sre-lead-01@senior-lead-01@cto-lead@ops-manager
Oct 2026@sre-lead-02@senior-lead-02@vp-engineering@team-lead
Nov 2026@sre-lead-01@senior-lead-01@cto-lead@ops-manager
Dec 2026@sre-lead-02@senior-lead-02@vp-engineering@team-lead

On-Call Duty Compensation

Primary On-Call (Active Duty)

Backup On-Call (Standby Duty)

On-Call Rotation Schedule

Usage: Use for on-call scheduling and escalation. Document primary, backup, and escalation paths. Works with oncall-escalation-slack-structure.md for communication.

Product links: /l/ops-starter-kit-vol-2 | /l/ops-starter-kit | /l/automation-starter-pack

From the HIVE80lab kit

Part of the five-pillar incident-response set: see the pillars overview and the blameless post-incident review template.