Ops notes
Practical notes on SOPs, checklists, AI workflows, and honest ops.
- Server Hardening Checklist for Small Teams (First 10 Servers)Ten controls, one afternoon, no security engineer: key-only SSH, default-deny firewall, automatic security patches, 2FA on the control plane, tested off-box backups — plus the 30-minute first-server card and the honest "what to skip" list.
- Server Monitoring Checklist for Small Teams (5 Signals, 1 Alert Rule)The five signals that catch 90% of small-team outages: outside HTTPS checks, the 80% disk line, cron heartbeats, cert/domain expiry, and a login-path test — plus the one alert rule that doesn't get the phone muted.
- Disaster Recovery Plan Template for Small Teams (The One-Page DR Plan)The one-page DR plan that answers the questions everyone asks during the outage: five fill-in blanks, the RTO/RPO/call-order numbers, failover as exact commands, the degraded mode decision made in advance, and the quarterly drill that makes the numbers real.
- The Uptime Budget: How Much Downtime 99.9% Actually BuysWhat 99.9% really allows (43m 50s/month), how deploys and maintenance spend the same wallet, the one-line ledger, and the burn-pace alert that beats perfection alarms.
- On-Call Handoff Checklist for Small Teams (No Monday Surprises)The written handoff that ends Monday surprises: state of the world, open items with owners, known flaky alerts with the tell, in-flight change landmines, and the seven-line note you can paste tonight.
- Patch Management Checklist for Small Teams (Monthly Cadence That Ships)The monthly patch cadence that actually ships: 4-tier CVE triage, snapshot-before-patch, one production host before the fleet, scheduled reboots, and the one-line patch log auditors ask for.
- Change Management Checklist for Small Teams (No CAB Required)The 10-minute change record that replaces a change advisory board: the 4 questions, rollback-as-commands, freeze windows, and why 12 dated records ARE the audit.
- Deployment Rollback Checklist (Tested Before You Need It)The pre-deploy card that makes every deploy reversible: rollbacks as exact commands, measured rollback time, forward-only migrations, and the tripwires you set before shipping.
- Backup Restore Test Checklist (The 20-Minute Drill)The quarterly restore drill that proves your backups work: pass/fail timing, the checksum rule, the credential-decay defect, and the one-page drill log auditors ask for.
- On-Call Rotation Template for Small Teams (2–6 People)The fair-rotation math for 2 to 6 people, the count-back rule, the post-incident recovery rule, and the Monday handover ritual that keeps the schedule honest.
- AI Agent Ops: 10 Guardrails Before You Let an Agent Run 24/7Ten operational guardrails — budget caps, append-only audit trails, kill switches, self-healing ceilings — that separate a useful 24/7 AI agent from an expensive accident.
- AI Automation Templates for Solo Founders: 5 Workflows Worth Setting Up FirstA practical guide to the five AI automation templates solo founders set up first — what each does, how to configure it safely, and where human review still matters.
- AI Prompt Templates for Content Creators: Build a Reusable Prompt LibraryStop rewriting prompts from scratch — a guide to building a reusable AI prompt template library for creators, with five patterns to start and how to refine them.
- Content Calendar Template for Solo Creators: Plan a Month in One SittingA content calendar template for solo creators — batch a month of posts in one sitting, with formats, slots, and a repurposing loop that halves the workload.
- Customer Support Escalation Checklist: When to Hand Off and HowA customer support escalation checklist for small teams — severity levels, the handoff information that must travel, and the loop that fixes root causes.
- Daily Operations Checklist Template (The 3-Zone Format)A daily ops checklist template in three zones — money, people, systems — so one tired person can still run the day. Copyable structure for 1-50 person teams.
- Employee Onboarding Checklist Template: The First 30 Days, PlannedAn employee onboarding checklist template structured by day one, week one, and the first 30 days — plus the failure modes that make new hires stall.
- The 3 Questions Your IR Plan Must Answer in the First 30 MinutesSmall teams don't need a 40-page incident response plan. They need three answers fast, in order: what is happening, what gets disconnected first, who says what. Log it, cut it, say it.
- Greyscale Mockup Workflow: Present App Designs Before Color Muddies the ConversationA step-by-step greyscale mockup workflow for app and web designers — why structure-first presentation reduces revision cycles, and how to run each stage.
- Incident Communication Templates: Customers, Status Page, and SlackCopy-paste incident comms: the first customer ack, status page updates in plain language, a five-line exec update, and the resolution note — with the two rules that make each one work.
- Incident Response Changelog Automation: Keep the 2am Log Without Relying on 2am MemoryHow to automate an incident response changelog for a small team: log templates, a copy-paste timestamp script, chat-channel logging, snapshot options, and what you should never automate.
- Incident Response Plan Template for Small Teams (1-50 People)A one-page incident response plan template small teams actually use at 2 a.m.: who declares, what to shut down first, who talks. Fill-in-the-blanks, no security staff required.
- Notion vs Spreadsheets for Ops Checklists: An Honest ComparisonNotion and spreadsheets both work as checklist systems — here's an honest comparison of where each wins, where each struggles, and how to decide for your team.
- Postmortem Template for Small Teams (Blameless, Fill-in)A fill-in blameless postmortem that fits on one screen: impact, timeline, contributing factors, and exactly three actions with owners and dates — plus why skipping it costs more than writing it.
- On-Call Handover Template (Copy-Paste, 10 Minutes)A fill-in on-call handover that fits on one screen: status line, alerts open, watchlist, deployment freeze, pending follow-ups, and escalation notes — plus the one rule that makes handovers actually work.
- Access Request & Offboarding Checklist for Small TeamsThe 6-field access request template, the grant recipe, the same-day offboarding checklist, and the 20-minute quarterly review that keeps them true.
- Weekly Status Report Template for Small TeamsThe 5-section one-screen format, what to cut, why async beats the Monday meeting, and the 30-minute Friday routine that writes it.
- Runbook Template (Free, Copy-Paste)The 7 sections a runbook needs, when to write one, and the 2am rule that separates a runbook from documentation theater.
- Process Documentation Best Practices: How to Write Docs People Actually UseProcess documentation best practices for small teams — why docs rot, the write-where-you-work rule, and a maintenance loop that keeps them alive.
- Escalation Policy Template for Small TeamsWho owns the incident, when to wake a human, and the three-level escalation ladder that replaces an on-call rotation you can't staff. One page, named names, hard time thresholds.
- Severity Levels: Why 3 Levels Beat 5 for Small TeamsA five-level severity matrix is five chances to argue instead of act when you have no SOC. Three levels — stop everything, contain, schedule — name the first move in seconds.
- Small Business Automation Ideas: 12 Workflows Worth Automating FirstSmall business automation ideas ranked by payback — 12 workflows small teams automate first, how to spot a good candidate, and when automation backfires.
- SOP Checklist Template: The 10-Part Structure That Actually Gets FollowedWhy prose SOPs go unread and checklist SOPs get used — plus a 10-part SOP checklist template you can copy for any recurring process in your business.
- Standard Operating Procedure Examples for Small Business (5 Formats That Get Followed)Real standard operating procedure examples for small business — five SOP formats, when each fits, and the two reasons most SOPs get ignored within a month.
- The 5 Failure Points a Small-Team Tabletop Will Find (Before a Real Incident Does)A small-team tabletop is just walking through 'every file server shows an extortion note' out loud. These five failure points show up in nearly every team: no decision-maker, backups in the same breach, no pre-written messages, stale access, unread plans.
- Incident Timeline Template: The 5 Timestamps Every Postmortem NeedsTurn a messy Slack scroll-back into a postmortem you can learn from: five timestamps, one-line format, where small-team timelines usually die.
- The 2am Test: Can Someone Else Run Your Incident Bridge Tonight?If your pager went off at 2am, could the person who picks it up run the first thirty minutes without you? Six failure points, a 20-minute fix, and a free pre-written card.
- Weekly Review Checklist: 30 Minutes to Close the Week ProperlyA weekly review checklist that fits in 30 minutes — metrics, loose ends, next week's top three — and how to keep it from becoming a skipped ritual.
- What to Automate First: the 4-Question FilterFour questions, three red flags, one worksheet — pick the recurring task worth automating before you spend a weekend building the wrong thing.
- Workflow Automation Checklist: What to Automate First (and What to Leave Alone)A decision checklist for small-team workflow automation: the 4 questions that pick your first candidate, the 3 red flags that mean don't automate yet, and how to keep a human in the loop.
- Status Page Template: What to Say When You Don't Know What's Broken YetWhat to publish ten minutes into an outage, the one-hour update cadence, and the phrasing that keeps customers calm while you debug.
- AI Agent Failure Postmortem: The 6-Line ChecklistWhen an autonomous agent breaks something: scope the damage, kill the trigger, preserve the log, name the missing guardrail, replay in dry-run.