Promise only what you can measure: the one-page SLA/SLO
Every small team gets asked "what's your SLA?" in the first serious sales call, and too many answer with a number someone once said out loud — 99.99%, response in one hour, 24/7 support — that nobody wrote down and nobody can honor. The fix is not a longer contract. It is a definition: three sentences that say what you measure, what you target, and what you promise, written once and kept honest by a quarterly review. This page is that definition. The internal half (SLI and SLO) is for the team; the external half (the SLA) is for the customer; and the seam between them is where small teams either build trust or quietly start lying. The uptime budget page covers how much downtime a number like 99.9% actually buys; this page covers how to pick the number in the first place.
1. SLI, SLO, SLA — three words, one sentence each
- SLI — the measurement. The thing you actually count: successful responses divided by total responses, the share of requests answered under 500 milliseconds, the minutes from page to first human acknowledgment. An SLI is a number your monitoring can produce without a human assembling it — if a person has to build the number by hand every month, it is not an SLI, it is a story.
- SLO — the target. The line the team holds itself to on that measurement: 99.5% availability over a rolling 30 days, p95 under 800 milliseconds, first human response inside 15 minutes for S1. The SLO is internal. Nobody outside the team needs to see it; the team needs to live inside it. The monitoring checklist is what turns an SLO from words into a dashboard.
- SLA — the promise. The external document: what the customer can hold you to, with named consequences when you miss. The test that keeps the three honest: every SLA line must trace to an SLO, every SLO to an SLI you can produce on demand. A promise with no measurement behind it is not a commitment — it is a future apology with a signature.
2. Pick SLIs you can actually measure
- Availability, from real traffic. Successful responses over total responses, measured by the thing that serves the traffic — not by how the founder feels about the week. For a small team this is usually one number from a health check or the load balancer, and that is fine. Crude and true beats elaborate and imagined.
- Latency, p95 not average. The average hides the tail, and the tail is what customers experience. If your average is 300ms and your p95 is 4 seconds, the customer on the tail does not care that "on average" you are fast. Promise on the percentile you can measure, and let the average be good news you do not sell.
- Time to first human response. The SLI small teams forget, and the one customers remember: minutes from ticket or page to a human who says "I am on it." It is the cheapest SLI to improve and the one that makes every other number forgivable — the rotation schedule is what makes it possible, and the onboarding checklist is what makes it survivable.
- The production test. Before you commit to any SLI, ask: who produces this number, how often, and where does it live? If the answer is "whoever is awake, monthly, in a spreadsheet," either automate it or pick a cruder number you can automate. The monitoring checklist covers the minimum stack that makes the big three (availability, latency, response time) self-reporting.
3. Set SLOs with headroom — the 2× rule
- Promise less than you deliver. Look at what you actually delivered last quarter and promise roughly half the failure rate. Measured 99.7%? Promise 99.5%, not 99.9%. The headroom is not sandbagging — it is the difference between a promise you keep in a bad month and a promise that turns every bad month into a breach. Headroom is where rollbacks, deploys, and Tuesday mistakes live without breaking faith.
- Write the window and the ledger. Every SLO needs two details or it is a vibe: the measurement window (rolling 30 days is the honest default — a calendar-month SLO lets a bad week hide in a good one) and where the number is published (a dashboard, a line in the weekly status report, anywhere the team sees it without opening a ticket). The downtime budget ledger is the one-line version: minutes spent vs minutes allowed, updated per incident.
- One SLO per promise, written as a sentence. "The API answers successfully 99.5% of the time, measured over each rolling 30-day window, on the status dashboard." If the SLO needs a paragraph, it is three SLOs — split it. If two SLOs can conflict (99.9% availability and deploys every day and no maintenance window), the conflict is the finding, and the freeze policy or a maintenance window is the fix.
4. The external SLA — one page, plain words
- Name the covered thing precisely. "The Dashboard web application and its API, hosted by us at app.example.com." Not "the service," not "the platform" — the customer should be able to point at a URL and ask "is this covered?" Ambiguity here is where every future dispute lives.
- State the hours of coverage. 24/7 paging for S1 with a named on-call human, best-effort business hours for everything else, is a legitimate small-team SLA — and a sustainable one, because it matches the compensation you actually pay. Do not promise 24/7 for S3 tickets; nobody sane wants a 3AM phone call about a billing export, and nobody should be paid to take one.
- Write the exclusions in the same voice as the promises. Scheduled maintenance announced 48 hours ahead; outages caused by third parties you do not control (name the categories: upstream clouds, payment processors, DNS registrars — the vendor outage runbook is the playbook for those); problems caused by the customer's own configuration, keys, or integrations; and anything outside your control entirely. Exclusions are not fine print — write them as plainly as the promises, because the customer who reads only the exclusions is the one who stays a customer.
- The plain-words test. Hand the draft SLA to someone non-technical and ask them to say back what they are entitled to and when. If they cannot, the SLA is written to win arguments, not to keep customers — rewrite it. A one-page SLA a customer can re-read beats a ten-page one a lawyer has to interpret; the status page template is the matching voice for the days you are living inside it.
5. Response times tied to severity — not vibes, the matrix
- Borrow the severity matrix you already run incidents with. S1 (production down, money path broken): acknowledge in 15 minutes, work continuously, status update every hour until restored. S2 (degraded, workaround exists): acknowledge in 2 business hours, daily updates. S3 (cosmetic, no workaround needed): next business day. The severity matrix already defines these levels for your own incidents — using the same ladder in the SLA means the customer's ticket and your pager speak the same language, and the timeline becomes your compliance record for free.
- Promise acknowledgment, not instant repair. The SLA clock starts at the ticket and stops at first human acknowledgment with a named owner. Resolution time is a different promise — one that depends on the bug, not on you — and small teams should not sign it. What you can always honor: a human answered, said what they are doing, and said when they will speak again. That is what "responsive" means to a customer at 2AM.
- Define the clock once, in one line. "The response clock starts when the ticket or page arrives and stops at the first substantive human reply; updates continue at the stated cadence until resolution." One line, no footnotes. The handoff checklist keeps the clock honest across rotation changes — a promise that dies at shift change is a promise the next shift inherits broken.
6. Consequences — what happens when you miss
- Service credits are theater at small scale. A 10% credit on a $19/month plan is $1.90 — nobody's churn decision turns on it, and refund math will not save the relationship. What actually saves it: the outage was announced before the customer noticed, updates were honest and rhythmic, and a written postmortem arrived after. The postmortem template you run internally is the artifact to share (redacted) — it converts a bad day into evidence that you are the kind of team that fixes things.
- When a customer's procurement genuinely needs a credit clause: cap it, window it, and tie it to the ledger. Credits capped at the last month's fees, claimable within 30 days, payable only for S1 breaches measured on the same dashboard the team watches. A credit clause you can operate in an afternoon beats an impressive one you would fight in a quarter.
- The consequence you should actually fear is churn, not the clause. Price the SLA into the plan: the cost of the on-call stipends, the monitoring that produces the SLIs, and the headroom you promised. A plan that carries an SLA it cannot afford is a subscription to resentment — and the quarterly review (next section) is where you notice.
7. The quarterly review that retires promises
- Pull the real numbers and put them next to the promises. Once a quarter, one hour: measured availability vs promised, p95 vs promised, first-response times vs promised, each from the ledger and dashboards — not from memory. Missed once with a story (the postmortem has it): note it. Missed twice: the SLO is wrong — either fix the system or lower the number. A promise you miss repeatedly is not a stretch goal; it is a lie on a schedule, and customers keep receipts.
- Let incidents and vendors edit the document. Every S1 postmortem action and every third-party outage that ate your budget is a candidate edit: a sharper exclusion, a moved threshold, a dependency the SLA should name. Feed the year's findings into the annual review as named open items, so the SLA gets one serious rewrite a year instead of a hundred guilty ones.
- Retire promises loudly, not quietly. When a number must come down — product changed, a dependency got fragile, a market got served worse than promised — tell affected customers before the next billing cycle, in the same plain voice the SLA was written in. Renegotiating before the breach is integrity; renegotiating after the invoice is churn with extra steps.
Small-team honesty note: if the team is one founder and a part-time admin, your SLA is one page with five lines and a phone that you actually answer — and that beats most enterprise SLAs, because the enterprise version promises a human and delivers a queue. Promise less than you deliver, measure what you promise, and never sign a number you have not seen yourself produce twice. And if the founder is also the on-call rotation, the first SLO to write down is the one protecting the founder's sleep — the compensation policy and the rotation schedule are that promise to yourself.
Related: uptime and downtime budget · severity matrix · server monitoring checklist · status page template · on-call rotation schedule · on-call compensation policy · postmortem template · vendor outage runbook · incident timeline · change freeze policy · weekly status report · on-call onboarding · vendor escalation ladder