Server room environment monitoring: temperature, humidity, water, and the closet nobody watches
Here is the story that repeats at every small office: the NAS and the server move into a telecom closet one summer because it's out of the way. The closet has no vents, shares a wall with the kitchen, and holds the switch, the UPS, two servers, and a mop someone stored there in 2022. It runs fine for fourteen months. Then a heat wave arrives on a Friday, the room hits 41°C by Saturday afternoon, and on Sunday someone notices the file share is slow — right before a drive that ran hot for thirty-six straight hours dies on Monday morning. Nothing was "down." Nothing alerted. The machine cooked quietly while everyone watched dashboards that never asked the room how it felt. This page is about adding the room to your monitoring: temperature and humidity thresholds that mean something, a $30 water sensor under the rack, heat and smoke you'll actually hear, a door alert for the closet everyone opens, and the one-glance monthly log that catches the slow drift before it becomes the Monday incident.
1. The audit: measure the closet before you trust it
- You cannot set thresholds for a room you've never measured. Put a cheap thermometer in the closet for a week — USB sensor, Wi-Fi sensor, whatever reads out. Record the afternoon peak on the hottest day, the overnight low, and what happens at 12:30 when the kitchen wall heats up. Small-team closets routinely swing 12°C between a Sunday morning and a Tuesday afternoon, and nobody knows until a drive complains. The week of readings becomes the baseline every threshold below is set against — the same evidence-first instinct as the power failure checklist's runtime re-measure: a number written down beats an assumption that feels right.
- Airflow is the cheap fix hiding behind the expensive one. Before anyone suggests an air conditioner, check what the closet actually does with its air: a door undercut, a vent, a fan pulling from the hall. Half the "closet is hot" problems are a door wedge that fell out in 2023 and a rack door closed since forever. Mark the airflow path on the closet card (in through the front, out through the undercut) the same way the cold-start card marks the power-on order — if it isn't written down, the next person who stores a box of cables in front of the intake doesn't know they've broken it.
- Inventory the heat sources with the same register you keep for everything else. The switch, the NAS, the UPS (a UPS idles warm; a UPS charging after an outage is the hottest box in the room), the old mini-PC someone repurposed and never documented. Every heat source belongs on the asset inventory with a power note, because the closet's thermal budget is a real constraint on what you're allowed to put in it — and "what's in the closet" is a question nobody can answer from memory once three people have shelved things there.
- Decide the maintenance class of the room. A closet with dust, a mop, and a recycling bin is a fire-risk room pretending to be a server room. Half a day: vacuum, remove combustibles, wipe intakes, check that nothing leans against the UPS, confirm the shelf brackets are screwed (not wall-anchored-in-plaster). This is the physical security walk extended by one door — same routine, different room, and it pairs naturally with the quarterly UPS unplug drill so one calendar entry covers both.
2. Temperature and humidity: thresholds that mean something
- Set one actionable number, not five aspirational ones. For a small-team closet: warn at 30°C sustained (30+ minutes), alert at 35°C, and act like an incident at 40°C (shut down the NAS yourself — thermal shutdown on its own terms is a dirty stop). Humidity: warn below 30% (static risk in winter) and above 60% (condensation and corrosion). One warning, one alert, one stop line — the same three-level severity shape as the severity matrix, because a room with eleven alerting thresholds is a room whose alerts everyone mutes.
- "Sustained" is the word that saves you from alert fatigue. A closet that hits 31°C for six minutes when the server runs a backup is normal; 31°C for three hours is a problem brewing. Every threshold needs a duration clause, and the sensor alert needs the same "did the thing that should have happened, happen" wiring as the cron monitoring checklist — it's not enough for the sensor to be capable of alerting; the alert has to actually arrive, and prove it arrives, the way a shutdown script proves it runs by being fired once on purpose.
- The trend matters more than the snapshot. A closet that peaked at 27°C last July and peaks at 33°C this July has a story: a fan died, a vent got blocked, a new box moved in, or the battery in the UPS is venting. The monthly log (section 5) is what turns a single reading into a trend line — one row per month, hottest reading and when. Without the history you'll keep re-deciding "is 33°C bad?" every summer from scratch, which is exactly the kind of question the server monitoring checklist exists to pre-answer.
- Humidity is the sensor everyone skips, in the season it matters. Winter offices run 20–25% humidity with baseboard heating; that's static-discharge territory for hardware and a shock waiting for the person who racks the next device. Summer offices with a leaky pipe run the other way. A combined temp/humidity sensor costs the same as temp alone. Set the band, log the extremes, and note the season when you set it — a threshold chosen in January looks wrong in August, and that's normal.
3. Water: the $30 sensor under the rack
- The leak is never in the server room; the water comes to it. Above a ceiling closet: the bathroom upstairs, the coffee machine's supply line on the shared wall, a roof drain. Beside it: the kitchen sink's trap. Under it: the floor drain nobody remembers existed. A spot-leak sensor on the floor under the lowest point of the rack, plus one on the wall facing the wet room if they share it, covers the arrival paths. Wi-Fi leak sensors with an app and a loud local beep are fine; the requirement is that it screams in the room and notifies a phone.
- Put a second sensor at the door level if the closet is in a basement. Basement closets flood from the bottom; a floor sensor under the rack misses the four centimetres of water that arrive around the door threshold first. This is the same "alarm where the threat actually enters" thinking as the badge access logic — sensors at the perimeter, not just at the valuables.
- Decide the response before the water, in one line on the wall card. "Water sensor sounds: kill power at the breaker to the closet, unplug the NAS after clean shutdown if reachable, do NOT step into standing water near powered equipment." One line. The cold-start card taught the format; the leak card teaches the discipline: the moment you're standing in water at 2 a.m. is not the moment to invent an electrical safety policy.
- Test it with a wet finger, log the test. Sensors fail silent the way batteries fail silent — green light, dead sensor. A five-second wet-finger test each quarter (same window as the UPS unplug drill) with a one-line log entry is the whole maintenance plan. The trap this section exists to prevent is identical to the power page's: three years of a green light on a sensor whose battery died in year two.
4. Heat, smoke, and the door: who notices when nobody's in the room
- Heat is the failure mode the temp sensor misses. A temp sensor tracks the slow climb; heat tracks the event — the fan that seizes, the PSU that starts cooking, the cable that melts. A small heat-rise or smoke detector in the closet, wired to the building's alarm if there is one, or a standalone unit at minimum, is the difference between "the room was 41°C last Tuesday" and "the room was on fire last Tuesday." Combustible storage is what turns heat into fire (see the audit, section 1); the detector is what turns fire into a Tuesday instead of a total loss.
- The door alert answers a different question: who was in the closet? A cheap door-open chime or contact sensor turns "someone unplugged the switch last week and it never came back" from an unsolvable mystery into a timestamped event. It also enforces the visitor rule: anyone non-team in the closet gets logged, the same rule as the visitor log — the closet holds the machines that run payroll, and payroll deserves a door that says who came through it. If the closet is on the badge system, tie the alert to the badge log instead of buying a second sensor.
- Physically secure the closet like the asset it is. A lock on the door (keyed to the key register), no ceiling-tile access from the next room, and a note on the door about what's inside and who to call. Small offices treat the telecom closet as shared storage until the day the shared storage takes the file server with it. The physical security checklist walks the office; add this door to its route permanently.
- Route every room alert into the same channel as everything else. Temp, water, smoke, door — if they can push to your alert channel, wire them there. One-glance triage beats four apps: the incident response system exists so that a 2 a.m. water event and a 2 a.m. cron failure land in the same inbox with the same severity labels. A room that alerts in a channel nobody reads is a room that monitors nobody.
5. The monthly glance: one row, five numbers, five minutes
- One row per month, in the log the whole team can see. Five columns: hottest temp (and when), coldest, humidity band hit, sensor tests done (UPS drill, leak, door), anything changed in the room. Five minutes, once a month, in the same sitting as the weekly review's monthly cousin or the server monitoring glance. The row exists to make the trend visible: three months of climbing peaks is a fan dying politely, giving you a whole quarter of warning.
- The log is also your evidence file. Insurers, landlords, and auditors ask the same family of questions as the UPS drill logs: do you monitor the environment, do you test the sensors, show the records. A year of monthly rows is a PDF answer. It's the same "evidence, not reassurance" posture as the annual security review — the page that gets read to you after an event should be one you wrote before it.
- Put the seasonal jobs on the calendar, not on memory. Winter: check the humidity floor and static risk, confirm heat doesn't fight the closet (radiator against the rack wall is a classic). Summer: re-run the afternoon peak measurement, clean intakes, confirm the closet door wedge survived the year. These land on the patch cadence calendar as environment rows because they're maintenance with a schedule, not chores with a mood.
- When the room fails the trend, treat it like any other incident. Threshold crossed for real: capture the timeline, triage with the incident plan, and if hardware took damage, walk the recovery the way the disk full runbook and ransomware recovery checklist structure theirs — stabilize, assess, restore, then write the one-paragraph postmortem that updates the threshold so the same event can't surprise you twice.
Small-team honesty note: you do not need a building management system, a raised floor, or a dedicated cooling unit — those belong to the data centre. What a ten-person office needs is a week of baseline readings, four thresholds written on the wall card, a $30 leak sensor under the rack, a smoke detector in the closet, alerts in the channel the team already reads, and a monthly row that makes the trend impossible to miss. The trap this page exists to prevent is the closet that "runs fine" — because it ran fine right up until the heat wave, the leaking trap, and the seized fan that nobody measured. You cannot alert on a room you never added to the system. Add it.
Related: power failure IT checklist · physical security checklist · asset inventory checklist · server monitoring checklist · severity matrix · badge access control · key inventory register · visitor log template · cron job monitoring · AI incident response system · incident response plan · disk full runbook · maintenance window policy · patch cadence calendar · weekly review checklist · annual security review · new admin's first week · server hardening checklist · office Wi-Fi security · runbook template · ransomware recovery checklist