I have watched more than one team wire disk-at-70%, certificate-in-30-days, and slight latency bumps into the same escalation as "API down." It feels thorough. It produces a culture where the loudest sound in ops is ignored.
What we changed
We applied a hard filter: night pages require customer impact or imminent data/security harm. Warnings became:
- Slack
#ops-warningsin business hours - Ticket or issue for capacity and expiry work
- Email digests where useful
SEV1-style pages stayed rare and sharp (page vs notify).
What improved
- Ack times on real incidents dropped; people believed the tone again
- Warning work still happened, just in daylight when thinking is cheaper
- On-call stopped bargaining with "is this another warning?"
What did not break
Certificates still got renewed because calendar alerts existed, not because we woke someone at 3am thirty days early. Capacity still got planned because weekly review looked at the warning channel.
How to try it for two weeks
- List last month's pages. Tag warning vs action.
- Demote the warning class to a non-escalating channel.
- Keep a written exception list (e.g. "disk >95% pages").
- Review whether any demoted warning became a surprise outage. If yes, promote that specific case, not the whole class.
Paging on warnings is how you spend trust on things that were never emergencies. Spend trust on the emergencies.