Back to blog
2 min read
AlertingOn-call

What changed after we stopped paging on warnings

Warning-level pages feel responsible until nobody sleeps. Moving warnings to daytime channels restored trust in the alerts that still wake people.

NK

Nabin Khair

Founder

I have watched more than one team wire disk-at-70%, certificate-in-30-days, and slight latency bumps into the same escalation as "API down." It feels thorough. It produces a culture where the loudest sound in ops is ignored.

What we changed

We applied a hard filter: night pages require customer impact or imminent data/security harm. Warnings became:

  • Slack #ops-warnings in business hours
  • Ticket or issue for capacity and expiry work
  • Email digests where useful

SEV1-style pages stayed rare and sharp (page vs notify).

What improved

  • Ack times on real incidents dropped; people believed the tone again
  • Warning work still happened, just in daylight when thinking is cheaper
  • On-call stopped bargaining with "is this another warning?"

What did not break

Certificates still got renewed because calendar alerts existed, not because we woke someone at 3am thirty days early. Capacity still got planned because weekly review looked at the warning channel.

How to try it for two weeks

  1. List last month's pages. Tag warning vs action.
  2. Demote the warning class to a non-escalating channel.
  3. Keep a written exception list (e.g. "disk >95% pages").
  4. Review whether any demoted warning became a surprise outage. If yes, promote that specific case, not the whole class.

Paging on warnings is how you spend trust on things that were never emergencies. Spend trust on the emergencies.

Keep reading

More from the Tallwatch blog

More on monitoring, alerting, and status pages.