Tallwatch
Back to blog
2 min read
AlertingOn-call

What changed after we stopped paging on warnings

Warning-level pages feel responsible until nobody sleeps. Moving warnings to daytime channels restored trust in the alerts that still wake people.

NK

Nabin Khair

Founder

What changed after we stopped paging on warnings

I have watched more than one team wire disk-at-70%, certificate-in-30-days, and slight latency bumps into the same escalation as "API down." It feels thorough. It produces a culture where the loudest sound in ops is ignored.

What we changed

We applied a hard filter: night pages require customer impact or imminent data/security harm. Warnings became:

  • Slack #ops-warnings in business hours
  • Ticket or issue for capacity and expiry work
  • Email digests where useful

SEV1-style pages stayed rare and sharp (page vs notify).

What improved

  • Ack times on real incidents dropped; people believed the tone again
  • Warning work still happened, just in daylight when thinking is cheaper
  • On-call stopped bargaining with "is this another warning?"

What did not break

Certificates still got renewed because calendar alerts existed, not because we woke someone at 3am thirty days early. Capacity still got planned because weekly review looked at the warning channel.

How to try it for two weeks

  1. List last month's pages. Tag warning vs action.
  2. Demote the warning class to a non-escalating channel.
  3. Keep a written exception list (e.g. "disk >95% pages").
  4. Review whether any demoted warning became a surprise outage. If yes, promote that specific case, not the whole class.

Paging on warnings is how you spend trust on things that were never emergencies. Spend trust on the emergencies.

Related

Keep reading

False alerts and status pages.

How to cut false downtime alerts without waiting forever to get paged

How to cut false downtime alerts without waiting forever to get paged

You should not have to choose between a quiet pager and a slow one. Confirmation — across time or across regions — is how you keep both.

NKNabin Khair
Webhook alerts that silently fail: how to know your pager is broken

Webhook alerts that silently fail: how to know your pager is broken

The scariest failure mode is not a loud false page. It is an outage with zero notifications because the webhook path died quietly.

NKNabin Khair
How to test your alerting before the real outage

How to test your alerting before the real outage

The first time your escalation policy runs should not be a customer emergency. Break something on purpose and watch who gets paged.

NKNabin Khair