Back to journal

ALERTING / 5 MIN READ

How to Get Alerts When Your Website Goes Down

An alert is useful when it reaches the right person with enough context to act. The goal is not to send an email for every unusual request; it is to notify your team when a meaningful page failure has been confirmed.

Choose the pages that deserve an alert

Start with routes that affect visitors, revenue, or support. A checkout failure may need an immediate response even when the homepage is healthy. A private admin route may need a different priority or no public status update at all.

Give every monitor a clear name. An alert that says “Checkout page” or “Customer sign-in” is easier to act on than an alert that only contains a long URL.

Confirm failures before notifying people

A consecutive-failure threshold helps prevent noisy alerts from temporary packet loss or a one-off remote error. Choose a threshold that balances speed and confidence for the importance of the page and its monitoring interval.

The alert should include the affected page, URL, first observed failure time, current error evidence, and a link to the incident history. This gives the recipient a starting point without claiming a cause that has not been verified.

Send a recovery alert too

Recovery messages close the loop. They tell the team that a successful check was observed after the outage and provide a useful time for post-incident review. A recovery alert does not prove every visitor is fixed, but it confirms the monitor can again receive the expected response.

If you publish a status page, update it with clear language during an incident and after recovery. This gives visitors a reliable place to check instead of asking support for every update.

Keep alerts actionable

Review your alert settings after a few incidents. Too many notifications make real issues easier to miss; too few monitored pages leave important failures invisible.

  • Enable alerts for active, visitor-facing monitors.
  • Use a realistic consecutive-failure threshold.
  • Keep outage and recovery alerts enabled for important pages.
  • Review incident evidence before declaring a root cause.
  • Use a public status page when customers need an update.