A status page
Somewhere to point people when something is wrong.
Build time ~3 hrs
When your service is down, users need somewhere that isn't your service to find out what's happening. A status page is that place — and the one design rule follows directly.
1. Host it somewhere else
If your status page is on the same server as your app, it goes down with your app, which is precisely when it's needed. Static hosting on a different provider, on a subdomain, is the whole trick.
2. Automate the checks
Build a status page. Static hosting, separate from my main app. - A checker running on a schedule, hitting <list the endpoints> and recording up/down plus response time. - Results written to a JSON file the static page reads. No database. - The page shows: current status per component, response time, and a 90-day history bar per component. - A manually-editable incidents file for posting updates during an outage. The checker must run somewhere independent of the thing it's checking. Say where you'd put it.
3. Check something meaningful
Hitting your homepage tells you the web server is running. Hit an endpoint that touches the database and any critical dependency — otherwise you'll show green while every page returns a 500.
4. Write incidents like a human
During an outage people want three things: acknowledgement that you know, what's affected, and when you'll next update. Not a root cause — that comes later.
"We're aware that logins are failing. We're investigating and will update within 30 minutes" is a complete and good update. Silence is what makes people angry, not the outage.
Post the resolution and a short explanation afterwards. It's the part that rebuilds trust.
5. Don't over-engineer it
A single page, a JSON file and a scheduled script is a genuinely good status page. Component hierarchies, subscriber notifications and SLA calculations are for later, if ever.
Test it by deliberately breaking something. A status page that shows green during an outage is worse than not having one.