Build Better Stack
YESreplaces $34/mosaves $408/yrback to the verdict
An uptime monitor you run yourself: a fetch on an interval with a real timeout, a state machine that only calls a site down after two consecutive failures, one chat message on down and one on recovery, and a public status page with uptime percentages, a latency sparkline and an RSS feed of incidents. Honest about the one thing a single box cannot do: tell your outage from your own network's.
Before step 1
Everything below is assumed from the first step. Tick each one when you actually have it, not when you plan to.
- installfree
Why Everything in this build runs on it: the server, the scripts, the tests.
Get it Download the LTS installer from nodejs.org, or install with your package manager (brew install node, or nvm install 22). Restart the terminal afterwards. open ↗
Verify
node --version prints v22 or higher - installfree
Why Every step below is a command you type or a file you edit.
Get it VS Code (code.visualstudio.com), Cursor or Zed. Open a folder for the project and use the editor's built-in terminal. open ↗
Verify
You can open a folder and run a command in its terminal - installfree
Why History for your code, and the way most hosts deploy.
Get it Install from git-scm.com or with your package manager, then run git init in the project folder once it exists. open ↗
Verify
git --version prints a version - API keyfree
Why Alerts go to a chat channel you already watch. A webhook URL is the only credential this needs.
Get it Discord: Server settings > Integrations > Webhooks > New Webhook, copy the URL. Slack: create an app at api.slack.com/apps, enable Incoming Webhooks, add to a channel, copy the URL. Telegram: create a bot with @BotFather and use the bot token plus your chat id. open ↗
Verify
curl -X POST -H 'Content-Type: application/json' -d '{"content":"test"}' <url> posts a message (Discord form; Slack uses a text field) - have readyfree
Why Each monitor needs a URL, an interval, an expected status and optionally a keyword the body must contain.
Get it List every site or endpoint. For each decide: check every 60 seconds or 300, expect 200, and a word that only appears when the page really works.
- accountabout $5 a month
Why A monitor hosted next to what it watches reports nothing when that host dies. Different provider, ideally a different region.
Get it Hetzner if your sites are elsewhere, or vice versa. Smallest Ubuntu 24.04 instance with SSH. open ↗
- accountroughly $10 a year, or free on an existing domain
Why A status page address you can give people, e.g. status.yourdomain.com.
Get it Register at Cloudflare Registrar, Porkbun or Namecheap, or use a subdomain of one you already own. You add one DNS record in the deploy phase. open ↗
- installfree
Why Automatic HTTPS in front of the Node process. Without TLS the browser features this relies on (and your visitors' trust) do not work.
Get it On the VPS: follow the install steps at caddyserver.com/docs/install for Ubuntu. One Caddyfile with your domain and a reverse_proxy line is the whole config. open ↗
Verify
caddy version prints a version on the server
Data model
Create these before the first phase that stores anything. Changing a table later is the expensive kind of change.
- `monitors`: id, name, url, method, interval_seconds, timeout_ms,
expected_status, keyword (nullable), enabled, status ('unknown' | 'up' |
'down'), consecutive_failures, last_checked_at, last_change_at
- `checks`: id, monitor_id, checked_at, ok (bool), status_code, latency_ms, error
- `incidents`: id, monitor_id, started_at, ended_at, cause
Index `checks(monitor_id, checked_at)`. All timestamps are UTC epoch
milliseconds. `incidents` is a separate table from `checks` on purpose: uptime
percentage comes from checks, but the human question ("how long was it down, and
why") comes from incidents, and deriving that from raw checks at read time gets
slow and wrong at the edges.Environment variables
These go in a .env file the app reads at startup. The pack's .env.example is this table as a file · copy it, never commit the filled-in version.
| Variable | Needed | Example | Where the value comes from |
|---|---|---|---|
PORT | required | 3000 | Any free port; Caddy proxies to it. |
DATABASE_PATH | required | ./data/uptime.db | SQLite file. |
ALERT_WEBHOOK_URLsecret | required | https://hooks.slack.com/services/... | The chat webhook from the prerequisites. |
ALERT_FORMAT | required | slack | discord, slack or telegram. |
FAILURES_BEFORE_DOWN | optional | 2 | Consecutive failures before a monitor is called down. |
RETENTION_DAYS | optional | 90 | Days of raw checks to keep. Incidents are kept forever. |
SITE_URL | required | https://status.yourdomain.com | Public base URL for the status page and RSS feed. |
ADMIN_USER | required | admin | Any username for the basic-auth admin pages. |
ADMIN_PASSsecret | required | change-me-to-a-long-random-string | Generate one: openssl rand -base64 24. Never reuse a real password. |
The build, in order
The check
One function from a monitor to a result, with a timeout that actually fires and distinct error kinds.
monitors (id, name, url, interval_seconds, timeout_ms, expected_status, keyword, enabled, status, consecutive_failures, last_checked_at, last_change_at), checks (id, monitor_id, checked_at, ok, status_code, latency_ms, error), incidents (id, monitor_id, started_at, ended_at, cause). Index checks(monitor_id, checked_at).
Files
server.mjsdb.mjscheck.mjsterminalmkdir uptime && cd uptime && git init && npm init -y && npm pkg set type=module mkdir data && cp .env.example .env
fetch with an AbortController timeout, follow redirects, read at most 64 kB of body for the keyword, record latency. Return ok plus one of: dns, tls, timeout, status, keyword, as the error.
- terminal
node scripts/check.mjs https://example.com 200 "Example Domain"
done when · tick each as it passesScheduler
Each monitor on its own interval, concurrently with a cap, and one bad monitor never stops the loop.
Every 5 seconds find monitors due (last_checked_at plus interval before now), run up to 10 concurrently, catch every exception per monitor. On boot treat every monitor as due.
done when · tick each as it passesTransitions and incidents
Down only after N consecutive failures, up on the first success, one incident per outage.
Increment consecutive_failures on a failed check; at FAILURES_BEFORE_DOWN flip to down. Any success resets the counter and flips to up.
Open a row on the flip to down with the first failure's error as cause; set ended_at on recovery. Incidents, not checks, are what people read later.
done when · tick each as it passesAlerting
One message down, one up, retries that never stall the loop.
Down: monitor name and the error. Three retries with backoff, then record the failure and move on.
Duration comes from the incident row's started_at and ended_at.
done when · tick each as it passesStatus page
A public page whose numbers can be argued with, and a feed people can subscribe to.
One row per monitor: dot, state, uptime over 24 h, 7 d and 30 d with the denominator printed (over 2,880 checks), a latency sparkline as inline SVG, recent incidents with durations. Meta refresh every 30 seconds.
Subscription without accounts.
- terminal
node scripts/seed.mjs 90
done when · tick each as it passesRetention and deploy
Old checks pruned, incidents kept, live behind HTTPS on the other provider.
README: the single-region caveat stated plainly, the instruction to host on a different provider than the things watched, and the RSS URL.
Files
deploy/uptime.serviceCaddyfileREADME.md
done when · tick each as it passesOperate it like a productproduct builder
Only for the product-builder path: know when the monitor is down, never lose the database, and keep the server patched.
Answer 200 with the build id and a quick database read. Point a free uptime monitor (or your own, from the Healthchecks entry on this site) at it so an outage is noticed before a user notices.
One JSON line per request: method, path, status, duration, no raw IPs. Rotate weekly with logrotate, keep eight.
SQLite's .backup command makes a consistent copy while the app runs. Copy it to object storage or a second machine; then, once, restore it into a fresh checkout and confirm the app reads it.
terminalsqlite3 data/app.db ".backup '/tmp/app-$(date +%F).db'" rclone copy /tmp/app-$(date +%F).db remote:backups/
Firewall allowing only 22, 80 and 443; unattended security updates on; the app running as an unprivileged user under systemd with Restart=on-failure.
done when · tick each as it passes
That is the whole plan for Better Stack. What it deliberately does not cover is below · check the gaps before you call it a replacement.
- Multi-region probes. One box cannot tell an outage from its own bad network; the status page says so.
- Subscriber email notifications, incident templates, on-call.
- global multi-region probes
- on-call scheduling & escalation policies
- incident timelines and postmortem tooling
- phone-call alerts
- A second probe location that must agree before an incident opens
- Maintenance windows that suppress alerts
Need the files? The project pack on the verdict page hands your agent the whole brief · more uptime.