Monitor User Guide
1. Executive Summary
The Monitor module (aegis-monitor) is the always-on health check for every
active route. It answers two questions on a schedule: is the page up? and does
the page still look right?
Note for anyone reading older docs or decks: there used to be a separate
visual-monitorworker. It was removed from the repo (commit24af897) and its job now lives here. Monitor holds the Browser Rendering binding, the screenshot bucket, and all the drift logic. There is one monitor, not two.
2. What runs, and when
Monitor is driven by two cron triggers:
| Cron | What it does |
|---|---|
15 * * * * (hourly) |
Uptime + content pass across active routes. Screenshots are skipped by design. |
30 3 * * * (daily) |
Retention cleanup — purges visual_audits rows older than 30 days and their R2 screenshots. |
Visual capture is the expensive part, which is why the hourly pass skips it. A single route can be re-shot on demand (see §5) without waiting for a full pass.
Which pages: every active page with Monitor this page ticked (the page's wizard, QA & Monitor step; on for every page by default). Untick it and the page is left out of the rotation and every sweep; QA is separate (QA this page). A check you start by hand on one page still runs.
3. Uptime and content checks
- Endpoint pinging: fires real HTTP requests at live landing pages and critical endpoints, following the route's configured origin.
- Status verification: confirms routes return
200, not a500or an unintended404. - Latency tracking: flags backend origins (Lovable SPAs, Shopify) that are degrading.
- Request pacing: requests are deliberately paced. Firing a full sweep flat
out triggered Cloudflare rate limiting and produced
429false positives, so the sweep is throttled rather than parallelised.
4. Visual drift detection
Monitor renders each page in a real headless browser (Cloudflare Browser Rendering), screenshots it, and compares it against a stored baseline.
- Baselines live in
route_visual_baselines; captures and diffs live in R2 (SCREENSHOTS_BUCKET) with rows invisual_audits. - Rolling baseline: on a clean capture (≤ 0.5% drift) the fresh screenshot is promoted to the baseline. This matters — without it, run-to-run rendering noise accumulates until it trips the threshold and reports drift that isn't real. On genuine drift the last-good baseline is kept until a human reviews it.
- Ignore selectors: dynamic overlays are hidden before capture, via a global
list (
GLOBAL_IGNORE_SELECTORS) and a per-route list editable in the Admin dashboard. The default global entry targets the Market-Distributor localisation popup ("We don't ship to …"), which only appears in the headless geo context and would otherwise register as drift on every run. - DOM hashing: a DOM hash is stored alongside the screenshot so structural changes can be distinguished from purely visual ones.
Reading an audit report
/api/report/:audit_id shows Baseline, Current and Diff side by side. If the
change is legitimate, Image OK promotes the current capture to the baseline in
one click. Report images are readable without a session so the links work
straight from Slack.
5. On-demand checks
GET /api/trigger runs a pass immediately. Useful parameters:
route_id— check a single route instead of everything.force_hourly— run the hourly configuration on demand.force_visual— capture screenshots even under the hourly config, which normally skips them. This is what the per-route "Run Visual Check" action in the Admin dashboard uses to re-shoot and re-baseline one route.force_slack— send the Slack report even if the run would normally be quiet.
The endpoint is behind Basic Auth. Credentials are validated against the users
table, then ADMIN_PASSWORD, then — if that isn't set — any value present in the
app_tokens table.
6. Incident alerting
- Failures are only escalated after consecutive misses, to rule out momentary network blips.
- Alerts go to Slack with the URL, the status code, and the time of failure; drift alerts include links to the three-image report.
- Alert links point at
aegis.purdyandfigg.dev.
7. Data retention
Monitor is the only module with a full retention story, and it's worth knowing what it does and doesn't cover:
| Table | Retention | Owner |
|---|---|---|
visual_audits + R2 screenshots |
30 days | Monitor daily cron |
route_logs |
48 hours | Monitor |
qabot_test_runs |
72 hours | Admin daily cron |
analytics_logs |
365 days | Admin daily cron |
click_events |
365 days | Admin daily cron |
The screenshot cleanup is reference-safe: it never deletes an object still referenced by a retained audit row or by an active baseline.
Analytics retention used to run inside the Analytics worker's queue consumer, on every batch. It now runs once a day from the Admin cron and is indexed. The window is 365 days rather than 90 on purpose: there is no archive behind the purge yet, so anything it deletes is gone once D1's 30-day Time Travel passes. Shorten it once day-partitioned exports are landing in R2.