Help & FAQ
Short answers to the things people ask most. If you want the detail behind an answer, each one points at the guide that covers it properly.
Something missing or wrong here? Add it — this page is document-module/src/faq.md
in the staticproxy repo, and it publishes on the next merge to main.
Getting oriented
What is Aegis, in one paragraph?
(The longer, non-technical version — the one to send somebody who has never heard of it — is What is Aegis?.)
Aegis is the layer that sits between a paid ad click and a landing page. It routes the visitor to whichever backend actually hosts that page (Shopify, a Lovable React app, Cloudflare Pages), injects the tracking the page needs, records the click first-party, and then keeps checking — every hour — that the page still works. It's built as a set of small Cloudflare Workers sharing one D1 database.
Which module do I need?
| I want to… | Module | Where |
|---|---|---|
| Add, edit or promote a route | Admin | aegis.purdyandfigg.dev |
| See ad clicks, channels, attribution | Admin → Analytics | aegis.purdyandfigg.dev |
| Know why a page went down or looks wrong | Monitor | aegis-monitor.purdyandfigg.dev |
| Write or review a QA test spec | Admin → route Options → QA Test | aegis.purdyandfigg.dev |
| See what a test found, or triage a failure | Admin → route Options → QA Status | aegis.purdyandfigg.dev |
| Check ad copy for compliance | Passmark | via Admin → Compliance |
| Chase a Trello landing-page project | Nudger | Nudger dashboard (low activity — see below) |
| The proxy itself | Router | no UI; it's the edge worker |
Where's the developer documentation?
Right here. User Guides for how each module behaves — if you are
taking over QA, start with the
QA Manager Handbook, which covers the whole
job end to end. Then Core Concepts for how the pieces fit and why,
and API Reference for the endpoints. The API reference is generated from
the worker source by npm run docs:generate --workspace=document-module, so if
an endpoint is in the docs it exists in the code.
Routes and pages
I added a route and it 404s / shows the wrong page.
Work through it in this order:
- Is the route active? A route that hasn't cleared sign-off is not serving.
- Is it signed off? Four approvals are required before a route can be activated — Page Author, Head of Growth, Head of Tech, and the AI Evaluator (compliance). This is deliberate: it stops paid budget pointing at an unfinished page. The first three are ticked by default on a new route; the AI sign-off is not, so it is usually the one blocking activation. A human can override the AI sign-off, but the override requires a written reason and is recorded.
- Is the origin reachable? Open the origin URL directly. If the origin is
down, the router falls back to
DEFAULT_ROUTE_URLrather than erroring, so a "wrong page" often means "origin is down". - React app showing its own 404? The SPA was compiled for a different basename than the path you mapped. See Router User Guide.
Why does a page have its own pixels when there's a global set?
Per-route pixels override the global set for that route. Historically several routes had per-route pixels attached by hand, which is why they can drift from the global config. New routes can have pixels auto-applied via the per-pixel auto-apply flag. If you're unsure whether a route should have its own, ask before removing them — a route may have been given a dedicated pixel for a specific campaign.
What does "locking" a route do?
It marks the route as locked and sets an edit password, so it can't be casually changed by someone who doesn't know it's load-bearing. It's a "don't-touch-this" guard, not a security boundary — treat it that way.
Can I point a route at somewhere other than Shopify or Lovable?
Today: Shopify, Lovable, and Cloudflare Pages origins. R2 and Render are requested but not yet available as configurable origins in the dashboard.
Analytics and tracking
The revenue / conversion numbers show £0. Is it broken?
No — it isn't connected yet. Aegis records clicks end to end. Order and
revenue data requires the Shopify orders/paid webhook to be exposed,
secretted, registered and enabled, and that work is still outstanding. Until
then, the "Performance by Page Type" panel deliberately ghosts itself rather
than showing a misleading £0 table.
Treat click data as reliable and revenue/ROAS as not yet live.
What does "GA4 Firing Rate" mean, and why isn't it 100%?
It's the share of clicks where the browser confirmed GA4 actually loaded. The server logs each landing as pending, and a heartbeat inside the GA pixel confirms it once GA4 initialises. A miss usually means an ad-blocker, a privacy-focused browser, or the visitor left before GA4 loaded — not a bug. Below ~95% is worth investigating; single-digit misses are normal.
Why do some clicks have no UTMs?
Either the ad creative wasn't tagged, or the visitor arrived without them. Aegis locks the marketing source into a server-side session store on first landing and restores UTMs to the address bar if they're lost on a refresh — but it can't invent tags that were never set. High untagged traffic is a signal to audit the ad platform's tracking templates.
A single visitor shows hundreds of sessions. Real?
Historically no. A cookie-collision bug (fixed in v1.0.56) could collapse many
browsers into one visitor_id via shared-cache prefetch. The fix is live and the
distribution panels now bucket the long tail so old collision rows don't distort
the charts. Rows in the flagged 10+ bucket from before the fix are noise.
How do I get the raw data out?
GET /api/analytics/export with format=csv|json, plus optional
source, medium, campaign, start, end, limit. Authenticate with an app
token. start and end accept either ISO timestamps or plain YYYY-MM-DD. See the Analytics User Guide.
Paging through a lot of rows? Use cursor, not offset. Pass
meta.next_cursor from each response back as cursor on the next request, and
stop when has_more is false. offset still works, but the database has to walk
and discard every row before the offset, so each page costs more than the last —
and worse every month as the table grows. On 2026-09-06 a consumer paging to
offset=500000 produced repeated 46–50 second queries and slowed unrelated parts
of Aegis, because D1 handles one query at a time per database. Cursor paging
returns the same rows in the same order at a constant cost per page.
QA and testing
Who actually runs the tests?
An AI agent, on a schedule — not a person. Test specs live in the database as Markdown; a Claude-based runner claims routes off a queue, drives a real browser, and posts a pass/fail report back. See the QA User Guide and the QA Runner Operations Guide.
Not every route is run every hour. A page that has not changed since it was last tested is skipped, and the skip is recorded carrying the previous verdict so the history stays continuous — with a backstop that forces a real check after a couple of days whatever the signal says.
A test failed. What do I do?
Read the report first and decide which of the two things happened:
- The page changed (price, copy, CTA label) but still works → the test is stale. Approve a refreshed baseline.
- The page is broken (add-to-cart fails, subtotal stuck, JS error) → the site is wrong. Treat it as an incident. Do not refresh the baseline.
Getting this backwards is the one genuinely dangerous mistake in QA: refreshing a baseline over a real bug makes the test go green and hides the breakage.
Then write the answer down. In Options → QA Status, a failed run takes Known defect (the page is wrong — files a defect against the route), Not a defect (the test is wrong), or Investigating. Accept is separate and answers what now?: the run stays red in the record, but everything downstream reads it as a pass.
The runner usually tells you what it thinks, as a proposal on the report — it had the page open and did the work already. It is never a verdict: a runner that can dismiss its own failure can make any test pass.
Can I run a test myself, right now?
Two ways:
- From Admin — on the route row, Options → 🧪 Run QA Test Now. This pushes the route to the front of the runner queue and bypasses the skip check, so it is really run rather than answered with last week's verdict; the next free runner picks it up.
- Through Claude, using the Aegis MCP connection, if you'd rather not open the dashboard. See Claude MCP Setup.
A run you asked for reports back to Slack either way — pass or fail. Scheduled runs stay failure-only, which is what keeps the channel readable. The report names you beside the agent that ran it, and the run does not count toward the 3-strike demotion, so you can re-run a page you are fixing without demoting it.
Note that Run visual check is a different thing — it captures a fresh screenshot via the Monitor (~15s), it does not run the QA spec.
Where do I see everything that's broken right now?
The day's findings — /api/report/findings, also linked from the user menu,
from the Forge report and from the daily Slack post. Every open defect grouped by
store, then severity, then owner, with the run that found each one linked.
Two things on it are worth knowing before you read it. Coverage is stated at the top and called partial below two thirds — a short list means less when half the estate has not been checked. And "not defects" are only the ones a person judged: an expiry after seven days is not a decision and is kept out, which is the difference between a list of 47 judgements and a list of 443 rows.
The severity, store and not-a-defect filters live in the URL, so a narrowed view can be sent to whoever owns it.
Why is a test I just created marked "create test" or "preview"?
Because a new test is a proposal until a human promotes it. A route is born in create test — nothing written yet — so an empty suite reads as work outstanding rather than as a test that passes because it asserts nothing. Agents may draft and snapshot tests freely; making a baseline live is always a human action.
A signed-off page went back to preview on its own. Why?
Three consecutive failures. A signed-off baseline is a claim that the page is verified, and three failures in a row make that claim false, so the suite goes back to "a human needs to look".
What does not count toward it: a skipped run, a run from someone's own Claude
session, a run you asked for with Run QA Test Now, a failure you triaged or
accepted, a preview route, and an error (the runner could not reach an answer,
which is not evidence the page is wrong). The counter measures failures nobody
has looked at, and each of those is either someone looking or not a failure.
What is the GA Framework, and what does it check?
Three levels, and reading them in order is the whole thing: the framework says
what each data-module carries, a page type says which modules it has, and
a page is measured against its type. Admin → GA Framework holds the
module list — elements, a description, a screenshot, and valid / unsure so a
name nobody here can decode becomes a list to take to the developers. It is a
copy of the Confluence SSOT and links back to it.
The check itself runs in the Monitor, not as a QA test: asserting that an attribute exists is a cheap job that was being done by the most expensive runner on the estate.
ga_check_mode decides what a mismatch does — off, log (record it, raise
nothing) or on (file a defect). It is log today, and on on only a changed
value files anything: all eleven page types are still seeded with all
thirty-eight modules, so a missing module usually means the page type has not been
trimmed yet rather than that the page is wrong. Defects it files carry a GA pill,
one per page.
What happens when I take a page live?
A modal walks the steps and each one reads a fact the server just returned, so it can come back amber rather than pretending. It checks sign-off, makes sure a test exists, adds the route to the monitor, and — on a Shopify-hosted page — scans it once to seed the GA framework from what is actually there.
Disable runs the same thing in reverse, with one honest ending: if the page is not proxied, disabling it in Aegis does not take it down. It stays live at its host until you disable it there, and the last step says so.
Why did the monitor flag a visual change when nothing changed?
Usually a dynamic overlay (geo/market popup, cookie banner) or rendering noise. Per-route and global ignore-selectors hide known offenders before capture, and a clean capture rolls the baseline forward so noise can't accumulate. If a specific route keeps drifting, add its overlay selector to that route's ignore list.
Access and setup
How do I get access?
Ask a superadmin to add you in Admin → Users. Permissions are per-area (routes, pixels, checkouts, analytics), so ask for what you need rather than blanket superadmin.
How do I connect Aegis to Claude?
Aegis is already registered in the Claude organisation settings — enable it from Customise in Claude. Full instructions, including how to refresh the tool list after a deploy, are in Claude MCP Setup.
I can't see a tool in Claude that I know exists.
Claude fetches the tool list on startup. Quit Claude completely (Cmd + Q) and reopen. If it's still missing, the change may not be deployed yet.
Can I run Aegis locally?
Yes — each module is an npm workspace with its own dev server
(npm run dev:admin, dev:router, and so on). You'll need a .dev.vars file per
workspace. The repo README.md lists exactly which variables each one needs and
where to get them.
Deploys and safety
How do I deploy?
Merge to main. GitHub Actions deploys Router, Admin, Monitor, Analytics,
Ecommerce, Nudger, QA Bot, the docs site, and the Hetzner fallback node — all of
them, every time, deliberately. On 2026-09-16 Admin shipped without Monitor and
the two ran different copies of the same function for an hour, writing
disagreeing verdicts into one table.
QA Bot is still in that list but no longer has a UI: it exists to redirect old links into Admin.
Tests run on every branch push, not just on main, and every deploy job is
gated behind them — so you find out a change is broken when you push it rather
than when you merge it. The gate runs the router unit tests, checks every worker
parses, checks every worker builds, and checks the committed admin stylesheet is
up to date.
Nudger is not in the CI pipeline — it deploys manually.
Can I let an AI agent deploy for me?
No. Deployment is manual by policy, and the same applies to database schema/data changes: an agent gives you the command, you run it. See Deploy Guardrails.
What happens if Cloudflare or D1 goes down?
The router can run entirely from a synced JSON config held in edge memory, and the Router plus Admin can be served from a Hetzner node. Background modules (Analytics, Nudger) are deliberately not mirrored — the priority is keeping customer-facing pages and live ad budget alive.
Modules you can mostly ignore
Nudger — is it still used?
Barely. Nudger automates the Trello → Slack → QA landing-page pipeline. It works, but the team's workflow has moved on and it isn't in the CI deploy pipeline. It's deliberately defocused until the wider automation work later this year, at which point its Trello/Slack orchestration gets revisited properly. Don't invest in it before then, and don't be surprised if its dashboard is stale.
Compliance scanning — is it a product?
Not on its own. Ad copy is scanned by Passmark, called from Admin, which returns real CAP rule ids, risk levels, the offending line and suggested wording. You reach it through Admin → Compliance rather than using it directly.
The old in-house Compliance AI worker was retired on 2026-09-05. It scanned copy with an LLM and passed almost everything, so a page could collect an AI-Evaluator sign-off it had not earned.