QA Runner Operations
Start with Testing Shopify pages. That is the short version of the whole flow — preview against the Range Plan, live against its own baseline, and what happens in between. This page is the longer reference behind it, and parts of it predate the current flow.
How Aegis QA actually executes: who runs it, on what hardware, under what authority, and how the results get back into the platform.
For the who is Agatha framing, see Meet Agatha. For how test specs are written and signed off, see the QA User Guide. This page is the operational mechanics.
1. The shape of it
There is no bespoke test-runner service. The runner is an AI agent — Claude, driving a real browser, holding a skill prompt, talking to Aegis over MCP.
Aegis (Cloudflare) Local runner (Mac Mini)
┌────────────────────┐ ┌──────────────────────────┐
│ qabot_test_queue │◄─ claim ────│ Claude + Aegis MCP │
│ qabot_test_versions│── spec ────►│ + Claude in Chrome │
│ qabot_test_runs │◄─ report ───│ (real browser, real DOM) │
└────────────────────┘ └──────────────────────────┘
│
└──► Slack (failures) ──► Admin dashboard (history)
Three properties follow from that design, and they're the reason it was built this way:
- The browser is real. Most Aegis routes proxy a client-rendered Lovable origin. A plain HTTP fetch returns an empty shell, so anything that isn't a real browser sees nothing. Claude in Chrome executes the page.
- The test is prose. Anyone can read it, run it, or change it — including people who will never open an IDE.
- The runner is portable. The same Markdown spec runs from a scheduled Mac Mini, from a developer's Claude session, or from CI. There is one source of truth for what "correct" means.
2. Where a run comes from
A run can be triggered three ways, all executing the same specs:
| Trigger | Cadence | Typical use |
|---|---|---|
| Scheduled runner on a Mac Mini | Hourly | The standing regression sweep |
| A person, via Claude | On demand | "Is this page OK right now?" |
| CI, on code change | Per push | Catch a regression before it ships |
The Mac Minis matter for a practical reason: Claude in Chrome needs a real, logged-in desktop browser session. A Cloudflare Worker can't provide that. So the control plane is at the edge and the hands are on a machine in the office.
What the queue will not hand you
Three gates sit between a route and a runner, and two of them surprise people:
- A page whose last run failed and nobody has triaged. It is not offered to the schedule or to a forced request. Running it again could only produce a second failure nobody has looked at, or a pass that overwrites the first before anyone saw it. Triage is the only thing that restarts it.
- A page that has not changed since QA last looked. The skip assessment passes it over, and forcing does not override this — the poll never tells the assessment a route was forced. So "Run QA Test Now" on an unchanged page runs nothing and reports success.
- The pass gate. Between passes the poll answers
pass-not-duewithout reading a single queued row, so a queued route can legitimately wait hours.
Anything held back for the first reason appears on Pages that require review by a human in the admin, with the reason spelled out. Up next and that report are complements: one is what the runner will be given, the other is what it will not, and nothing falls between them.
3. What a scheduled run does
The runner skill (aegis-qa-runner-hourly.md, stored in qabot_test_templates
and served by the setup_qa_runner_skill MCP tool) runs two phases. Order
matters — running every existing test comes before authoring new ones, so a
backlog of missing tests can never starve the regression sweep.
Feeding a lesson back (push_playbook_rule)
Added 2026-09-09. When a run learns something that is true of the estate
rather than of one page, it can add it to the fleet-wide playbook itself instead
of writing it into that route's content.test.md — a rule left in one route's
test is a rule the next route has to learn again.
The worked example is the rounding rule: a runner failed /collections because
37.66% was badged 37%, and the correction applied to every route with a sale
price. Before this existed, moving it fleet-wide took a hand-written SQL file, a
branch and a deploy.
The test is whether the correction would apply to a page nobody has looked at yet. "A save percentage may round either way" is fleet-wide. "The hero on /summerhosting says Save 30% and the Range Plan says 28%" is a defect on one page.
It is a proposal, not a change. The rule goes to a superadmin for review and does not apply until they approve it — Settings → Playbook Rules, with a Slack notification when one arrives, because nobody watches a queue. A runner that has proposed a rule keeps following the playbook as it stands.
The gate is the same one a new QA test goes through: tests land as preview and a human promotes them. Append-only alone was not enough, and the reason is worth stating — it limits the blast radius of a bad rule without stopping it taking effect. Between a run adding a rule and someone noticing, every runner on every route is following it.
Append-only still holds past the gate: approving appends, and nothing can alter or remove a rule already in the document. Every entry is dated and attributed to the run that produced it. Proposing the same rule twice is safe — you will be told whether it is already in the playbook or already waiting.
Phase 1 — run every queued test
- Claim the next route from the queue using this run's
runner_id. - Resolve the real target URL. Trust the URL in the latest spec, not a
hardcoded root — the root legitimately varies (UK apex,
us.domain, theget.proxy). - Execute the checks in a real browser: content, GA/pixel, and cart.
- Submit a Markdown report with per-check pass/fail and shown-vs-expected for every failure.
- Repeat until the queue is empty, or until a route repeats — which means the queue has wrapped and the run is done.
Phase 2 — draft tests for routes that have none
Only once Phase 1 has fully stopped. For each route still missing a suite: open
the live page, inspect the sections, and draft a content.test.md from the actual
rendered copy, layout and CTA destinations, following the route's test-brief.md.
Save it as preview.
Cleanup
Clear the Shopify cart so no test items are left behind, and close the browser tabs the run created.
4. Non-negotiable guardrails
These are enforced at the scheduling level, not left to individual test files:
- Never fabricate a result. If the browser is unavailable or a step can't be
verified, submit status
errorwith an explanation. A green run that didn't happen is worse than no run. - If the browser isn't connected, stop. Report it and do nothing else. No partial runs, no assumed passes.
- Hard stop at cart verification. Never proceed to checkout or payment.
- Exclude test traffic. Use a test parameter or flagged session, or target a staging store, so hourly runs don't pollute GA4, Elevar or Meta CAPI.
- Clear the cart between batches, and click real CTAs rather than posting to
/cart/add.js— the raw post skips the on-page discount flow and manufactures false price failures.
5. Authority: propose, don't activate
The single rule that makes an autonomous QA agent safe to run against production:
The agent may propose. Only a human may activate.
| Action | Agent | Human only |
|---|---|---|
| Read a test | ✅ | |
| Scrape a live page to author a test | ✅ | |
| Dry-run a test in the browser | ✅ | |
| Create a test for a route with no baseline | ✅ (lands as preview) |
|
| Snapshot the current baseline as a version | ✅ | |
Save a draft (make_default=false) |
✅ | |
| Make a version the live baseline | ✅ | |
| Quarantine a route | ✅ | |
| Un-quarantine / resume a route | ✅ |
Pausing is safe, so agents may pause. Resuming and activating are judgement calls, so they are not delegated. This holds even when the agent is confident a change is benign.
6. Triaging a failure: drift vs breakage
This is the part that most matters, and the easiest thing to get wrong. A failing test has two opposite causes, and the correct response to each is the opposite of the other.
Never wire "fail → refresh the baseline" together. That heals the alarm instead of the fire: the baseline is rewritten to match a broken page, the test goes green, and a real bug is now invisible. This is the classic self-healing-test failure mode and it is explicitly forbidden.
| Classification | What you're seeing | Correct action |
|---|---|---|
| Drift | Content, price or label changed, but the flow still works end to end — item adds, subtotal updates, removal resets. | The test is stale. Snapshot the current baseline, propose a refreshed one as a draft, quarantine as QUARANTINE-DRIFT, and ask a human to approve the diff. |
| Breakage | The flow itself fails — add-to-cart throws, subtotal stuck at zero, cart won't open, JS error. | The site is wrong. Do not touch the baseline. Quarantine as QUARANTINE-INCIDENT and raise it as an incident. |
| Unsure | Classification is ambiguous. | Quarantine as QUARANTINE-ESCALATE and escalate. Never auto-refresh on uncertainty. |
The quarantine tag tells the human which conversation they're having: approve a new baseline, investigate a live bug, or classify this for me.
7. How this rolls into the wider platform
The runner is one input into a system that already exists, which is why QA results show up in places you'd expect:
- Slack — failures are reported where the team already is.
- Admin dashboard — run history and reports alongside the routes themselves.
- Monitor — uptime and visual drift run on their own hourly cron, feeding the same alerting.
- Route sign-off — a route cannot take live paid traffic until Page Author, Growth and Tech have all approved. QA evidence is what makes that approval meaningful rather than ceremonial.
8. The agentic-employee model
The QA runner is the first instance of a pattern the platform is built around: a named agent, with a defined job, a bounded authority ceiling, and an audit trail — reachable in the tools the team already uses rather than behind a new dashboard.
Concretely, that pattern is:
- A named role, not a feature. Agatha owns QA; Alfred is the tech-lead role that reviews code quality at build time. Naming the role sets the expectation that it has responsibilities and limits, like any other colleague.
- A skill prompt as the job description. The runner skill lives in the database, is versioned, and is served over MCP. Changing how the job is done is an edit to prose, not a deploy.
- MCP as the only hands. Every action an agent can take is an MCP tool. That's what makes the authority ceiling in §5 real: what isn't exposed as a tool cannot be done.
- Human activation gates on anything irreversible. Propose freely; activate never.
- Reachable where the team is. Claude and Slack, not a bespoke UI.
Where this goes: a page is built, added to Aegis, and its tests are written as it's built. Anyone in the company can run a test on any page at any time via Claude. Tests run on every code change and every hour. Failures land in Slack instantly. The human role shifts from running the checks to managing the system that runs the checks and fixing what it finds — less manual verification, more time on what actually broke.
9. Known gaps
Honest open items, so nobody discovers them the hard way:
- The runner is a schedule, not a service. If the Mac Mini is asleep, offline or logged out, runs silently don't happen. There is no alert for "no runs in the last N hours".
- Stuck queue rows. A runner that dies mid-route leaves its row in
processinguntil the queue is reset. There's no automatic reclaim of a stale claim. - Volume limits are unproven. Nobody has established how many behavioural runs per hour production tolerates before bot protection or analytics noise becomes a problem. This may yet force the behavioural lane onto staging only.
- Notification format isn't settled. A
QUARANTINE-INCIDENTalert needs to be actionable at a glance, and it currently isn't clearly distinct from a drift notice. - Report pages are unsanitised and the report-submission tool is unauthenticated. See the handover security notes; this needs closing before the QA surface is widened.
- Run QA Test Now is offered when it cannot work. On a page awaiting triage, or one that has not changed since QA last looked, it reports success and runs nothing. The behaviour is intended; the silence is not. The button should be unavailable with the reason in its place.
- A drafting claim is never released. Nothing sets
qabot_tests.drafting_runner_idback toNULL, so every page ever drafted carries the id of the runner that did it, for ever. Harmless where it is read today — the poll checks the lock age — but anything reading the column bare will conclude the page is being drafted when it is not. - Signing off does not remove the page from the review report until a refresh. The queue picks it up within a minute, which is fine; the report underneath keeps showing it until the screen is reloaded.