QA User Guide
Start with Testing Shopify pages. That is the short version of the whole flow โ preview against the Range Plan, live against its own baseline, and what happens in between. This page is the longer reference behind it, and parts of it predate the current flow.
The standalone QA Bot app is gone. Everything below now lives in the Admin dashboard.
qabot.purdyandfigg.devstill answers and redirects, and any old Slack link or bookmark lands on the right screen, but there is nothing to open there any more. This page kept its URL so those links keep working.
1. Executive Summary
QA is where Aegis test specifications are written, versioned and signed off. It is the authoring and review surface; the tests themselves are executed elsewhere, by AI runners (see the QA Runner Operations Guide).
The core idea: a test is a Markdown file in the database, not code in a repo. That's what lets a marketer, a developer and an AI agent all read, run and amend the same test without anyone learning a test framework.
How to open it
On the route row in Admin, Options gives you two entries, and they are two different jobs on purpose:
| What it is for | What you see | |
|---|---|---|
| ๐งช QA Test | Writing and changing the test | The file editor, version history, Create test and Run QA Test Now |
| ๐ QA Status | Reading what the tests found | Run history with triage, and the route's open defects |
Both open over the dashboard rather than in another tab, so you keep the route list behind you. They are two tabs of one screen: switch without losing your place.
A link from outside opens the right screen. ?qa=<route id>&tab=status (or
&tab=test) opens Admin straight onto that route, which is what every Slack
alert, report and MCP reply now links to. A stale qabot.purdyandfigg.dev link
is translated to the same thing.
There is no separate left-hand list of routes any more. Cheap and expensive tests became separate things, so a route has exactly one script and there was nothing for a second pane to list.
2. How a test is stored
| Table | Holds |
|---|---|
qabot_tests |
One suite per route, with a status (preview / signed_off) |
qabot_test_versions |
Every version of every file, with one flagged is_default |
qabot_test_templates |
Fleet-wide templates and briefs shared across routes |
qabot_test_queue |
Which route a runner should pick up next |
qabot_test_runs |
Submitted results and reports (retained 72 hours) |
A suite is a set of named Markdown files, and they do two different jobs. This is the single most important thing to understand about how QA is stored:
content.test.mdandga-tracking.test.mdare tests. A runner executes them.test-brief.mdis not a test. It is the instructions a runner reads in order to writecontent.test.md. Runners are explicitly told to ignore it when running tests.
Anything ending .test.md is a test the runner executes. Nothing else is. The
brief used to be called mcp.test.md, which is why it was so often mistaken for
a third test โ a file that was not a test, wearing a test's file extension.
What each test covers:
content.test.mdโ behavioural and content checks. What the page says, what the CTAs do, whether add-to-cart works and the subtotal updates.ga-tracking.test.mdโ tracking checks. Whichdata-modulesections exist and whichdata-*attributes they must carry. Generated, not hand-written, and no longer edited from here at all: the modules now come from the page type and the framework behind it (see ยง5). It only appears on Shopify Landing pages, because the modules it asserts are Shopify theme sections that a Lovable-rendered page has nothing to match.
Keep the two apart. Tracking assertions do not belong in content.test.md โ
mixing them means a copy change and a tracking regression trip the same alarm and
triage gets muddled.
Files you will not see on a route, because they are fleet-wide config rather
than any one route's tests: aegis-qa-runner-*.md (runner task definitions),
lovable_*_autofixes (rule sets), test-brief-shopify-preview.md and the
per-page-type brief_*.md files (brief templates โ see below). They live in
Global Test Templates, under the profile menu in Admin, with the same version
history, diffs against the live version, and a description and type on each one.
They used to be reachable only through MCP tools, which meant the fleet-wide
config was the one part of QA a person could not read.
A new route asks for its test. Creating a route puts its suite in
create_test rather than seeding a placeholder, so an empty suite reads as "this
needs writing" instead of as a test that passes because it asserts nothing.
Enabling an existing route skips that when a real test is already written.
3. The brief, and how a test gets written
A route's test is not written by hand. The sequence is:
route created โโโบ test-brief.md and content.test.md are seeded as stubs
โ
runner reads test-brief.md, opens the live page,
and drafts a real content.test.md from what it sees
โ
suite goes to preview, awaiting a human
โ
signed off โโโบ runner executes content.test.md each pass
Because the brief and the test are two ends of one pipeline, only one of them is
ever shown. While no real test exists you see test-brief.md, because editing
the brief is how you change what gets written. Once a test has been drafted, the
brief steps aside and you see content.test.md, because that is what now runs.
The banner above the editor always tells you which file you are looking at and what it does.
Asking for a new test. The โจ button beside the brief queues the route for the runner to draft one, jumping it to the front of the queue. Use it after editing the brief, or when a page has changed enough that its test should be rewritten rather than patched. On a route that already has a real test, that button becomes โถ, which runs the existing test instead.
Preview routes and the Shopify brief
A route is in preview whenever it has a Preview URL set (Routes โ edit
route โ Preview URL). The routes list and the QA Bot header both show an amber
PREVIEW pill, and the preview URL becomes the address that gets tested.
Preview pages need a different brief. An unpublished Shopify preview has no live
URL to compare against and no other system holding its intended prices, so its
brief cross-checks prices against the Range Plan Forge app instead of scanning
copy and CTAs. That brief is chosen automatically when a route has a Preview URL
and is a Shopify page (shopify_url set, a shopify:// origin, or a
/pages/ path).
You do not pick it, and it does not appear as a second file. The route still has
one brief called test-brief.md; which template filled it depends on the route,
and the banner says which one you are reading.
There is now a brief per page type (2026-09-09). Each is named after its
type โ brief_pdp.md, brief_landing.md โ and they were cloned from the
original test-brief.md, so they started identical and diverge as people edit
them. A page type with no brief of its own falls back to test-brief.md.
The order is: Shopify preview brief if the route is in preview, otherwise the
page type's brief, otherwise test-brief.md. The preview brief comes first
because it is chosen by the route's state, not its type.
4. Versions and the sign-off gate
Test State has three settings, and only one of them causes anything to happen:
- Create test โ nothing has been written yet. A route is born here, so an empty suite reads as work outstanding rather than as a test that passes because it asserts nothing.
- Preview โ drafted, not running. This is the human review gate: an agent may draft as much as it likes, and nothing it wrote runs until a person signs it off. A suite demoted by three consecutive failures lands back here.
- Sign Off โ the runner executes this route's tests each pass. This is the only state that runs.
A pass is not hourly. qa_sweep_interval_seconds decides it (Settings; 10800 โ
three hours โ on 2026-09-26, not the 21600 this line used to state as fact) โ
and a pass has to drain before the next is due, so a signed-off page is tested
roughly four times a day rather than twenty-four. Within a pass the runner polls
every fifteen minutes and takes one route at a time. Whatever the cadence, no
page goes longer than qa_force_check_days without a full check.
To stop a route being tested, either pause QA on it โ which keeps it live, keeps it in the coverage count and says who paused it and why โ or move the suite back to Preview. (There used to be a third state, Paused. It did nothing โ the runner only ever ran signed-off routes, so pausing changed a badge and stopped nothing. It has been merged into Preview.)
Version history. Every save keeps the previous content. The Version History tab lists each saved version of the open file, newest first, with the live one marked. Click any entry to read it.
Viewing an old version does not change what runs. There is deliberately no one-click rollback: to restore an old version, open it, copy what you need, and save it as the current version. That leaves an honest trail of what changed and when, rather than silently swapping the baseline underneath the runner.
Most files show a single version. History only accumulates from the point the snapshot behaviour was added โ before that, saving overwrote the previous content in place, so nothing older survives.
5. What the tests found โ the QA Status screen
A red run used to have exactly one affordance: read the report. Whatever you concluded lived in your head, the run pruned after 72 hours, and three failures demoted the suite regardless of what the failure turned out to be.
Triage is that conclusion, written down. On a failed run, in QA Status โ Runs:
| What it means | |
|---|---|
| Known defect | The page is wrong. Files a route defect, which is now the thing being tracked. |
| Not a defect | The test is wrong โ a false positive. The most useful of the three, because it is the only one engineering can act on. |
| Investigating | Somebody is looking. |
| Accept | Sits across all three and answers a different question: what now? The run stays red in the record โ overwriting the status would make "this passed" and "somebody accepted this" identical forever โ but everything downstream reads it as a pass. |
Flattening those into one "make the red go away" button is how a defect list turns into a make-the-red-go-away button, which is worse than no button.
The runner proposes; a person decides. A runner that found a failure often already worked out which of the three it was, with the page still open in front of it. It records that as a proposal โ suspect, confidence and a note โ and nothing else. It cannot dismiss its own failure, because a runner that can do that can make any test pass.
Defects are on the second tab. They carry a severity, an owner (content / growth / dev), the run that found them and an occurrence count; filing one notifies Slack, and so does closing it. The fleet-wide view of all of them is the findings report, in ยง8 below.
The same fault, filed twice, is worse than useless โ it makes a list of seventeen problems look like twenty-six. Two things keep it down:
- A fingerprint. A short slug naming the fault โ
meta-description-missing,checkout-marketing-consent-pretickโ carried on the defect and matched before the wording is. Word overlap catches a re-report that reuses most of its words and nothing else: the pre-ticked UK checkout consent box was filed five times estate-wide, and two of those score 0.20 against each other, which is the same ground as two unrelated defects. - A dispute. When a runner re-checks a defect and cannot reproduce it, it says so on that defect rather than filing a retraction as a new one. The defect stays open โ a dispute is a proposal, not a closure โ and shows its reason on the card and on the findings report, so whoever picks it up knows the last look disagreed with the filing. Dismissing the dispute is a separate action from closing the defect, because "the runner was wrong" and "the runner was right and this is fixed" are different conclusions.
The 3-strike demotion. Three consecutive failures on a signed-off suite move
it back to Preview โ a signed-off baseline is a claim that the page is verified,
and three failures in a row make that claim false. What does not count toward
it: a skipped run, a run from someone's own Claude session, a run you asked for
with Run QA Test Now, a triaged or accepted failure, a preview route, and an
error (which means the runner could not reach an answer, not that the page is
wrong). The counter measures failures nobody has looked at, and each of those
is either someone looking or not a failure.
A route with an open defect is not demoted again for the same problem every three sweeps.
6. GA tracking is configured somewhere else now
Tracking checks are no longer a QA test you open here. They moved to the Monitor,
which visits the page every hour anyway โ asserting that a data-module exists
is a cheap job that was being done by the most expensive runner on the estate.
What replaced it is a three-level model, and reading it in the right order is the whole thing:
- The GA framework (Admin โ GA Framework) says what each module
carries. One row per
data-moduleโannouncement_bar,multi_step_buybox,membership_blockโ listing thedata-*elements that belong to it, with a description, a screenshot and a valid / unsure state so a name nobody here can decode becomes a list to take to the developers. It is a copy of the Confluence SSOT, linked from the top of the table. - A page type says which modules it has. Tick them on the page type; every route of that type inherits the list.
- A page is measured against its type.
Removing a module archives it rather than deleting it, and archived rows can be restored.
ga_check_mode decides what a mismatch does: off, log (record it, raise
nothing) or on (file a defect). It is log by default. On on, only a
changed value files a defect โ a missing attribute does not, because all
eleven page types are still seeded with all thirty-eight modules and an FAQ page
is currently asked for a buy box. One open defect per page, updated rather than
re-filed, and it carries a GA pill so you can tell it from a runner's finding.
Going live teaches GA about the page. When a Shopify-hosted route goes live, Aegis scans it once and records what it actually found, so the framework starts from the page rather than from an empty row.
ga-tracking.test.md still exists on Shopify Landing routes and still holds the
generated assertions. route_ga_config โ the old per-route override โ is empty
and unread: a route being different from its page type turned out to be a
question about the page type, not about the route.
7. The runner queue
Runners don't scan for work โ they claim it:
- A runner asks for the next route, identifying itself with a
runner_id. - The queue hands out one route and marks it
processingagainst that runner. - The runner executes and submits a report; the row moves to
completed. - The queue refills itself rather than waiting for a sweep to end.
This is what makes multiple runners safe to run at once โ two Mac Minis will never
pick up the same route. If a runner dies mid-route, that row stays processing
until the lock ages out.
Not every route is handed out every cycle. A route whose page has not changed
since it was last tested is skipped, and the skip is recorded as a run carrying
the previous verdict so the history stays continuous. qa_skip_mode is visual.
A route is force-checked after qa_force_check_days whatever the signal says.
A page waiting on a person is not handed out at all. If the last run failed and nobody has triaged it, the queue will not offer that route โ not to the schedule, not to a forced request. Running it again could only produce a second failure nobody has looked at, or a pass that overwrites the first before anyone saw it. Triage is the only thing that puts the page back in the queue.
The consequence worth knowing: a page no longer heals itself. A transient failure used to clear on the next pass with nobody looking. Now it stops dead and stays on Pages that require review by a human until somebody decides.
"Run QA Test Now" jumps the queue and says who asked. It moves the route to
the front and marks it forced. Your name travels with it, and at submission it
lands on the run as requested_by. Three things read it:
- The report names you beside the agent that ran it โ ๐ค who ran it, โ who asked.
- Slack answers you whatever it found. A scheduled pass is silent, because that is what keeps the channel readable; a run you asked for reports back either way.
- It does not count toward the 3-strike demotion. That counter exists to catch a signed-off suite going quietly stale on the schedule, and someone pressing the button is the loudest possible evidence it is not. Otherwise investigating a broken page by re-running it three times would demote it mid-investigation. A requested run that passes still clears the streak.
Runs from someone's own Claude session are excluded from that counter for the same reason, and more strongly: failing is often the point when you are reproducing something you are about to fix.
What forcing does NOT do: override the skip assessment. This guide used to say it did. It does not, and has not since polling replaced sweeping โ the poll asks the skip assessment without ever telling it the route was forced. So pressing the button on a page that looks unchanged runs nothing, reports success, and says nothing about it.
That is now a deliberate choice rather than an oversight. If the page genuinely has not changed, a run returns the answer already on record and spends the allowance to learn nothing. What is still wrong is the silence: the button should not be offered when it cannot work, which is on the todo.
Found 2026-09-26 on a page that sat pending and forced for two days, first
in the queue order, while twenty-five other pages ran past it โ every poll
recording skipped, visual_reason: unchanged.
7a. Pages that require review by a human
A second table under the queue, and the two are complements: Up next is what the runner will be given; this is what it will not.
A page is on it when nothing is testing it, and the Why column says whether that is a fault or a decision:
| Why | what it means |
|---|---|
| Untriaged failure | the last run failed and nobody has decided what that means |
| Unknown state | the suite carries a status the system does not recognise |
| Stub only | signed off on a test that asserts nothing |
| Awaiting sign-off | a draft nobody has accepted โ including anything demoted after three failures |
| QA paused | somebody stopped testing it on purpose |
| PIM โ on request | request-only by design |
Anything in the runner queue is not on it, including a test being drafted:
that page is already in Up next as CREATE TEST, waiting on a runner rather
than on a person.
Two clocks, because they are different facts: Waiting is how long the page has been in that state, On list is how long this report has known.
The list is written when a runner hands work back, and refreshed whenever the screen is opened โ so an empty list means nothing is outstanding, and a failed refresh says so in amber rather than showing you a confident blank.
8. Reports
Submitted reports render at /api/report/qa/:id with a pass/fail badge, the
route and the URL that was actually fetched, the submitting runner_id, who
requested the run, and the runner's own suspicion if it offered one. These pages
are readable without a session so the links work from Slack.
The Markdown is rendered server-side, by shared/miniMarkdown.js, which
escapes the input before it does anything else โ a runner's write-up routinely
quotes the page's own HTML. The QA report, the audit report and the findings page
all run that same renderer and share one light palette, so the three read as one
set.
The day's findings โ every open defect grouped by store, severity and owner โ
is at /api/report/findings. Coverage is stated at the top and called partial
below two thirds, each finding links the run that found it, and the severity,
store and not-a-defect filters live in the URL so a narrowed view can be sent to
whoever owns it. It is linked from the daily Slack post (behind its own toggle),
from the Forge report and from the user menu.
The security note that used to sit here is resolved. Report submission goes through the MCP router, which requires a bearer token scoped per app, and report pages escape their content before rendering. The pages are still deliberately readable without a session โ that is what makes a Slack link work โ so treat the URL as shareable within the team and nothing more.
9. Practical guidance
Writing a good baseline. Capture real current values with a date: product name, live and struck-through price, the CTA label verbatim, the option selectors, the empty-cart state, the free-shipping threshold, and cart-drawer upsells. A baseline of vague intentions catches nothing.
Hard stop at the cart. Tests verify that an item lands in the cart and the subtotal updates, then remove it and confirm the reset. They never proceed to checkout or payment.
Don't pollute analytics. Runs fire against production, so every test must use test-traffic exclusion (a test parameter or a flagged session) or target a staging store. Otherwise QA traffic shows up in GA4, Elevar and Meta CAPI as real traffic.
Click the real button. Adding to cart by posting straight to /cart/add.js
bypasses the on-page discount flow and produces false price failures. Click the
actual CTA, then read /cart.js to confirm SKU, plan, original price and final
line price. Clear the cart between batches.
Flag, don't fail, on ambiguity. Internal inconsistencies (mismatched review counts, price maths that doesn't add up, CTAs pointing at different targets) are worth surfacing to a human rather than hard-failing a run. Same for a route that legitimately uses a first-party analytics stack instead of Google tags.