Testing Shopify pages

How a page is tested, from the day someone starts building it to the day it is quietly still being checked eight months later.

Read this once before you touch anything. It is short on purpose.


The one thing to understand first

A page is tested against two different things at two different stages, and almost every misunderstanding about this system comes from mixing them up.

Stage Checked against Run by
Preview (PIM) the Range Plan you, in your own Claude
Live the page's own baseline the runner fleet, automatically

A runner never tests a live page against the Range Plan. That is deliberate and it is not going to change. The Range Plan can be edited by anyone, at any time, for any reason, with no audit log. A baseline that can be silently changed is not a baseline — it makes a failure impossible to read (did the page change, or did the Range Plan?), and it opens a way for a broken page to be made green by editing the thing it is judged against.

The Range Plan check does happen. It happens before the page goes live, over and over, while the page is in preview. That is what the preview stage is for. By the time you flick the switch you have checked against the Range Plan more times than a machine would have.

Launch day is still a human job. Aegis does not test a page onto the site and is not trying to. The team's launch process is unchanged. Aegis picks the page up from the moment it goes live and keeps checking it long after everyone has moved on — which is the part no person was ever going to do.


1. PIM — while the page is being built

A page in PIM has a PIM URL — that is the field's name in Aegis, and filling it in is the whole distinction. A route with a PIM URL is a PIM page; a route without one is live.

Setting it does more than label the page. Admin's links, the QA screens, compliance scanning and the monitor all start using that URL instead of the live one. Clearing it is what sends the page live.

PIM pages are deliberately kept out of the runner fleet. They are under active development, they change by the hour, and a fleet filing failures against a page somebody is still building is noise, not QA. The monitor still watches them — at the PIM URL — so a preview that stops answering is still noticed.

Their data comes from the Range Planner: https://pf-range-plan.purdyandfigg.app/

Creating a PIM test

Pages is the Aegis section that lists every route the platform knows about — live and PIM together — and a route is simply the platform's record of one page: its path, its store, its type, where it lives, and who should be told when something happens to it. Everything else in Aegis hangs off that record. A page with no route is a page Aegis has never heard of.

So: add the route in Pages, and fill in the PIM URL. That is it.

"Why does a preview page need to be in Aegis at all?"

It is the first question everybody asks, and it is a fair one.

Because it is the same page. The route you add today is the record that will still be there on launch day, after launch, and eight months later. You are not registering a preview — you are registering the page, early, in the state it happens to be in. On launch day nothing is recreated: the PIM URL comes off and the same record carries on as a live page, keeping its history, its defects and its identity.

It is no stranger than a preview page on Shopify. You would not want a preview page sitting on the live site either, which is exactly what the PIM URL field is for — it is how one record can be "not live yet" without pretending to be something else.

You can skip it, and here is what you lose. Nothing stops you pointing Claude at a URL and asking for an ad-hoc check; that works, and sometimes it is the right thing. But Aegis does not know what that page is, so there is no audit log, no history, no defect record, and nothing carries forward to launch day. You get an answer and nothing else. Adding the route is what turns a one-off answer into a record that accrues.

Running a PIM test

PIM tests run in your own Claude, through the Aegis MCP server — not on the runner fleet. If you have not set that up yet, see Using the Aegis MCP Server in Claude.

Ask for it in plain language. Something like:

test https://purdyandfigg.com/products/<the-preview-url>
against the PIM / range planner

Reading the result

The result comes back in your own session, in front of you: a pass, or a list of what does not match the Range Plan — prices, SKUs, copy, variants. Fix the page, run it again. Repeat until it is clean.

Do not ask for the results in Slack. A page in preview changes by the hour and you will run it many times before it is right; every one of those in a channel buries the thing Slack is actually for, which is the live estate telling you something has gone wrong. While a page is in preview it belongs to you, and the only person who needs the output is the one reading it.

This is the loop that earns the page its live switch.


2. Launch day

The human process does not change. Everyone's eyes are on the page, the same as they have always been.

What changes is what happens the moment it goes live in Aegis:

  1. Aegis flips the route to live — the PIM URL comes off.
  2. A test is created automatically. It reads the live URL and stores what it finds: prices, copy, structure, a screenshot.
  3. That captured state becomes the page's baseline, and from here the page belongs to the machines.

The create step reads the live page and not the Range Plan for the reason above: the live page is what customers see, and a crack team of devs has just spent weeks making it match the plan. The plan's job is finished. The page's own state is now the thing worth defending.

From here, start at the daily report

Once a page is live it stops being something you go and look at, and becomes one line among many. The daily report is the jump point — open that first, every day, and let it send you where it needs to.

It now counts the whole estate in four tabs, not just what is wrong — which matters more than it sounds, because until recently a quiet morning and an untested one both rendered as an empty page:

Tab What is in it
PIM (preview) Pages with a preview URL. They run only when asked, so they are never overdue. It opens here — this is what somebody is working on.
Live (no issues) A test is running, its last run was answered, no defect is open.
Live (failed QA) QA cannot vouch for these: an untriaged failure, or nothing testing them at all.
Live (with defects) A known fault on the page. These do not stop QA running, so a page can be in this tab and the one before it.

Every count is out of the same total, and every row links through to the page, its test, its defects and its live URL. It is also what gets posted to Slack each day.

The rest of this guide is what the daily report will send you into.


3. The monitor — the cheap check

The monitor runs before the runner fleet in the flow, and it is much cheaper.

It checks a page every 2 minutes, in a circle: whoever was checked longest ago goes next, so every page is seen once a lap. It asks two questions:

  • Does the page answer? Uptime. A page that 404s or times out repeatedly is auto-disabled after a run of failures.
  • Does it still look the way it did last time? A screenshot compared against the stored baseline image. A difference above the drift threshold raises a visual audit for a person to accept or reject.

A visual audit gets its own page — an audit report, at /api/report/<audit id> — showing the stored baseline beside the new capture so you can see what moved. If the change is correct, accepting it there installs the new capture as the baseline from then on.

The accept button only appears on the latest audit for a page. Accepting an older one would install an old photograph as the thing every future check is measured against, which is a quiet way to make a page permanently wrong.

The monitor is a different cadence and a different unit from QA, which is why its numbers are kept apart rather than averaged into the QA figures.

It proposes, it never activates. The monitor has no path to change a test, a status, or a baseline on its own.


4. The runner fleet — the expensive check

Runners are Claude instances that pick up one live page at a time and actually test it. You can watch the whole thing at https://aegis.purdyandfigg.dev/#qarunners.

The screen shows, left to right: how many runners exist, what is running now, runs in the last 24 hours, the pass rate, and how many pages are queued. Below that, the last five pages finished, then a card per runner with its own totals and when it last reported.

The loop, end to end:

  1. A test is created from the live page and stored.
  2. The test goes into the pool.
  3. A runner takes the next page from the pool when it is free.
  4. It runs the test. No meaningful difference from the baseline → it submits a pass.
  5. Any difference — a price, a copy change, a missing block — is either a failure (the page has moved and the test must be looked at again) or a defect (something is wrong with the page itself). See below.
  6. A human reviews the failure and does one of two things: accept that the page has legitimately changed and let a new test be drafted, or say it is fine and to carry on.
  7. Back to step 2 — until the page itself changes season. See below.

A page that has not changed does not need an expensive run, so an unchanged page is skipped cheaply. To stop that hiding a page that nothing has looked at properly in months, a forced check runs the full test regardless after a set number of days — 7 by default, configurable in Settings.

The playbook

Runners do not judge a page on the baseline alone. Playbook rules carry the standing knowledge that is true across the whole estate — things like "if the page says out of stock, it is out of stock, do not fail the route for it" — so a runner can tell a real fault from a thing that merely looks unusual. A playbook rule beats a baseline assertion: where the two disagree, the rule wins and the test does not need redrafting.

How a rule gets in. Two ways, and neither lets a machine change what the fleet does on its own.

1. Proposed from a run. A runner that learns something fleet-wide proposes it through the Aegis MCP. It does not take effect. It lands on the Playbook Rules screen in Aegis and waits for a superadmin, who can edit it before approving — proposals carry a full revision history, and several of the live rules were amended several times before they went in. Anyone can propose this way, not just the fleet: it is the same tool from your own Claude, and a good share of the current playbook came from people rather than runners.

A proposal is either an append (a new rule) or a replace (correcting text already in the playbook). A replace has to name the exact text it stands in for, and is checked both when proposed and again at approval, so the playbook cannot end up carrying two rules that contradict each other. A runner can also withdraw its own proposal — usually on realising the lesson belonged to one route rather than the estate.

2. Edited directly. The playbook is a QA template — aegis-qa-runner-hourly.md — and a superadmin can edit it in the template editor like any other, with versions. No proposal, no queue.

If it is true of one page, it is not a playbook rule. It belongs in that route's own content.test.md, where it is versioned with the page it describes. This is the most common reason a proposal gets withdrawn or rejected.

There is always another launch day

The loop above is not a straight line, and this is the part that is easy to miss. /pages/completestarterkit in autumn is not the same page as /pages/completestarterkit in summer. Same URL, same route, different page.

When the season turns, the page goes back to PIM: set the PIM URL again and it leaves the runner fleet, and you are back at section 1 — building against the Range Plan, testing in your own Claude. Then it relaunches, the PIM URL comes off, and it is a live page again.

Two things happen on the way back out that are worth knowing:

  • It goes to the front of the queue for an immediate live run. The page was baselined against the preview, which is the state it was meant to publish in, so the first run against the live page is the comparison that catches publish itself changing something. The interesting window is right now, not at the next rotation.
  • The old test will fail, and it should. The test signed off for summer does not describe the autumn page. That is a legitimate change: it goes to triage, you accept it, a new test is drafted from the page as it now is, and you sign that off. Section 7 is how.

So the lifecycle is a circle, not a line:

PIM  →  live  →  tested forever  →  back to PIM  →  live again  →  …

Which is also the shortest way to say why the two baselines are what they are:

The Range Plan is the source of truth before publish. The live page is its own truth after.


5. Triage

Triage is where a machine hands a decision to a person. Nothing leaves triage on its own.

A test that has just been drafted goes to preview and waits. A run that failed waits. Both sit until somebody answers them.

The full path for a new test:

create test  →  draft  →  preview  →  a human signs it off
                                   ↓
                             the real test
                                   ↓
                       a human signs THAT off
                                   ↓
                         into the runner fleet
                                   ↓
                     tested automatically, forever
                                   ↓
                        fails  →  back to triage

Two sign-offs, both by a person. A drafted test is a proposal; it does not test anything until somebody has read it and agreed.

And nothing redrafts itself. It has been asked — if forty routes fail the same way after one template edit, should the system stop asking and just redraft? Decided no, 2026-09-26. Accepting a failure rebuilds the baseline from the page as it is today, so a system that did it automatically would adopt whatever the page currently says — including a regression the runner had misdiagnosed as a stale test. A page that genuinely broke and keeps failing is exactly the shape that would reach any threshold you set. The forced check already stops a stale test going unexamined for ever, so the only thing automation would buy is removing the person, which is the part worth keeping.


6. Failures and defects are not the same thing

This distinction matters more than any other in the system.

Failure Defect
What it means the page no longer matches its test the page itself is wrong
Whose problem QA's — the test may be out of date whoever owns the page
Stops QA? Yes. The page waits for a person No. The page keeps being tested
Example the price changed, deliberately "member" is rendered as "emmeberd"

A page can carry an open defect and still be tested cleanly every pass. That is not a bug — the defect is logged, somebody owns it, and meanwhile the page is still watched for anything new going wrong.

Defects are filed from live runs. In preview, a mismatch against the Range Plan is just something to fix before launch; it does not need a defect record.


7. Working a failure

A failure in triage is a question, not a verdict. The runner is telling you the page no longer matches its test; it is not telling you which of the two is wrong.

The way this is actually done:

  1. It lands in triage and stops there. Nothing moves until a person answers.

  2. Take it to your own Claude. Point it at the page and at the run that failed — the runner has already written up what it saw — and work out what actually happened. This is the part that needs a person: the runner reports, you decide what it means.

  3. Answer it. The screen gives four, and they are not four ways of saying the same thing:

    What it means
    Known defect The page is wrong. Files a defect, which becomes the thing tracked. QA carries on.
    Not a defect The test is wrong — a false positive. The most useful of the three, because it is the only one engineering can act on.
    Investigating Somebody is looking.
    Accept Sits across all three and answers a different question: what now? The run stays red in the record — overwriting it would make "this passed" and "somebody accepted this" identical forever — but everything downstream reads it as a pass.

    Flattening those into one make-the-red-go-away button is how a defect list becomes a make-the-red-go-away button, which is worse than no button.

    A runner that found the failure has often already worked out which of these it is, with the page still open in front of it. It records that as a proposal and nothing more. It cannot dismiss its own failure — a runner that can do that can make any test pass.

  4. If it is the page that is wrong, file it as a defect. This is the move that gets the page back to work: the run is marked triaged with your classification and note, the defect (or several — a run can fail on more than one thing) is linked to it, and the page stops waiting on a person. The problem is not forgotten, it has just moved somewhere that does not block QA. Only a failed run can be triaged; there is nothing for a pass to explain.

    If a runner had already proposed a triage and you decide for yourself, its proposal is marked superseded rather than rejected. "Nobody answered it" and "somebody disagreed with it" would otherwise read identically, and only one of those says the runner was wrong.

  5. If you learned something, write it down — in the same session, while you still have it in front of you:

    • True of the whole estate? Propose a playbook rule. Every runner on every page picks it up once a superadmin approves it.
    • True of this page only? It belongs in that route's own content.test.md, versioned with the page it describes.

Step 5 is the one people skip, and it is the one that compounds. A rule proposed after a triage is the reason the next fifty runs do not raise the same question.

Leave it, and the page falls out of the schedule

Three consecutive failures on a signed-off suite demote it back to preview. A signed-off test is a claim that the page is verified, and three failures in a row make that claim false — so the claim is withdrawn rather than left standing.

The counter measures failures nobody has looked at. A triaged failure does not count. Neither does an accepted one, a skipped run, a run from somebody's own Claude session, a run you asked for with Run QA Now, a preview route, or an error — which means the runner could not reach an answer, not that the page is wrong.

So the way to stop a page being demoted is to answer it. Ignoring a failure three times takes the page off the schedule, which is the opposite of what ignoring it was meant to achieve.


8. Defects: what each action does

Open a page's Defects tab and every one has an Actions menu. They are not four ways of saying the same thing.

Action What it does
Investigate Puts your name on it. It stays open — this is "I am on it", so two people do not start the same thing. Reversible.
Mark fixed Closes it as fixed. It stays on the Closed list rather than vanishing, because a recurrence needs to read as a recurrence and not as a fresh discovery.
Won't fix Closes it as a decision: we are living with this.

Severity is one of critical, high, medium, low, and only open defects count towards a page needing attention. A won't-fix is a decision already taken and a fixed one is history; neither should put a marker on a route.

There was a fourth action, Hide, and it is gone. It refused to hide an open defect, which meant it could only ever act on a closed one — and a closed defect taken off the reports is what Archived already is, with a worse name. Two controls for one outcome is how somebody picks the one that does not do what they meant.

Open, Closed, Archived

Defects sit in one of three lists, and failures now use the same three — the two screens had different vocabularies for the same idea, which made "where did it go?" a question you had to answer twice.

What it holds
Open Nobody has decided. This is the only list that counts against a page.
Closed Decided — fixed, or accepted. Still visible, still counted if it recurs.
Archived Out of the way. Nothing is deleted on a schedule; things arrive here and stay.

Nothing is ever removed automatically. Items move between these lists on a timer, and that is all the timer does. The only way anything leaves is the trash icon on an archived item, which deletes it permanently and asks first.

The windows are settings, not constants

Settings → QA carries three, each a number of days. The defaults are 7, and none of them deletes anything:

Setting What it does
Open defect closes after An open defect nobody picked up is closed as accepted. It still counts if a runner sees it again.
Closed defect archives after Moves off Closed and onto Archived. A defect closed in March is the argument when the fault returns in June.
Answered failure archives after A failure somebody has decided about moves to Archived. An unanswered failure never moves on its own, whatever the number says.

That last row is the important one. An earlier version of this pruned old runs with no check on whether anybody had answered them, so untriaged failures were being deleted along with their verdicts and report bodies. Nothing does that now.

How they stack

A defect found again does not create a second row. The sighting is added to the one that exists: occurrences goes up and last seen moves. What happens next depends on the state it was in:

  • Open → counted. Nothing has changed.
  • Won't fix → counted, and it stays closed. A decision that unmakes itself every time the runner sees the problem again is not a decision. The evidence still accrues underneath, so the record reads "we chose to live with this, and it is still happening, 40 times, last seen today" — which is a better argument for fixing it than an open row nobody looked at.
  • Fixed → reopened, loudly. Somebody said this was fixed and a runner is still finding it, so either the fix did not work or it has regressed. That is the one case where finding it again is genuinely new information.

The automatic close

Any defect left open longer than the window is closed automatically as won't fix, with a note saying how long it sat. The window is Open defect closes after in Settings → QA, and it defaults to 7 days.

This is an expiry, not a judgement. Nobody looked at it and decided it was acceptable — it simply aged out. A defect open that long was not going to be picked up by being left open longer, and closing it says the true thing: we are living with this. Every later sighting still counts against it.

It cannot become a treadmill of closing and refiling, because a won't-fix is never reopened by a re-report. The auto-closed ones are also kept out of the "not a problem" section of the findings digest, so the handful a person actually looked at and dismissed are not buried under hundreds that closed themselves.


9. Run history

The third tab on a page is its log — every run, and every change to the test alongside them: passes, failures, sign-offs, demotions.

The most recent run is shown; the rest collapse behind Previous runs and changes. View report opens the QA run report — at /api/report/qa/<run id> — which is what the runner actually wrote for that run: what it checked, what it found, and why it called it a pass or a failure.

That report is the evidence behind a failure, and it is the thing to read first when you pick one up in triage. It is a public page, so it can be sent to whoever owns the fix without giving them an Aegis login.

The summary line above it is worth reading properly. "1 failure · 0 checks skipped" — a skipped check is a run where the page had not changed since it was last tested, so the expensive test was not re-run. A skip is not a pass and is counted separately, which is exactly why it is on that line rather than folded into the totals. A page showing nothing but skips for weeks is a page nothing has genuinely looked at, and the forced check is what stops that running forever.


10. The test screen

The Editor tab is where the test itself lives, and the first thing to know is that a test is a plain text file — content.test.md, ordinary Markdown. Not code, not a config format, nothing that needs a developer to read. It opens with what it is testing and how the baseline was captured:

**Target URL:** https://us.purdyandfigg.com/pages/completestarterkit?aegis_qa=1
**Page type:** LP (Shopify-rendered lander, US storefront)
**Baseline captured:** 2026-09-22, viewport 1280x813 (desktop), theme `191433212218` role `main`.
**Drafted by:** agatha-qa-runner

and then says, in numbered prose, what to check and what counts as a pass. If you can read it you can correct it, which is the point — when a triage turns up something the test got wrong, the fix is editing a sentence.

Version history keeps every version. One is the default: the one runners actually use. A new version does not become the default until a person signs it off, so editing is never the same thing as deploying.

Save, and the Manual actions menu

The foot of the editor has Save and one menu, with a line beside it saying those actions are outside the ordinary triage flow.

What it does
Save Saves your edit as a new version. A button, not a menu item — a save behind a menu is a save people lose.
⏸ Move to preview Takes the page out of the polling queue. Greyed out, with the reason on hover, when the page is already in preview.
📝 Ask runner to create test A runner reads the live page and writes a new version from scratch. Nothing you have open is edited — the new version arrives in preview for you to read and sign off, and the current one keeps being tested until you do.
▶️ Ask runner to run test Puts this page at the front of the queue; the next free runner tests it against the current script. Changes neither the script nor the test's state.

The three are behind a menu because they are not the ordinary way to work: the normal route through this screen is a failure, a decision, a redraft and a sign-off, and none of those touch them. The note beside the menu matters for the same reason — a collapsed menu with no label reads as "the other actions", which is how somebody asks a runner to redraft a test that triage is already redrafting.

Move to preview first, before you edit anything. The editor opens on a signed off test, and a signed off test is by definition still on the schedule — so a scheduled run can land while you are halfway through rewriting it and judge the page against a half-finished test. Moving to preview is what makes editing safe.

That gives the working loop for a test, and it is worth learning as one thing:

park it → edit it → ask the runner to run it → sign it off.

Note that nothing here runs in your browser. "Ask runner to run test" queues the work; a runner picks it up. The button used to say Run now and promised an immediacy it could never deliver.

Pausing QA on a page

Separately from any of the above, you can pause QA on a page — the pause icon on the runners screen, on both the up-next list and the human-review list.

A paused page is not tested until somebody resumes it. It is not the same as either of its neighbours:

  • Move to preview parks the test, because you are rewriting it.
  • Pause stops QA on the page, for a reason that usually has nothing to do with the test — a known migration, a sale weekend, a supplier changing copy hourly.
  • PIM and disabled are different again: a paused page is still live and still monitored. Only QA stops.

Pausing asks for a reason, and the page then sits in the review list reading QA paused by — . That is deliberate: a page nothing is testing should be visible and attributable, so a pause cannot quietly become permanent. Resume puts it straight back on the schedule.


11. Where to look

What is running, what is next, what needs a person The runners report
Today's outstanding work, and every page carrying a defect The daily report
Every route's health The daily status report
Every finding, searchable The findings report

The daily report is the one that gets posted to Slack each day. It is a page rather than a message, so it stays correct after somebody closes a defect.

What lands in Slack, and where

Settings → Slack has two halves and they answer different questions.

What Aegis Sends is the list of every kind of update and a switch for each. Turning one off here silences it everywhere, for every channel. Each row also says which channels carry it, with the scope beside the channel name — because "this goes to #shopify-ai-testing" is rarely true on its own; what is true is "it goes there for Shopify pages".

Routing Rules decides where each one goes. A rule is a channel, a scope (page type, origin, store — any of them optional) and the update types it carries. Every matching rule fires, so an update can reach more than one channel; anything no rule matches goes to the default channel, which cannot be deleted and cannot be left unset, because it is the floor that stops an update being dropped in silence.

Two things worth knowing when a message turns up somewhere unexpected:

  • A page can carry a channel override in its own settings, which skips the rules entirely rather than narrowing them.
  • The route's Slack panel has a Where does this go? control. Pick an update type and it reports the channels, the rules that matched, and how the page was classified — answered by the same code that sends, so it cannot disagree with what actually happens. Use it before reading channels.

Those four are the estate. Two more cover a single thing, and both are public pages you can send to somebody:

What a runner found on one run /api/report/qa/<run id> — the QA run report
What the monitor saw change on one page /api/report/<audit id> — the audit report

Known gaps, stated honestly

  • A baseline is never re-validated after it is captured. Whatever the page looked like when it went live is what it is defended against from then on. The forced check above cures a page being skipped for too long; it does not re-examine the baseline itself. This is covered today by the human launch process, which is why it has not been automated.
  • The monitor is not PIM-aware yet. A cheap PIM check on preview pages is the next planned upgrade.