QA Manager Handbook
Start with Testing Shopify pages. That is the short version of the whole flow β preview against the Range Plan, live against its own baseline, and what happens in between. This page is the longer reference behind it, and parts of it predate the current flow.
This is the usage guide for whoever owns QA. It assumes no prior knowledge of Aegis and walks through the job in the order you will actually meet it.
It deliberately does not repeat the architecture. Where something is explained better elsewhere, this page links to it:
| For | Read |
|---|---|
| Why the runner is an AI agent, guardrails in depth | QA Runner Operations |
| Writing and versioning test specs | QA User Guide |
| Connecting Claude to Aegis, full tool list | Claude MCP Setup |
| Uptime and visual monitoring internals | Monitor User Guide |
| Ad-compliance scanning | Compliance Scanning Guide |
The four places you will live:
| System | URL | What it is for |
|---|---|---|
| Admin | https://aegis.purdyandfigg.dev |
Routes, QA, sign-off, audit reports, settings |
| The day's findings | β¦/api/report/findings |
Every open defect, grouped and shareable |
| Route Status | https://aegis-route-status.purdyandfigg.app |
At-a-glance coverage: is QA actually running? |
| Docs | https://aegisfaq.purdyandfigg.dev |
This site |
QA Bot is no longer a place. It was a separate app at
qabot.purdyandfigg.dev; it is now two entries on the route's Options menu in Admin β π§ͺ QA Test to write the test, π QA Status to read what it found. The old domain still answers and redirects, so existing links and bookmarks land on the right screen.
0. The one rule
The agent may propose. Only a human may activate.
Everything below is a variation on this. An AI runner can write a test, run it, snapshot a version and pause a route. It cannot make a version the live baseline and it cannot un-pause. Those are your decisions.
The reason is not caution for its own sake. Pausing is safe, so it is delegated. Resuming and activating are judgement calls, so they are not β including when the agent is confident the change is benign.
1. What a QA runner is, and how to become one
What it is
There is no test-runner service. Nothing is deployed that "runs tests". The runner is an AI agent β Claude, holding a skill prompt, driving a real Chrome browser, talking to Aegis over MCP.
Aegis (Cloudflare) Local runner (a Mac Mini)
ββββββββββββββββββββββ ββββββββββββββββββββββββββββ
β qabot_test_queue βββ claim βββββ Claude + Aegis MCP β
β qabot_test_versionsβββ spec βββββΊβ + Claude in Chrome β
β qabot_test_runs βββ report ββββ (real browser, real DOM) β
ββββββββββββββββββββββ ββββββββββββββββββββββββββββ
Three consequences worth understanding, because they explain most of the design:
- The browser must be real. Most routes proxy a client-rendered Lovable origin. A plain HTTP fetch returns an empty shell. Claude in Chrome executes the page, so the runner sees what a customer sees.
- The test is prose. A Markdown spec can be read, run and amended by anyone, including people who will never open a code editor. That is what makes "anyone can test any page" real rather than aspirational.
- It runs on a Mac Mini, not in the cloud. Cloudflare Workers cannot hold a logged-in browser session. The control plane is at the edge; the hands are on a machine in the office. That machine is Agatha.
Becoming a runner (about 10 minutes)
Step 1 β Connect Claude to Aegis. Aegis is already registered in the Claude organisation settings, so there is no config file to edit:
- Open Claude
- Go to Customise
- Find Aegis Admin MCP Edge Router
- Enable it
The endpoint behind it is POST https://aegis.purdyandfigg.dev/api/mcp. One
endpoint gives you everything β the Admin worker hosts its own tools and proxies
the rest to the other workers internally.
If a tool exists but Claude cannot see it, Claude caches the tool list at startup. Fully quit (Cmd + Q β not just closing the window) and reopen.
Step 2 β Enable Claude in Chrome. The runner needs a real browser. Without
it, the agent can read specs but cannot execute a page, and every run should
return error rather than a guessed pass.
Step 3 β Load the runner skill. In Claude, invoke the MCP tool:
setup_qa_runner_skill
This returns the current runner instructions. It is not hardcoded β it reads
aegis-qa-runner-hourly.md from the qabot_test_templates table. See
section 5 for why that matters.
Step 4 β Run. Ask Claude to begin. It will claim work from the queue itself.
Agatha and Tailscale access
Agatha is the Mac Mini that runs QA on a schedule. The agentic employees β Agatha (QA) and Alfred (tech lead) β are accessed over the company Tailscale account. Ask the outgoing owner to add you to the tailnet; once you are on it you connect to Agatha's desktop and use the Claude session running there exactly as you would locally.
Agatha is a scheduled session, not a service. That is the single most important operational fact about it β see section 8.
2. Adding a route and getting it tested
Add the route
In Admin β Routes β + Create New Route, the fields that matter for QA:
| Field | Why QA cares |
|---|---|
| Path | What the test targets, e.g. /pages/completestarterkit |
| Origin | Where it proxies to β Lovable, Shopify, Cloudflare or external |
| Store | uk or us. Routes are matched per store; the same path can exist in both |
| Page type | Drives Slack routing rules and reporting |
| Preview URL | If set, used instead of the computed live URL everywhere β including tests |
A new route is inactive. It cannot be activated until it has passed sign-off (section 4).
Run a test immediately
You do not have to wait for the next pass β which is every six hours, not hourly. On any route row:
Options β π§ͺ Run QA Test Now
This pushes the route to the front of the runner queue and records that you asked. The next free runner picks it up, and results land in Options β QA Status β Runs.
Two things it will not do, both worth knowing before you rely on it:
- It does not override the skip check. If the page has not changed since QA last looked, the runner passes it over β the request reports success and nothing runs. That is deliberate: an unchanged page returns the answer already on record. What is wrong is that the button does not say so, which is on the todo.
- It does not get past a failure nobody has triaged. A page whose last run failed is not handed out at all until somebody decides what the failure means. Triage it first, then run it.
Because you asked for it, three things follow: Slack tells you the result either way (pass or fail β scheduled runs stay failure-only), the report names you beside the agent that ran it, and the run does not count toward the 3-strike demotion, so you can re-run a page you are fixing without demoting it mid-investigation.
There is a separate Run visual check, which is not the same thing β it captures a fresh screenshot through the Monitor and takes about 15 seconds. Use that when you want to refresh a visual baseline, not when you want a QA run.
The queue makes this safe
A runner does not scan for work, it claims it: ask for the next route β the row
is marked processing against that runner_id β execute β submit report β row
goes completed. Two Mac Minis will never test the same route at the same time.
The queue refills itself rather than waiting for a cycle to end, and one poll
hands out both tests and diagnostics.
A route whose page has not changed is skipped, and the skip is recorded as a
run carrying the previous verdict so the history stays continuous. A backstop
forces a real check after qa_force_check_days whatever the signal says.
If the queue looks jammed, the MCP tool reset_qa_test_queue puts every route
back to pending.
Where results appear
- Options β QA Status β run history, the full Markdown report per run, and the route's open defects
- The findings report β every open defect across the estate, grouped by store, severity and owner, with coverage stated at the top
- Admin β route audit reports and the audit log
- Route Status β coverage across every route, flagging anything with no recent audit or QA run
- Slack β failures (section 6)
3. What each check actually does
Two independent systems produce QA evidence. They are often confused.
QA checks β the agent, in a real browser
Driven by the Markdown spec for the route. Typically:
- Content β the page renders and the expected copy, prices and labels are present
- GA / pixel β tracking fires correctly (GA4, Meta, TikTok as configured)
- Cart β clicking the real CTA adds the right SKU, plan and price
The cart check has rules that exist because they were got wrong before:
- Click the real CTA. Posting directly to
/cart/add.jsskips the on-page discount flow and manufactures false price failures. - Read
/cart.jsto verify SKU, selling plan, original price and final line price β do not trust the on-page display alone. - Hard stop at cart verification. Never proceed to checkout or payment.
- Clear the cart between batches, so one run cannot contaminate the next.
- Exclude test traffic (test param, flagged session or staging store) so QA runs do not pollute GA4, Elevar and Meta CAPI.
Shopify preview checks β the Range Plan price cross-check
A preview page is usually unpublished, so there is no live URL to compare it against and no other system holding the prices it is supposed to show. Normal content assertions can't help: there is no previous version to diff. The Shopify preview brief closes that gap: on a preview route it replaces the generic brief, so the test that gets written checks the page against the system that actually owns the numbers rather than scanning copy and CTAs.
The route still has a single brief called test-brief.md β which template filled
it is decided by the route, and the banner above the editor says which.
Since 2026-09-09 there is a template per page type, named after the type:
brief_pdp.md, brief_landing.md, and so on. They were cloned from the original
test-brief.md, so they all started identical and diverge as people edit them β
a product page and a collections page need different drafting instructions.
The Shopify-preview brief still wins over the page type's one, because it is chosen by the route's state rather than its type: a route in preview has no live page to compare against, and the Range Plan owns the numbers until it is published. So a Shopify route uses both over its life β the preview brief while it is in preview, its page type's brief once it goes live.
It runs only when both are true:
- The route has a Preview URL set (Routes β edit route β Preview URL). This
is what puts the route into preview state β it shows an amber
PREVIEWbadge in the routes list. - The route is a Shopify page β
shopify_urlset, origin startingshopify://, or a path beginning/pages/.
Page type is deliberately not part of that gate. A preview page keeps its real
page type (Landing, Collectionsβ¦); preview is state, carried by the Preview
URL, not a type of its own. If either condition is false the test reports
skipped, not failed β it is seeded onto every route but only applies to some.
Once a Preview URL is set, Aegis serves it in place of the live URL everywhere
β Admin, QA, compliance scanning and the Monitor. So the URL the runner is handed
already is the preview URL; the agent is told not to reconstruct a live
purdyandfigg.com address.
Where the prices come from
The source of truth is the Range Plan FY27/28 Forge app:
| App | pf-range-plan.purdyandfigg.app |
| Data | GET https://pf-range-plan.purdyandfigg.app/api/data |
| Shape | JSON array of product cards. One product can appear as several cards β one per launch month β sharing a pid |
The endpoint sits behind Forge auth. If it returns 401/403, or HTML instead of
JSON, the agent must report error with "Range Plan not reachable β check runner
auth" rather than failing the price assertions. It is explicitly forbidden from
guessing prices from memory β a confident wrong price is worse than a failed run.
The fields it reads: rrp, ncSubPrice, ongoingSubPrice, membersOtpPrice,
nonMembersOtpPrice, subPriceMembers, ongoingSubPriceMembers, joined on pid
and sku_uk. Older cards use legacy names (subPrice, otpPriceβ¦); the API
migrates them on read, so the canonical names are what you should see.
How a page is matched to a product
In order, because the reliability drops at each step:
sku_ukagainst any SKU on the page or indata-skuattributes β the reliable joinproductNameagainst the page's product title, case-insensitive- If several cards match, prefer the one whose
season/monthmatches the page's stated drop; otherwise the latestlaunchDate - If nothing matches β report
failednaming the title it looked for and the closest candidates. It must not invent a match
The matched pid and sku_uk are reported in the run so a human can check the
join was right. Worth actually reading: a test that passes against the wrong
product is the failure mode this design is most exposed to.
Two exclusions that exist because they caused false failures
- Prices in buy-box bullet points are ignored. "FREE Bottle for Life (worth Β£20)" is a marketing claim about an included extra, not a selling price, and has no Range Plan counterpart. Skipped ones are listed in the report under "not asserted" so they stay visible without affecting the result.
- A bare number is not automatically a malformed price. Multi-market themes
render the currency symbol in a separate
.market-place-currencyelement, so19.95beside<span class="market-place-currency">Β£</span>is a correctΒ£19.95. Missing currency is only reported when neither the amount nor that element supplies a symbol β and the report has to say which was checked.
Hard stops
Reports failed immediately, skipping the remaining assertions, if the preview URL
doesn't load or shows a Shopify password page, or if any price renders as Β£0,
Β£NaN, undefined or an empty placeholder.
The lifecycle gotcha worth knowing
Test templates are copied onto a route when the route is created. Seeding or
editing a template therefore reaches new routes only β every existing route keeps
the file list it was born with. That is why changes to this test shipped as
migrations (0026 backfilled it onto existing suites, 0027 and 0028 pushed
later rewrites). All 51 suites carry it today, and 5 routes are currently in
preview state.
So if you edit the template and an existing route's behaviour doesn't change, that is expected, not a bug β the route is running its own copy. Edit the route's copy in Options β QA Test, or backfill.
Monitor checks β automated, on a cron
Configured per route in monitor_config, across three tiers:
| Check | What it catches |
|---|---|
active |
Uptime β the route responds |
pixel |
Expected tracking pixels are present |
large_images |
Oversized images hurting page weight |
visual |
Visual drift against the stored baseline screenshot |
checkout |
Checkout link resolves |
cap |
Passmark ad-copy compliance scan |
Defaults differ by tier, which is deliberate β the cheap checks run often, the expensive ones run once a day:
| Tier | Runs | Default checks |
|---|---|---|
| admin | On demand from the dashboard | active, pixel |
| hourly | 15 * * * * |
active, pixel, large_images |
| daily | Daily sweep | active, pixel, large_images, visual |
You can override any of these per route on the route form.
4. The draft and sign-off flow
There are two separate sign-off gates. Keeping them apart avoids most of the confusion new starters hit.
Gate 1 β the test spec (Options β QA Test)
Every test has a status:
| Status | Meaning |
|---|---|
create_test |
Nothing written yet. A route is born here, so an empty suite reads as work outstanding rather than as a test that passes because it asserts nothing |
preview |
A draft. Written by an agent or a person, not yet trusted. A suite demoted by three consecutive failures comes back here |
signed_off |
A human approved it. This is the live baseline |
(paused is gone. It did nothing β the runner only ever ran signed-off routes,
so pausing changed a badge and stopped nothing.)
Specs are versioned. Each version is a row; exactly one carries the
is_default flag, and that is the live baseline. An agent may create a new
version and save it as a draft. Only a human may make a version the default.
When the runner drafts a test for a route that has none, it lands as preview.
It does not become the baseline by writing itself.
Gate 2 β the route (Admin)
Before a route can be activated, all four roles must sign off:
| Role | Field |
|---|---|
| Page Author | sign_off_author |
| Head of Growth | sign_off_growth |
| Head of Tech | sign_off_tech |
| AI Evaluator (Passmark) | sign_off_ai |
Attempting to activate without all four returns:
Cannot activate: Page Author, Head of Growth, Head of Tech, and AI Evaluator must all sign off completely.
The AI sign-off can be overridden by a human, but the override requires a written reason, and it is recorded. That is intentional: an override is a decision someone owns, not a checkbox.
This gate is what makes QA consequential rather than ceremonial β a route cannot take paid budget until all four approve.
Drift vs breakage β the one dangerous mistake
When a test fails, it has two opposite causes, and telling them apart is the core QA judgement call.
| Classification | What you see | Correct action |
|---|---|---|
| Drift | Content, price or label changed; the flow still works end to end | The test is stale. Snapshot, propose a refreshed baseline as a draft, tag QUARANTINE-DRIFT, a human approves the diff |
| Breakage | The flow itself fails β add-to-cart throws, subtotal stuck, JS error | The site is wrong. Do not touch the baseline. Tag QUARANTINE-INCIDENT and raise an incident |
| Unsure | Ambiguous | Tag QUARANTINE-ESCALATE. Never auto-refresh on uncertainty |
Never wire "test failed β refresh the baseline" together. That rewrites the baseline to match a broken page, turns the test green and hides a real bug. It is the classic self-healing-test failure mode, and it is why refreshing a baseline is human-only.
The quarantine tag tells you which conversation you are about to have.
Write the judgement down β triage
That judgement used to live in your head. The run pruned after 72 hours and three failures demoted the suite regardless of what the failure turned out to be.
On a failed run, in Options β QA Status β Runs:
| Use it when | |
|---|---|
| Known defect | Breakage. Files a route defect with a severity and an owner (content / growth / dev) β the defect becomes the thing being tracked, not the run |
| Not a defect | The test is wrong. The most useful of the three, because it is the only one engineering can act on |
| Investigating | You are looking |
| Accept | A different question β what now? It sits across all three: the run stays red in the record, but everything downstream reads it as a pass |
A triaged failure stops counting toward the 3-strike demotion, because the demotion exists to catch failures nobody has looked at.
The runner usually tells you what it thinks. It had the page open and had already done the work, so it records a proposal β suspect, confidence, a note β on the report and in the alert. It is marked as a proposal in the text itself, because a confident sentence in an alert reads as a conclusion. It can never decide: a runner that can dismiss its own failure can make any test pass.
Defects, not runs, are the backlog. The Defects tab holds the route's
open ones; /api/report/findings holds the whole estate's, grouped by store,
severity and owner, with coverage stated at the top and filters that live in the
URL so you can send someone exactly their slice.
5. The prompting system
The thing that surprises people most: changing how QA is performed is an edit to prose in a database, not a code change and not a deploy.
Test specs and runner instructions both live in the database as Markdown:
| Template | Role |
|---|---|
aegis-qa-runner-hourly.md |
The runner's job description β served by setup_qa_runner_skill |
content.test.md |
Starting point for a new route's spec β copy, layout, CTAs, cart |
ga-tracking.test.md |
GA/pixel verification steps |
test-brief.md |
The brief a runner reads to write content.test.md. Not a test. The fallback, used when a page type has no brief of its own |
brief_<page_type>.md |
That page type's brief β brief_pdp.md, brief_landing.md. Named after the type, so renaming the type renames the brief |
test-brief-shopify-preview.md |
The brief used instead, for Shopify preview routes, whatever the page type |
lovable_proxy_autofixes |
Rules fed to Lovable for proxied apps |
lovable_shopify_autofixes |
Rules fed to Lovable for native Shopify pages |
Two layers:
- The runner skill (
aegis-qa-runner-hourly.md) is how to be a runner β the two-phase cycle, the guardrails, how to report. Edit this to change the behaviour of every runner at once, with no deploy. - The test spec (per route, versioned) is what to check on this page.
The runner cycle has two phases, and the order is deliberate:
- Phase 1 β run every queued test. The regression sweep.
- Phase 2 β draft tests for routes that have none. Only after Phase 1 has fully stopped, so a backlog of missing tests can never starve the regression sweep.
- Cleanup β clear the Shopify cart, close the tabs the run created.
Never fabricate a result. If the browser is unavailable or a step cannot be verified, the run must be submitted as
errorwith an explanation. A green run that did not happen is worse than no run at all.
Where to edit them: Admin β profile menu β Global Test Templates. Each one has version history, a diff against the live version, a description and a type, and archiving rather than deletion. They used to be reachable only through MCP tools, which meant the fleet-wide config β the runner's own job description included β was the one part of QA a person could not read without an agent.
Because the prompt is prose in a table, editing it is a QA-manager task rather than an engineering one. It is versioned, but it is not code review: change it deliberately and tell someone.
6. Slack: the routing system
Aegis events are routed to Slack channels by rule, not hardcoded.
Channels (Admin β Settings β Slack) each have a name, a Slack channel ID,
an active flag, and optionally is_default.
Aegis-QA must be invited to the channel. Aegis posts through a Slack bot,
and a bot can only post where it is a member β adding the channel in Aegis is
not enough. In Slack: channel β Integrations β Add apps β Aegis-QA.
Then press Test on the channel row; not_in_channel means the invite is
missing. Slack failures are swallowed by design, so a missing invite is silent
apart from that Test error.
The channel ID comes from the channel URL (app.slack.com/client/Tβ¦/C0AUDTK337U)
or channel β About β bottom. Webhook URLs are the legacy transport and are being
retired: a webhook is a credential, and messages posted through one cannot be
threaded onto or reacted to.
Routing rules match an event to a channel:
| Rule field | Matches on |
|---|---|
event_type |
The kind of event, or all |
match_page_type |
Landing, AI Landing, Blog, FAQ⦠|
match_origin_kind |
Lovable, Shopify, external⦠|
match_store |
uk or us |
priority |
Order rules fire in β not a tie-break. See below. |
The authoritative list is on the screen, not here. Settings β Slack β What Aegis Sends shows every update type Aegis can send, grouped, with a switch for each and a line saying which channels carry it.
That used to be a table in this document, and by 2026-09-27 it listed 14 of 33 β two of which no longer existed. A hand-copied list of something the code owns drifts the week after it is written, and a stale list is worse than none: it is the one somebody checks before concluding an update does not exist.
Worth knowing about the screen rather than the list:
- The switch there is the only lever that reaches the summaries β the morning report included. Everything else is per-channel routing.
- Each row says where it goes, with the scope beside the channel name. "It goes to #shopify-ai-testing" is rarely true on its own; what is true is "it goes there for Shopify pages".
- An update no rule carries says so explicitly, and names the default channel as a fallback rather than as a destination.
Every QA message names the store β π¬π§ or πΊπΈ since 2026-09-26, (uk) and
(us) before that. That is not decoration: a path is not unique across the two
storefronts (/pages/completestarterkit exists in both), so without it the two
copies failing in the same hour produce two identical lines that read as one
duplicated message.
Every message about a page is built the same way β glyph, the page as a link, the flag, then who did what, and the estate's own icons throughout: π for a defect, π for a test, π for a report. The channel is scanned by page, so the page leads; Slack also names a thread after the opening words of its parent, and these messages are thread anchors.
Controlling QA volume
qa_test_failed and hourly_failure_summary deliberately cover the same
events β one live, one batched. Left on the same channel that means every
failure is announced twice, and an attended /qa-run across the whole estate
can post a burst of a dozen or more.
You have two levers, and they do different things.
Route them to different places (Settings β Slack β Routing Rules). This is the usual answer β it keeps every update but puts each one where its audience is:
| Channel | Events | Why |
|---|---|---|
| Shared team channel | hourly_failure_summary, qa_test_failing_repeatedly, qa_test_signed_off |
Low volume, and everything on it is worth someone's attention. One grouped message per hour, plus escalations. |
| A QA-runs channel, or a DM to whoever runs QA | qa_test_failed, qa_test_created |
High volume and only useful while a run is happening. This is where an attended /qa-run burst belongs. |
The distinction that matters is attended versus unattended, not urgent versus routine:
- An unattended hourly sweep wants the digest. Nobody is watching, and one grouped message an hour is the right cadence.
- An attended run β someone typed
/qa-runand is sitting there β wants the live messages, because a digest forty minutes later is useless to them. Those belong wherever that person is looking, not in the shared channel.
qa_test_failing_repeatedly should stay on the shared channel whatever else you
do. It fires on the third consecutive failure and demotes the suite out of signed
off, so it is the one message that always means something changed.
Or turn one off entirely (Settings β Slack β What Aegis Sends). Each
update type has a switch there, and off means off everywhere: no channel, no
rule, not the legacy webhook. Reach for it when there is genuinely no audience
for an update β a team that never runs attended QA can switch qa_test_failed
off and keep the hourly digest, and stop announcing every failure twice.
Two things that look like this lever but are not:
- A routing rule cannot silence anything. Rules only say where an update goes. An update no rule matches still lands in the default channel.
- A route's own Slack settings cannot reach the digests.
hourly_failure_summary,daily_summaryandmonitor_summaryare sent for the whole fleet rather than for one route, so there is no route whose preferences apply to them. Muting a route stops its own alerts; it does not remove it from the digest.
If a specific route is noisy for a known reason, mute the route rather than the event β see the per-route overrides below.
Every matching rule fires. This is the part people get wrong: priority sets
the order, it does not pick a single winner. If three rules match an event, it
posts to all three channels. If no rule matches, it falls back to the channel
flagged as default.
Per-route overrides sit on the route form:
slack_notify = 0silences all Slack updates for that routeslack_channel_idsends that route's events to a specific channel and replaces rule matching entirely β no other rule fires for it
A muted route stays muted: when nothing resolves, Aegis deliberately does not fall back to the global webhook. Slack failures are also swallowed by design β a Slack outage never breaks the action that triggered the message.
Use the per-route mute for noisy test routes rather than disabling a rule that other routes depend on.
7. Slack: the monitor app
The Monitor posts on its own schedule, separately from QA runs.
- Hourly sweep (
15 * * * *) β uptime, pixel and image checks across monitored routes - Daily run β adds visual drift, and appends a traffic summary (yesterday, last 7 days, this month, all time)
A sweep-wide summary posts as monitor_summary. Because it belongs to no single
route, only rules without route filters can match it β a common source of
"why isn't the summary arriving in my channel?".
A typical message reads:
β Aegis Monitor Audit 24 passes / 1 failures (25 sites monitored)
with failed routes listed, and rate-limited routes (HTTP 429/503 after retries) called out separately β those are treated as transient, not failures.
You can trigger either run by hand from the Admin profile menu: Trigger Hourly Run and Trigger Daily Report.
8. When something looks wrong
"Is QA even running?" β Open Route Status. It shows uptime, last audit and last QA per route, and flags anything missing either. This is the first place to look, because the runner is a schedule, not a service: if Agatha is asleep, logged out, or Claude in Chrome has lost its session, nothing errors β runs simply stop happening. Silence is not success.
A test fails β classify it as drift or breakage before touching anything
(section 4). When unsure, tag
QUARANTINE-ESCALATE.
The queue looks stuck β rows stay processing if a runner dies mid-run. Use
reset_qa_test_queue to put everything back to pending.
Claude cannot see a tool β fully quit Claude (Cmd + Q) and reopen. If it is still missing, check the change is deployed; a local change is invisible to Claude.
A run reports error β that is the system working. It means the agent could
not verify something and refused to guess.
First week checklist
- Get added to the Tailscale account and confirm you can reach Agatha
- Enable the Aegis MCP connector in Claude and confirm the tools appear
- Enable Claude in Chrome
- Run
setup_qa_runner_skilland read the runner prompt end to end - Trigger Run QA Test Now on one route and follow it through to the report
- Open Route Status and note which routes are missing QA or audits
- Read one
previewspec and onesigned_offspec side by side - Find the Slack channel where
monitor_summarylands - Walk one route through all four sign-offs on a test route, not a live one