Status: Phase 1 complete · Phase 2 screenshot coverage complete: 83 automated and 86 skipped Date: 2026-09-04 Owner: Docs / Frontend
Resuming work? Jump to Current state & next steps at the bottom — it lists exactly what is done, what is next, and the workflow for adding new screenshot entries.
After each saas-frontend release, automatically:
Two manual audits (DOCS_AUDIT_2026-07-31.md, FAQ_AUDIT_2026-08-18.md) each took days of human effort and found ~30 stale screenshots, dozens of outdated statements, and orphaned images. This system runs that audit automatically after every release.
flowchart LR
A[saas-frontend release] -->|repository_dispatch| B[docs-refresh workflow]
B --> C[Playwright capture run<br/>HAR replay — no live backend]
C --> D[screenshot-manifest.json<br/>route + steps + selector per image]
D --> E[pixelmatch diff vs<br/>committed PNGs]
E -->|changed| F[Auto-commit new PNGs]
B --> G[claims-check run<br/>assert UI text matches docs claims]
G -->|drift| H[Drift report]
B --> I2[reference hygiene<br/>broken refs + orphans]
F & H & I2 --> I[PR to firetail-documentation<br/>human reviews & merges]
The core pattern: screenshots and factual claims stop being static assets and become
build artifacts generated from manifests — the same approach this repo already uses
for API specs, tags, findings and dynamic variables (scripts/*/download-*.js +
build-*.js).
Static analysis, zero risk, pure win.
scripts/images/check-image-refs.js:
docs/**/*.md (and releases/, api/) for image references
(<img src="..."> and ).docs/**/images/ never referenced by any page
(orphans). Use --fail-on-orphans to make orphans fail CI too.yarn check:imagespull.yaml) so every PR is checked.Make every doc screenshot reproducible from a recipe.
Implemented (pilot, 2026-08-21):
screenshots/manifest.json + screenshots/manifest.schema.json in this repo —
capture recipes with shared defaults (viewport 1440×900 @2x, maxDiffPixelRatio,
fixedTime, orgId) and per-entry route, steps, capture, mask.saas-frontend/e2e/docs-screenshots/ (own Playwright config +
capture.spec.ts), reusing the e2e auth setup and harReplay.ts.saas-frontend:
npm run docs:screenshots — replay mode (deterministic, hermetic; CI-safe)npm run docs:screenshots:record — record HARs against a real env, then
sanitize credentials out of themsaas-frontend under
e2e/docs-screenshots/fixtures/<entry-id>.har — not in this repo. Two reasons:
the recordings contain internal API data (org/user identifiers, usage numbers)
that must not live in the public docs repo even after credential sanitization,
and a HAR records API response shapes that track app versions, so it belongs
with the app code (next to the existing e2e HAR fixtures) and is re-recorded by
frontend developers. The manifest (docs-owned: which screenshots exist, where
they land) stays here; the fixtures (app-owned: what the API returned) live
with the runner.pixelmatch; PNG overwritten only above maxDiffPixelRatio.saas-frontend/e2e/docs-screenshots/capture-report.md
(feeds the future docs-refresh PR body).unchanged in the report) for the two pilot entries
(workforce-platforms-list, workforce-platform).Scaled out — workforce section complete (2026-08-21):
docs/workforce/images/*.png are now manifest-managed: platforms list +
drawer, employee drawer + edit form, group drawer + edit form, device drawer +
edit form, application drawer, Create Policy form, scoped Create Policy form
(opened from an employee's Policies tab, showing the read-only "Scope set to"),
provided-guardrail detail modal, Platform Rules tab (clipped to tabs + table).
All recorded, captured, and verified deterministic (4+ replay runs →
byte-identical PNGs, all unchanged).ai-policy-precedence.png — an illustration/diagram of policy evaluation
logic, not an app screenshot; cannot be captured from the UI.within (scopes a step's locator to
a role/name container, e.g. a drawer, to disambiguate from the page behind),
capture.target: "clip" (viewport-relative rectangle for partial-page shots),
and capture.target: "role" for dialog-only screenshots.Workload screenshots (2026-09-02):
docs/workload/images/*.png are now manifest-managed: model details,
agent details, prompt messages, the services listing, Available Integrations,
and the four AWS Bedrock service-detail views. Each has a sanitized HAR
fixture and passed hermetic replay. The service-detail routes use a 30-day
dateTime query range to stay within the sandbox plan's 90-day retention
window.Dashboard screenshots (2026-09-02):
dateTime query range so the fixed capture clock stays within the
sandbox plan's 90-day retention window.Alerts, FAQ download and integrations catalogue screenshots (2026-09-04):
docs/alerts/), the Findings and Events
Download Data dialogs (docs/faqs/), and the Settings > Integrations >
Available catalogue (docs/integrations/). Each has a sanitized HAR fixture
and passed hermetic replay twice with byte-identical PNGs.AI Static Alerting; the
delete entry opens the confirmation popconfirm but never confirms it.Current screenshot inventory (2026-09-04):
docs/**/images/; 86 intentionally skipped images
are third-party flows, generated pages, logos, or illustrations. No images
remain in planned or needs-triage.screenshots/image-inventory.json; ai-policy-precedence.png
is intentionally excluded because it is a diagram, not a UI screenshot.Lessons learned in the pilot:
page.clock.setFixedTime freezes requestAnimationFrame, hanging Ant Design
drawer animations and page.screenshot — use page.clock.install({ time })
instead (clock starts fixed but keeps ticking).animations: 'disabled' on screenshots must be avoided for the same reason.harReplay.ts now replays duplicate-key responses in
recorded order (repeating the last), otherwise the app refetches page 1 forever.networkidle wait after the steps is needed in record mode so lazy panel
requests are fully recorded rather than captured as aborted; the same wait is
needed after the capture (before the context closes) or trailing requests
record as aborted too.harReplay.ts leaves them pending on replay (an
instant abort beats the app's own cancellation and pops the global "Network
Error" modal into the screenshot — which then silently overwrites the good
PNG). The runner now also refuses to capture any frame showing the app's
error dialog, failing the entry instead.waitForLoadState('networkidle') is not enough to guarantee complete
recordings — it only needs a 500 ms lull and times out on slow endpoints,
leaving unfinished (status -1) HAR entries that replay as permanent spinners.
The runner now tracks in-flight API requests explicitly and waits until
all have settled (plus a quiet period) before capturing and before closing a
recording context, and refuses to capture frames with visible
spinners/skeletons..../policies/platform-rules) over clicking
through tabs: intermediate tab mounts fire extra requests non-deterministically,
producing racy HAR misses.[har] no exact query match … falling back but does not fail. Treat any
occurrence as a re-record signal: grep the run output for it. The suite
currently triggers it zero times.fallbackMac originally recognised
only 02:00:00:00:00:xx as already-masked while generating
02:00:00:xx:xx:xx, so every re-run re-hashed masked MACs and rewrote 15
untouched fixtures. docIp had the same latent bug for the RFC 5737 ranges.
Both now short-circuit on their own output.maxDiffPixelRatio — a name or an
email is far below 0.1% of the pixels, so real data silently survived into
the committed image. The runner now stops after the guards in record mode.docs:screenshots:sanitize must run both scrubbers: sanitize-har.mjs
(auth headers and bearer tokens) and then beautify-docs-fixtures.mjs (PII
and scruffy names). The npm script previously ran only the latter, so
hand-recorded fixtures kept live JWTs. The beautifier's "review leftovers:
jwt-like token" line is the signal that the first pass was skipped.wait step after the dialog appears and before opening the select;
stretching the post-step settle does not fix it.defaults.fixedTime is 2026-08-21; fixtures recorded on
2026-09-04 rendered every timestamp in the future ("Created in 14 days").
Entries now carry an optional fixedTime override — set it whenever you
record on a date far from the default.maxDiffPixelRatio (0.001) is too coarse to catch text. A changed name,
email or relative date is a few hundred pixels out of ~5M, so replay reports
unchanged and keeps the stale PNG. Delete the output file to force a
rewrite after any change that only affects text. This also weakens Phase 2's
drift detection and should be revisited with cross-environment evidence
before tightening the default.dateTime query
(?dateTime=%7B%22value%22%3A2592000%7D). The default range is 90 days and
the sandbox plan's retention is exactly 90 days, so recording fails with
"Bad Request — Start date is older than your plans retention period".Projects and organizations screenshots (2026-09-04):
docs/projects/), and the Members
table, Edit Team Member role selector and the all-organizations grid
(docs/organizations/).fill (role/name +
value, with optional nth for repeated key/value rows), press (a single
keyboard key, e.g. Enter to commit an Ant Design tag chip) and wait
(ms). nth also works on clickRole/waitForRole, which is how the
Members table's per-row edit buttons are addressed — they all share the
accessible name edit.Getting-started and posture-management screenshots (2026-09-04):
docs/getting-started/), and the Events listing, Provided
Policies grid, the empty Create Resource Policy drawer and its Add check
menu (docs/posture-management/).docs/posture-management/images/event-details.png was pulled from this batch:
the Events page has neither a working free-text search for event types nor a
filter, so a RESOURCE_POLICY:CREATED event cannot be reached
deterministically. Events are deep-linkable
(/platform/events/<uuid>), so the entry becomes trivial once a stable
sandbox event of that type is identified.Integration forms and suggested checks (2026-09-04):
docs/integrations/), plus the Resource
Policy Suggested Checks picker (docs/posture-management/). All replay from
sanitized HAR fixtures.ServiceNow Issue,
Jira Issue, Lambda Function, PagerDuty Issue, and HTTP Webhook (HMAC Signed). The latter replaces the documented generic HTTP webhook name.Remaining: add the docs-refresh workflow — see Current state & next steps.
Turn fragile factual claims in the docs into machine-checkable assertions.
claims/<doc-slug>.json per high-churn page:
{
"doc": "docs/workforce/workforce-platforms.md",
"assertions": [
{ "route": ".../platforms", "role": "columnheader",
"names": ["Name", "Capabilities", "Employees", "Logs"] },
{ "route": ".../platforms", "role": "button", "name": "Download" }
]
}
A Playwright run asserts these against the released frontend (using the same HAR
replay fixtures as Phase 2, so no live backend is needed) and emits a markdown
drift report: "workforce-platforms.md claims column 'Risk Level'; UI now
shows 'Risk Score'".
Assertions use getByRole-style semantic selectors, matching the saas-frontend
e2e conventions.
Anything that is a list of facts should be rendered by Eleventy from downloaded data, like findings/specs/tags already are. Candidates from the audits:
Once generated, this content can never go stale.
saas-frontend release workflow fires repository_dispatch on this repo
(fallback: nightly cron).The system detects all three (manifest coverage vs. app routes, failing claims, orphaned images) and surfaces them in the PR report.
When/if an LLM is added, it slots between drift report → PR: given the precise, grounded drift findings, it drafts the prose updates for human review. All the manifest/diff infrastructure built here is exactly the grounding it needs — nothing is throwaway.
| Phase | Effort | Status |
|---|---|---|
| 1 — Reference hygiene | days | ✅ Implemented |
| 2 — Screenshot capture via HAR replay (~170 PNGs, staged) | 1–2 weeks | ✅ Complete; 83 automated, 86 skipped; triage complete |
| 3 — Claims manifests (high-churn pages first) | 1 week after P2 | Planned |
| 4 — Generated content | ongoing | Planned |
flowchart TB
A["🚀 New product release goes live"] --> B["🤖 Docs Checker starts automatically"]
B --> C["📸 Re-takes every screenshot<br/>by clicking through the product,<br/>exactly like a user would"]
B --> D["🔍 Reads the docs and checks<br/>every fact against the product<br/>(button names, menus, columns...)"]
B --> E["🧹 Tidy-up check:<br/>finds broken or unused images"]
C --> F{"Did anything change?"}
D --> F
E --> F
F -- "No" --> G["✅ Docs are up to date.<br/>Nothing to do."]
F -- "Yes" --> H["📦 Prepares an update package:<br/>• fresh screenshots<br/>• list of outdated sentences<br/>• list of images to remove"]
H --> I["🧑💻 A person reviews the package<br/>and approves with one click"]
I --> J["📚 Documentation website<br/>is updated"]
Everything needed to resume this work in a fresh session.
| What | Where |
|---|---|
| This plan | firetail-documentation/internal-reports/DOCS_AUTOMATION_PLAN.md |
| Screenshot manifest (docs-owned) | firetail-documentation/screenshots/manifest.json (+ manifest.schema.json) |
| Complete image inventory | firetail-documentation/screenshots/image-inventory.json (+ image-inventory.schema.json) |
| Capture runner | saas-frontend/e2e/docs-screenshots/capture.spec.ts (+ own playwright.config.ts) |
| HAR fixtures (app-owned, sanitized) | saas-frontend/e2e/docs-screenshots/fixtures/<entry-id>.har |
| HAR replay helper (shared with e2e) | saas-frontend/e2e/utils/harReplay.ts |
| HAR sanitizer | saas-frontend/e2e/scripts/sanitize-har.mjs [harDir] (credentials) + beautify-docs-fixtures.mjs [harDir] (PII); both run by npm run docs:screenshots:sanitize |
| Capture report (PR-body input) | saas-frontend/e2e/docs-screenshots/capture-report.md |
| Image hygiene check (Phase 1) | firetail-documentation/scripts/images/check-image-refs.js (yarn check:images) |
saas-frontend dev server running (npm start, serves https://localhost:3000)e2e/.env with BASE_URL=https://localhost:3000 and TEST_USER_EMAILftauth) — the auth setup fetches the test user
password from SSM; a still-valid token in e2e/.playwright/.auth/user.json
skips login entirely94520fbd-7863-465b-bc77-70038e014aea (manifest defaults.orgId)npm run docs:screenshots) needs none of the above except the
dev server (or any deployed frontend via BASE_URL)grep -rn "<image>.png" firetail-documentation/docs/)playwright-cli -s=firetail-sandbox-authenticated-session (or
npm run e2e:auth) to find the route, step selectors
(role/name), and the dialog's accessible name.screenshots/manifest.json. Prefer deep routes over tab
clicking; use within to scope steps inside drawers; pick capture.target:
viewport (full page), role (a dialog only), or clip (rectangle).npm run docs:screenshots:record (records HAR + captures + sanitizes).
To re-record a single entry:
UPDATE_HAR=1 npx playwright test --config=e2e/docs-screenshots/playwright.config.ts -g "<entry-id>" && npm run docs:screenshots:sanitize
A record run never writes the PNG — it only proves the recipe resolves and
produces the fixture.npm run docs:screenshots twice — everything must pass and all
PNG hashes must be identical (shasum docs/**/images/*.png). If a change
only affects text, delete the output PNG first — the diff threshold treats
it as unchanged.yarn dev in the
docs repo) — data, framing, crop.node -e "console.log(JSON.parse(require('fs').readFileSync('e2e/docs-screenshots/fixtures/<id>.har','utf8')).log.entries.filter(e=>e.response.status<0).length)"
(should print 0).screenshots/image-inventory.json lists every image under docs/**/images/,
its documentation references, and one of four statuses: automated,
planned, skipped, or needs-triage.automationId and references are derived from the screenshot manifest and
Markdown files. Reviewers maintain planned decisions with required capture
notes and skipped decisions with required reasons; manifest outputs are
always marked automated.yarn build:image-inventory after adding, removing, or reclassifying an
image. CI runs yarn check:image-inventory and fails when the inventory is
stale, incomplete, or points to a missing manifest output.build + pull.yaml CI.
Current warning: 4 broken refs in draft browser-extension docs. Published pages
have no broken references, and there are no orphaned images.ai-policy-precedence.png, a hand-drawn diagram). All 13 entries are verified
deterministic. Sandbox data the steps depend on: employee
charles.miller@tropical.com, group sales team, device DESKTOP-7804TS5,
application Claude Desktop, platform ChatGPT, guardrail
Personal PII Blocking — if the sandbox org loses these records, re-pick
targets and re-record.Claude Sonnet 4, agent
Langchain Agent Python, prompt langchain_agent_python/main.py:84-88, and
service AWS Bedrock. The four service-detail routes use a 30-day query range
to remain within the sandbox plan's 90-day retention window.test-new-project
(94057967-a46a-4ef6-9c2d-e9b37de0ba5d, beautified to customer-portal) and
the second row of Settings > Members. The whole suite (60 tests) was run
twice end to end with byte-identical PNGs and no updated rows.updated rows. Sandbox data the steps depend on: the sandbox must have
at least one custom resource policy so the Create Resource Policy button is
present, and the Events listing must be non-empty.AI_WORKFORCE_USER_CONSUMPTION:DISCOVERED and Resource
Policy history uses a preselected matching-resource history entry.userUUID, which the workforce-employee,
workforce-employee-form and workforce-policy-scoped fixtures predated.
All three were re-recorded and the whole suite (54 tests) passes.
workforce-employee.png changed as a result: the Consumptions tab is now
empty, because charles.miller@tropical.com genuinely has no
consumptions. The previous screenshot's populated table was org-wide data
substituted by the shape-key fallback — it never belonged to that employee.
Open decision: only three sandbox employees have consumption data and all
three are personas for real staff accounts, so targeting one would put a real
colleague's address in this public manifest. Either keep the accurate empty
state or seed consumption data for a fictional employee
(charles.miller@tropical.com) in the sandbox.saas-frontend — it has the
runner, fixtures and app build; triggered by release + nightly cron):
checkout both repos → build/serve frontend → npm run docs:screenshots
(replay) → if PNGs changed, open a PR against firetail-documentation with
the changed images + capture-report.md as the PR body. Never auto-merge.saas-frontend: current branch is FIRE-4033-navigation-system; the docs
screenshot runner, fixtures, HAR replay changes, npm scripts, and dependencies
are present in the working repository.firetail-documentation: current branch is FIRE-4037-docs-automation; Phase 1,
the manifest, and refreshed screenshots are present here.