# E2E against production without a test environment

An end-to-end suite can run against a live storefront when every test stops before it causes a side effect. My Playwright suite for a production Shopify storefront has 45 tests on desktop and mobile: forms are filled and validated but never submitted, only one dedicated test account is written to, live data is discovered rather than pinned, and checkout is covered only up to the handoff.

Published: 2026-09-28
Canonical: https://umar.codes/e2e-against-production

## Revisions

- 2026-09-28 — Created; body written from the suite's specs, plans, handoff notes, and 39 commits between 2026-06-29 and 2026-08-04.
- 2026-09-28 — Published on its slot; auto-drafted by the publishing run because no draft existed for the row.

---

You can run a meaningful end-to-end suite against a live storefront if every
test stops at the point where it would cause a side effect. Over five weeks I
built 45 Playwright tests against a production Shopify storefront, run on
desktop and mobile, that fill forms without submitting them and write only to
one dedicated test account.

## Why the suite points at the deployed site

The suite lives in its own repository and contains no storefront code. Its
only link to the application is a `BASE_URL`, which defaults to the live
headless storefront customers shop on. The behaviours I needed to test came
from the live catalogue, the live search index, and the live third-party
scripts, and the suite had no other environment to point at. Shopify's test-mode
checkout would have needed store access the project did not have, so the
boundary was set on day one: coverage stops at the cart-to-checkout handoff.

The suite started on 2026-06-29 with a foundation and 12 tests across the
guest shopping journey and search. On 2026-07-09 I compared it against the
storefront source and wrote a coverage plan: 9 suites in 4 waves ordered by
revenue risk, one issue per suite, each one run
[spec-first](/spec-driven-development) as design, then plan, then pull request.
By 2026-08-04 six of the nine had merged and the suite stood at 45 tests. The
plan had estimated 32 new tests for those six suites; 33 shipped.

Each test runs in two Playwright projects, desktop Chromium and a Pixel 5
profile, so a full run is 90 test executions. The last full run was 86
passed, 4 skipped, 0 failed. The skips are by design: a desktop-only
navigation test does not run on mobile, its mobile twin does not run on
desktop, and tests that depend on live data skip when the data is absent.

## Assert up to submit, never submit

The first ground rule in the plan: any flow with an external side effect is
tested by filling the form, triggering validation, and asserting that the
submit control is enabled. Nothing is ever submitted. That covers customer
registration, password-recovery emails, newsletter signup, back-in-stock
requests, and review submission, each of which would write to a real system
that the storefront's owners read.

Up-to-submit is weaker than a full submission, so each test has to find a
signal that the form really would have worked. The gift-card form's submit
button is disabled until the form library reports the form valid, so the test
asserts disabled and then enabled. The newsletter form's button is never
disabled; its handler returns early at a captcha check before any request, so
the test asserts the error chip and confirms zero requests to the signup
endpoint. The password-recovery form is a native form, and its only
validation signal is a field error that appears and disappears.

One test asserts on a request that the page itself cannot show. A logged-in
checkout goes through a multipass hop chain across three domains, and it arrives
at the same final URL shape as a guest checkout. Only the intermediate hops
show that the customer stayed logged in, so the test starts a
`waitForRequest` for the hop before it clicks checkout.

## One test account, two browser projects writing to it

Some tests do write: the address book and the profile form act on a
dedicated test customer, and nothing else. The credentials live in a
gitignored `.env`; a global setup logs in once and saves the session to a
gitignored storage-state file, and authenticated specs opt in to it.

Two projects running in parallel against one account will collide unless the
tests are written for that. Address rows carry an `E2E <project>` tag, and a
test only touches rows with its own tag. The profile test splits its fields
by project, last name on desktop and first name on mobile, and reverts to a
baseline pinned in the test data. The tests also assume a crash can leave
residue from an earlier run, so they converge on the expected state instead
of assuming a clean start.

Logout is the exception to the shared session. Signing out can invalidate
the token that every other authenticated test is using, so the logout test
does its own login through the UI and never touches the saved session.

## Live data churns, so tests discover instead of pinning

A production catalogue changes daily, so the resilience rule is structural:
pick the first available product, never assert a specific name, price, or
count. Where a test needs a specific state, it finds one. The sold-out
product tests run one cached Storefront API scan per worker over the first
100 products of the main collection, looking for a variant that is sold out
with a back-in-stock flag, or a product with both in-stock and sold-out
sizes. When the scan finds nothing, the test skips with a reason. It never
fails, because an empty result is a fact about today's stock, not a bug.

Handles that must be pinned carry a dated verification comment and, where it
helps, the command that re-verifies them. The one path the suite requests to force a 404
ends in `-e2e`, so it stays easy to identify in the storefront's analytics.

## Vacuous passes are the real production risk

A test against a live site fails loudly when the site breaks. The quieter
risk is an assertion that can never fail. I found four of them.

The wishlist page renders its count unconditionally, so an empty wishlist
shows the literal text "0 products". The planned assertion, "empty state or
product count visible", would have passed forever; the test now asserts on
structure instead. The reviews widget renders nothing at zero reviews and
also nothing when its API errors, so the smoke test pins a product with 9
reviews, which keeps an absence loud. Collection grids render a bare empty
list with no empty state, and the filter drawer only offers facets with one
or more results, so "filter down to zero" cannot be reached through the UI.
That test was replaced by one for over-range page numbers, which redirect to
the last valid page. And the desktop mega menu had no panel data on any of
its 4 top-level items, so the panels have no test; the navigation test
covers the links.

The site's [live regions](/react-aria-live) set a similar trap. They render
once at the body level, so a bare `getByRole('status')` matched a site-wide
polite region that is routinely non-empty, and the locator had to exclude it
by id.

## When the flake is the storefront, not the test

The gift-message test failed on roughly 25–30% of desktop runs. The
diagnosis showed that the click landed and no network request followed: the
submit handler was intermittently detached while the component re-mounted.
That is a real storefront bug, and customers hit it too. I filed it with the
storefront team, and the test carries a documented `toPass()` retry that
names the bug and should be removed when the fix ships.

Production also has bad days that are not bugs. During one session, listing
pages took more than 30 seconds to load and filters stalled. Before I blamed
the refactor on the branch, I ran the unchanged main branch in a separate
worktree against the same site and it failed the same way. That check told
me the code was fine before I spent a day debugging it. In CI
the config allows 2 retries on 1 worker and records a trace on the first
retry, which keeps an environmental failure distinct from a deterministic
one.

## What a production suite cannot test

Some gaps are permanent under these rules. Payment and order placement sit
past the checkout handoff on a domain the suite does not control. Real
registrations and real signups are out. Server-to-server surfaces, such as
webhooks, are not browser journeys. Any state the UI cannot reach, like the
zero-result filter, stays untested unless someone adds a way to reach it.

Some gaps are only unfinished. Three of the nine suites have not shipped:
content-page smokes, accessibility scans with a committed baseline, and
visual regression with a scheduled run. The scheduled run is the important
one. Production regressions arrive without a pull request, so a suite that
only runs on a pull request cannot catch them. Until a nightly run exists,
the suite protects changes to the storefront, not the storefront itself.
