E2E against production without a test environment

You can run a meaningful end-to-end suite against a live storefront if every test stops at the point where it would cause a side effect. Over five weeks I built 45 Playwright tests against a production Shopify storefront, run on desktop and mobile, that fill forms without submitting them and write only to one dedicated test account.

Why the suite points at the deployed site

The suite lives in its own repository and contains no storefront code. Its only link to the application is a BASE_URL, which defaults to the live headless storefront customers shop on. The behaviours I needed to test came from the live catalogue, the live search index, and the live third-party scripts, and the suite had no other environment to point at. Shopify’s test-mode checkout would have needed store access the project did not have, so the boundary was set on day one: coverage stops at the cart-to-checkout handoff.

The suite started on 2026-06-29 with a foundation and 12 tests across the guest shopping journey and search. On 2026-07-09 I compared it against the storefront source and wrote a coverage plan: 9 suites in 4 waves ordered by revenue risk, one issue per suite, each one run spec-first as design, then plan, then pull request. By 2026-08-04 six of the nine had merged and the suite stood at 45 tests. The plan had estimated 32 new tests for those six suites; 33 shipped.

Each test runs in two Playwright projects, desktop Chromium and a Pixel 5 profile, so a full run is 90 test executions. The last full run was 86 passed, 4 skipped, 0 failed. The skips are by design: a desktop-only navigation test does not run on mobile, its mobile twin does not run on desktop, and tests that depend on live data skip when the data is absent.

Assert up to submit, never submit

The first ground rule in the plan: any flow with an external side effect is tested by filling the form, triggering validation, and asserting that the submit control is enabled. Nothing is ever submitted. That covers customer registration, password-recovery emails, newsletter signup, back-in-stock requests, and review submission, each of which would write to a real system that the storefront’s owners read.

Up-to-submit is weaker than a full submission, so each test has to find a signal that the form really would have worked. The gift-card form’s submit button is disabled until the form library reports the form valid, so the test asserts disabled and then enabled. The newsletter form’s button is never disabled; its handler returns early at a captcha check before any request, so the test asserts the error chip and confirms zero requests to the signup endpoint. The password-recovery form is a native form, and its only validation signal is a field error that appears and disappears.

One test asserts on a request that the page itself cannot show. A logged-in checkout goes through a multipass hop chain across three domains, and it arrives at the same final URL shape as a guest checkout. Only the intermediate hops show that the customer stayed logged in, so the test starts a waitForRequest for the hop before it clicks checkout.

One test account, two browser projects writing to it

Some tests do write: the address book and the profile form act on a dedicated test customer, and nothing else. The credentials live in a gitignored .env; a global setup logs in once and saves the session to a gitignored storage-state file, and authenticated specs opt in to it.

Two projects running in parallel against one account will collide unless the tests are written for that. Address rows carry an E2E <project> tag, and a test only touches rows with its own tag. The profile test splits its fields by project, last name on desktop and first name on mobile, and reverts to a baseline pinned in the test data. The tests also assume a crash can leave residue from an earlier run, so they converge on the expected state instead of assuming a clean start.

Logout is the exception to the shared session. Signing out can invalidate the token that every other authenticated test is using, so the logout test does its own login through the UI and never touches the saved session.

Live data churns, so tests discover instead of pinning

A production catalogue changes daily, so the resilience rule is structural: pick the first available product, never assert a specific name, price, or count. Where a test needs a specific state, it finds one. The sold-out product tests run one cached Storefront API scan per worker over the first 100 products of the main collection, looking for a variant that is sold out with a back-in-stock flag, or a product with both in-stock and sold-out sizes. When the scan finds nothing, the test skips with a reason. It never fails, because an empty result is a fact about today’s stock, not a bug.

Handles that must be pinned carry a dated verification comment and, where it helps, the command that re-verifies them. The one path the suite requests to force a 404 ends in -e2e, so it stays easy to identify in the storefront’s analytics.

Vacuous passes are the real production risk

A test against a live site fails loudly when the site breaks. The quieter risk is an assertion that can never fail. I found four of them.

The wishlist page renders its count unconditionally, so an empty wishlist shows the literal text “0 products”. The planned assertion, “empty state or product count visible”, would have passed forever; the test now asserts on structure instead. The reviews widget renders nothing at zero reviews and also nothing when its API errors, so the smoke test pins a product with 9 reviews, which keeps an absence loud. Collection grids render a bare empty list with no empty state, and the filter drawer only offers facets with one or more results, so “filter down to zero” cannot be reached through the UI. That test was replaced by one for over-range page numbers, which redirect to the last valid page. And the desktop mega menu had no panel data on any of its 4 top-level items, so the panels have no test; the navigation test covers the links.

The site’s live regions set a similar trap. They render once at the body level, so a bare getByRole('status') matched a site-wide polite region that is routinely non-empty, and the locator had to exclude it by id.

When the flake is the storefront, not the test

The gift-message test failed on roughly 25–30% of desktop runs. The diagnosis showed that the click landed and no network request followed: the submit handler was intermittently detached while the component re-mounted. That is a real storefront bug, and customers hit it too. I filed it with the storefront team, and the test carries a documented toPass() retry that names the bug and should be removed when the fix ships.

Production also has bad days that are not bugs. During one session, listing pages took more than 30 seconds to load and filters stalled. Before I blamed the refactor on the branch, I ran the unchanged main branch in a separate worktree against the same site and it failed the same way. That check told me the code was fine before I spent a day debugging it. In CI the config allows 2 retries on 1 worker and records a trace on the first retry, which keeps an environmental failure distinct from a deterministic one.

What a production suite cannot test

Some gaps are permanent under these rules. Payment and order placement sit past the checkout handoff on a domain the suite does not control. Real registrations and real signups are out. Server-to-server surfaces, such as webhooks, are not browser journeys. Any state the UI cannot reach, like the zero-result filter, stays untested unless someone adds a way to reach it.

Some gaps are only unfinished. Three of the nine suites have not shipped: content-page smokes, accessibility scans with a committed baseline, and visual regression with a scheduled run. The scheduled run is the important one. Production regressions arrive without a pull request, so a suite that only runs on a pull request cannot catch them. Until a nightly run exists, the suite protects changes to the storefront, not the storefront itself.

Revisions

  1. Created; body written from the suite's specs, plans, handoff notes, and 39 commits between 2026-06-29 and 2026-08-04.
  2. Published on its slot; auto-drafted by the publishing run because no draft existed for the row.