Tests that pass alone and fail in parallel CI usually share something mutable outside the test: a backend record, a user account, a file, a database namespace, or a global setting. Parallel workers isolate process memory. They do not isolate the world those processes talk to. The fix is a sequence: find the shared state, decide who owns it, isolate it, and restrict concurrency only where a resource truly requires it.
Why parallel workers expose hidden coupling
Playwright Test runs test files in parallel by default, each in its own worker process. Tests within a single file run in order by default. Workers don’t share process state or globals, so a variable set in one worker is invisible to another. That is real isolation, but only of memory.
As an Amazon Associate I earn from qualifying purchases.
Two workers can still log in as the same account, edit the same backend record, write to the same file, or query the same database rows. Likewise, a fresh browser context gives each test clean cookies and storage, but it does nothing to separate what the server holds. The pytest documentation makes the general point in its “Flaky tests” guidance: “Broadly speaking, a flaky test indicates that the test relies on some system state that is not being appropriately controlled – the test environment is not sufficiently isolated.” It also names ordering dependencies and missing cleanup as causes that parallel runs make visible.
That also explains “passes locally, fails in CI.” Locally you often run one file, one worker, or a fresh database, so order and timing hide the coupling. CI adds concurrency and leftover state from earlier runs, and the coupling shows up.
#1 Best Overall
Step 1: Find the shared state
For each failing test, ask whether it touches any of these:
- A hard-coded identity: a fixed username, email, order number, or record name that more than one test creates, edits, or deletes.
- A shared account: one login whose settings, cart, or permissions tests change.
- A file path: a fixed download, export, or output location.
- A database or namespace: rows, tables, or queues that tests read as well as write.
- A precondition set by another test: a test that only passes because an earlier one created something.
- Skipped cleanup: teardown that doesn’t run when a test fails, leaving stale data for the next one.
As a diagnostic, rerun the failing tests with different worker counts or in a different order. If failures appear or vanish, contention or order dependence is likely. This points you toward a cause; it doesn’t prove one, so confirm by finding the shared resource itself.
Rank #2
Step 2: Assign ownership
Every piece of mutable state should have one clear owner: a test, a worker, or, deliberately, nobody (read-only data). Playwright’s guidance maps onto these levels, and each has a different cost.
| Approach | Isolation granularity | Cost | Use when |
|---|---|---|---|
| Unique record per test | Per test | Setup and cleanup on every test | Tests create or edit the same kind of record |
| Data set per worker | Per worker | Setup once per worker; tests in a worker still share it | Creating data is expensive and tests can safely reuse it |
| Named lock | Per external resource | Tests needing the resource wait in line | A resource can’t support concurrent access |
| Single worker | Whole run | Slowest wall-clock time | Stability and reproducibility come first |
| Sharding | Across CI jobs | More CI jobs | The problem is total duration, not contention |
The sources don’t quantify the speed or cost of these options, and they give no universal rule for the right worker count. It depends on your infrastructure capacity, any rate limits on external services, and how costly isolated data is to create.
Rank #3
Step 3: Isolate the data
Per-test records
When tests create or edit the same kind of record, derive a unique identifier for each test. Playwright’s documentation illustrates using testInfo.testId for this. A minimal sketch:
import { test as base } from '@playwright/test';
export const test = base.extend({
projectName: async ({ request }, use, testInfo) => {
const name = `project-${testInfo.testId}`;
await request.post('/api/projects', { data: { name } });
await use(name);
await request.delete(`/api/projects/${name}`);
},
});
The endpoint is a placeholder for your own API. Putting creation and deletion in a fixture keeps setup inside the test’s own scope, so it doesn’t depend on another test. Teardown code after use still runs when the test fails, which addresses the skipped-cleanup problem.
Per-worker data
If creating data per test is too costly and tests can safely share it, make it worker-scoped and distinguish each worker’s users by worker index, for example testInfo.workerIndex in the account name. Clean it up in the worker-scoped fixture’s teardown. Each worker owns its own set, so workers never collide, though tests inside one worker still share that data and must not corrupt it for each other.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsFiles
Give each test a unique file path rather than a fixed one. Playwright’s guidance covers this under avoiding shared state, and deriving the path from the test’s identifier or its per-test output directory is a natural way to do it.
Best Value
Databases and setup
Playwright’s best-practices guidance says to control the data you test against, using a staging environment that doesn’t change, instead of depending on whatever a live database contains. Set up what each test needs inside that test or its fixtures. If one test’s side effect is another test’s precondition, parallelism and reordering will break it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Step 4: Constrain concurrency only where required
Isolation is preferable to serialization, because serialization removes the speed you wanted from parallelism. When a resource truly can’t handle concurrent access, Playwright documents named test locks. Tests that need that resource take turns, and unrelated tests keep running in parallel. This is narrower than dropping the entire suite to one worker.
For the whole run, Playwright’s Continuous Integration documentation recommends a single worker in CI to prioritize stability and reproducibility. That is the framework’s guidance, not a rule for every runner or environment, and it trades away speed. If you want more throughput, sharding distributes tests across CI jobs, each running its own subset. Sharding addresses job duration. It does not fix shared state: if two shards use the same account or database, they can collide just as workers do, so isolate data across shards too.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →A practical order of operations
- Reproduce: rerun the failures at different worker counts and orders.
- List the shared records, accounts, files, and global settings the failing tests touch.
- Replace hard-coded identities with per-test identifiers; use per-worker data where per-test creation is too expensive.
- Move required setup and cleanup into fixtures so teardown runs on failure.
- Give files unique paths.
- Lock only the resources that still can’t be shared.
- Only then consider a single worker or sharding, based on stability needs and job duration.
No reviewed official source offers a statistic on how common these failures are, and the worker and shard numbers in official examples are configuration illustrations, not measured results.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




