To set up a trustworthy online experiment, define who is eligible, how each eligible unit is assigned, the intended share for every variant, and what event counts as exposure. Keep assignment stable for that unit and compare observed counts with the configured allocation. A sample-ratio mismatch (SRM) is a warning that the experiment’s data may be skewed; investigate its source before relying on an effect estimate.
What sample-ratio mismatch means
SRM occurs when the observed number of units in experiment arms differs from the configured allocation by more than ordinary random variation would explain. For example, a test configured for a 50/50 split might appear as 60/40. Statsig uses that as an illustrative example in a 2025 product update; it is not a universal cutoff or a finding about a particular experiment.
A mismatch does not prove that the treatment caused harm, and an imbalance alone does not automatically make a test unusable. It does mean you should check the assignment and measurement pipeline before interpreting results. Microsoft Research describes passing an SRM check before effect analysis as a safeguard against untrustworthy conclusions. In its September 14, 2020 article, “Diagnosing Sample Ratio Mismatch in A/B Testing,” Microsoft Research writes: “To prevent that harm, at Microsoft, every A/B test must first pass this Sample Ratio Mismatch (SRM) test before being analyzed for its effects.”
Choose the assignment unit to fit the experiment
The assignment unit is the entity that receives a variant. Choose it to match both the product journey and the outcome you intend to measure; there is no single identifier that suits every experiment. Statsig’s overview uses user IDs, device-level stable IDs, and session IDs as examples of the tradeoffs.
#1 Best Overall
| Assignment unit | When it can fit | Tradeoff to check |
|---|---|---|
| User ID | When outcomes should follow a signed-in person across visits and devices. | It cannot assign a visitor before sign-in. (Statsig overview) |
| Device-level stable ID | When anonymous or first-visit behavior must be included. | It is device-bound, so one person using multiple devices may be represented by multiple units. (Statsig overview) |
| Session ID | When the outcome is contained within one visit and treating sessions as independent fits the use case. | It does not provide persistent assignment across visits. (Statsig overview) |
Before choosing, check whether the identifier persists for the journey, covers the population you need, can be null or regenerated, and can be logged consistently with the outcome. Microsoft Research identifies faulty IDs and incorrect bucketing as possible assignment-stage causes of SRM.
Set up assignment and exposure measurement
- Define eligibility and allocation. Record the eligible population, targeting and exclusion rules, and configured share for every arm. Allocations need not be equal: the expected counts must follow the actual intended proportions. Keep eligibility rules stable and account for any ramp changes when evaluating counts.
- Assign consistently. Use a documented identity and bucketing policy so a returning unit stays in the same variant unless the design intentionally specifies otherwise. Decide how missing, duplicated, or changing IDs are handled rather than letting fallback behavior vary silently.
- Separate assignment from exposure. Log which variant a unit was assigned and define the event that means it actually encountered the treatment. Assignment does not guarantee that the person saw the experience. If you monitor exposure counts, make sure the exposure event is possible in both arms.
- Preserve the randomized unit through the data pipeline. Confirm that joins, deduplication, and analysis retain the same unit used for randomization. A user-level assignment should not accidentally become a row-level count where repeat events inflate one arm.
- Validate before interpreting results. Check assignment records, variant rendering, exposure events, identity handling, and arm-specific collection. Monitor allocation while the test runs and before interpreting metric lifts. Automatic exposure logging can help, but it does not validate the full pipeline by itself.
Check observed counts against the configured split
Count unique units at the same randomization level used by the experiment, then compare the observed counts with the configured proportions—not with an assumed 50/50 split. If there are N eligible assigned units and arm i is configured for proportion pi, its expected count is N × pi. A chi-squared check compares observed counts with those expected counts; Statsig documents this approach for its SRM checks.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
Be explicit about which stage the counts describe. Assignment counts help check bucketing and eligibility; exposure counts can reveal problems between assignment and seeing the experience. Do not compare counts built from different populations, units, or inclusion windows and treat the result as an assignment test. Statistical alert thresholds and procedures vary by platform and policy, and the available guidance does not establish one universal p-value cutoff.
Investigate an SRM alert in data-path order
- Verify the comparison. Confirm the configured allocation for the relevant period, eligibility rules, analyzed unit, and count definition. Check whether the result is a persistent trend or a transient fluctuation rather than treating a single alert as a diagnosis.
- Inspect assignment. Look for incorrect bucketing, null or faulty IDs, identity churn, collisions, overlapping experiments, manual overrides, or a ramp that differs from the allocation used in the check. Microsoft Research identifies incorrect bucketing, faulty IDs, and carry-over effects among possible causes.
- Inspect execution and exposure. Check whether the treatment redirects behavior or changes who remains observable, and whether a client crash or rendering failure prevents exposure logging in one arm.
- Inspect logging and processing. Compare event loss, truncation, duplicates, joins, and inclusion windows across arms. A logging or processing difference can make a correctly assigned test appear imbalanced.
- Inspect analysis choices. Review filters, segment definitions, and any conditioning on behavior that occurred after assignment. These choices can select units differently between arms.
- Localize the imbalance. Examine counts and time trends across recorded dimensions such as platform, operating system or browser, SDK version, region, or bot status. Statsig documents these as diagnostic dimensions; a concentrated imbalance can help narrow the investigation, but does not alone establish the cause.
Decide what to do before using the result
If you find and fix a cause, decide whether the affected data can be defended or whether the experiment needs a clean restart. Statsig recommends investigation and commonly restarting after a fix; it notes that excluding a segment may sometimes be considered when the problem is clearly isolated. Exclusion changes the population the result describes, so document why it is justified and which units remain in scope.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
If the cause is unresolved, do not treat the effect estimate as decision-ready: Microsoft PlayFab guidance says analyses with unresolved SRM should not be used to make decisions. At the same time, report the alert as a reason to investigate rather than as proof that the treatment failed; Optimizely cautions that imbalance alone does not automatically make an experiment unusable.
When stratification may help
Stratification balances groups on selected characteristics before assignment. It may be worth considering for low-volume or high-variance experiments, such as B2B tests where a few large accounts can dominate a metric. Statsig says standard random assignment generally suffices for large consumer populations. Its documentation reports around 50% lower variance in its own simulations for the described setting; this is not an independent benchmark or a general guarantee.
Rank #4
Stratification adds setup and compute work, and allocating a smaller share to a test can reintroduce imbalance. It is an optional design choice, not a substitute for stable identity, reliable exposure measurement, or SRM diagnosis.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Further reading
For a broader treatment of experiment reliability, see Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing by Ron Kohavi, Diane Tang, and Ya Xu (Cambridge University Press, 2020). Its contents include a dedicated chapter, “Sample Ratio Mismatch and Other Trust-Related Guardrail Metrics.”
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




