Sample Ratio Mismatch

A sample ratio mismatch (SRM) is a red flag: it fires when the traffic split you actually see does not match what you configured, by more than chance would explain, which usually means a bug rather than a real result. Say your experiment is set to 50/50. If you actually see 60% of visitors in control and 40% in the variant (one specific version being tested, control or a challenger), something is skewing who ends up where. When that happens, the experiment's data cannot be trusted.

The SRM guardrail

A guardrail is a check A vs B runs to catch a broken experiment, not a number you are trying to improve. A vs B shows an SRM guardrail warning on the results page when your traffic split is off by more than random chance would explain. The warning has two severity levels:

  • Yellow warning: the split is noticeably off but not extreme. Proceed with caution and investigate.

  • Red warning: the split is severely imbalanced. The results cannot be trusted. Stop the experiment and fix the underlying problem before continuing.

  • Red means the split is severely imbalanced: stop the experiment and fix the cause before continuing.

  • This line updates with the specific reason, including a paused-variation delivery failure.

The guardrail runs a chi-squared test, a standard statistical check for whether an observed split could plausibly come from your configured split. The result is a p-value: the probability of seeing a split this skewed by chance alone, if nothing were actually wrong. A p-value below 0.01 triggers the yellow warning. Below 0.001 triggers the red. This one test covers any number of variations. A badly broken split on a 3- or 4-variation test is caught exactly as reliably as on a simple A/B. A red SRM also fires the experiment.srm_failed alert to your connected Slack or Teams channels.

Paused (0%) variations

A variation whose weight is set to 0% is excluded from the ratio test. It is not supposed to receive traffic, so it contributes no expectation either way. But if visitors DO arrive on a 0% variation, something is badly wrong: delivery is not following your configured split at all. The guardrail goes straight to red, and tells you how many visitors landed on the paused variation. This happens regardless of how healthy the rest of the split looks.

The measurement window restarts when you change the split

The SRM check measures from your most recent traffic-split change, and editing weights restarts it. A visitor counts as "exposed" the moment they are first counted in the experiment: usually the moment they see the part being tested. Visitors first exposed under the old weights are excluded from the ratio test. Only visitors first bucketed since the change are compared against the current split. Right after a change, the window is small, so the check has little power until new traffic builds up. That is expected, not a malfunction. For feature-flag rules, any edit to the rule restarts the window the same way.

Why SRM matters

If your traffic split is skewed, the two groups are not equivalent to begin with. Any difference in conversion rates between them might be caused by the variation. Or it might be caused by one group having a different mix of visitors than the other. You cannot tell which. An SRM makes the experiment's results meaningless: you cannot conclude the variation caused anything you observed.

Common causes

Variation code crashes for some visitors

If your variant has a JavaScript error that crashes for some visitors, the snippet catches the error and deactivates the experiment for that visitor. Those visitors then drop out of the experiment. This disproportionately removes visitors from the variant, so its count falls relative to control.

How to investigate: Open DevTools → Console on the target page and force yourself into the variant:

JavaScript
avsb.forceVariation(42, 1487); // Experiment ID, then Variation ID
JavaScript1 line

Look for any JavaScript errors in the console after the variation is injected. Any error here could be the cause of the SRM.

Bot traffic affecting one variation more than others

Web crawlers, monitoring bots, and automated scripts can end up consistently bucketed into one variation. This inflates that variation's visitor count without producing real human conversions. Most bots do not store cookies. So they can be assigned a new variation on every visit, and land disproportionately in certain buckets depending on how they cycle through.

How to investigate: Look at the visitors per variation over time. Bot traffic often shows up as spikes of visits with zero conversions during off-hours.

Redirecting visitors away from a variation

Say your variation includes a redirect: it sends visitors to a different URL right after they are bucketed. The exposure event can fire before that redirect happens, but the redirect then stops the visitor from ever seeing the actual variation experience. This can cause SRM, because some visitors are counted as exposed but never actually took part.

Fix: If you are testing a redirect, structure the experiment so the exposure event fires only after the redirect completes. Or use a redirect experiment type.

Say visitors' cookies are being cleared frequently: strict browser privacy settings, cookie-blocking policies, or a short cookie expiry can all do this. The same visitor may then get bucketed more than once. Each bucketing is random and independent, so over many re-bucketing events the distribution should even out. But in the short term, this can cause SRM.

How to investigate: Check whether the SRM correlates with a particular browser (Safari with ITP, Firefox with Enhanced Tracking Protection) or a particular traffic segment.

How to investigate any SRM

  1. Compare visitor counts per variation: look at the raw numbers on the results page. Is the imbalance consistent over time, or did it spike at a specific point?
  2. Check for JavaScript errors: open DevTools and force yourself into each variation. Look for any console errors.
  3. Look for patterns in the excess traffic: if nearly all of it is in the control, that suggests the variant is crashing or being deactivated for some visitors. If it is spread randomly, it may be a bucketing or cookie issue instead.
  4. Check for deployment timing: did the SRM appear right after you changed the variation code? A code deploy that crashes one variation shows up as a mismatch that starts at the deploy time.
Changing traffic splits mid-experiment restarts the SRM window

Weight changes on a running experiment restart the SRM measurement window. Only visitors first bucketed since the change are tested against the new percentages. So a legitimate ramp, say 10/90 to 50/50, does not raise a false alarm from mixing two different distributions. Two caveats remain. The check has low power until the new window builds up traffic. And your RESULTS still mix visitors from both periods, so a mid-experiment ramp can bias measured lift even when delivery is healthy (a pattern called Simpson's paradox). Prefer setting your split once. If you do ramp, treat the pre-ramp data with care.

Do not ignore SRM warnings

It can be tempting to proceed with an experiment despite an SRM warning, especially if the results look good. Resist this temptation. An SRM means the experimental groups are not equivalent, so the results might be a statistical artefact rather than a real effect. Fix the underlying problem first, then re-run the experiment.

Was this helpful?