Health Guardrails
Health guardrails are automatic checks A vs B runs on your experiment data to flag potential problems. They appear on the Results page as three cards with green, yellow, or red indicators. A red guardrail does not mean your experiment is broken: it means you should investigate before trusting the results.
The same three checks run on feature-flag rule results, using the same code. A flag rule and an experiment with identical data always get identical verdicts.
The three guardrails
Sample Ratio Mismatch (SRM)
Checks whether visitors actually landed in your variations in the proportions you configured.
A vs B runs a chi-square test comparing the observed visitor counts against your configured traffic split, and reports the resulting p-value. The p-value answers: "if delivery were working correctly, how likely is a split at least this lopsided, purely by chance?" A small p-value means chance is an implausible explanation: something is systematically skewing the split.
| Status | Condition | Meaning |
|---|---|---|
| Green | p ≥ 0.01 | No mismatch detected. |
| Yellow | 0.001 ≤ p < 0.01 | Possible mismatch. Monitor closely. |
| Red | p < 0.001 | Significant mismatch. Results may be unreliable. |
These thresholds match the industry standard used by LaunchDarkly, GrowthBook, Statsig, and Eppo.
The test scales with your experiment: two-variation and 3+-variation experiments run one shared implementation, so a badly broken split on a 4-way test is caught exactly as reliably as on an A/B.
There are two special cases that report red without reference to the p-value:
- No traffic yet. Before any visitor is exposed, there is nothing to test, so the card reads "Insufficient data: no visitors yet, so sample-ratio mismatch cannot be assessed." This is deliberately red rather than green: an empty experiment has not passed the check, it simply has not taken it.
- Traffic on a paused variation. A variation set to
0%is excluded from the ratio test: it should receive no traffic, so it contributes no expectation. But if visitors do arrive on a 0% variation, delivery is not following your split at all. The card goes straight to red and reports how many visitors landed there, however healthy the rest of the split looks.
The SRM check measures from your most recent traffic-split change. Only visitors first bucketed since the change are compared against the current percentages, so a legitimate ramp (10/90 → 50/50) does not raise a false alarm by mixing two different expected distributions. Right after a change the window is small and the check has little power: that is expected, not a malfunction. For feature-flag rules, any edit to the rule restarts the window.
A Sample Ratio Mismatch means the data is biased in some way. The results may show an apparent winner, but that winner could be an artifact of the mismatch rather than a genuine effect. Always investigate and resolve the root cause before drawing conclusions. See Sample Ratio Mismatch troubleshooting for common causes and how to investigate.
A red SRM also fires the experiment.srm_failed alert to your connected Slack, Teams, Jira, and webhook destinations. The "no traffic yet" red is the one exception: it never alerts, because "we have no data" is not an incident worth waking anyone for.
Statistical Confidence
Checks whether your results point strongly enough in any direction to act on. What counts as "strong enough" depends on the stats engine your experiment uses, so this card reads differently across the three engines.
Bayesian: reads the highest probability-to-beat-control across your variations, against your configured confidence level:
| Status | Condition (at the default 95% confidence) |
|---|---|
| Green | Probability ≥ 95%: the threshold is met. |
| Yellow | Probability ≥ 80%: trending, not yet there. |
| Red | Probability below 80%: collect more data. |
The green threshold is your confidence level (1 − α), so lowering confidence to 90% moves the green line to 90% and the yellow line to 75%. The yellow band is always the 15 points below the threshold.
Frequentist: reads the lowest p-value across your non-control variations, against your significance level α:
| Status | Condition (at the default α = 0.05) |
|---|---|
| Green | p < α: significant. |
| Yellow | p < 2α (below 0.10): close to significance. |
| Red | p ≥ 2α, or no variation has a p-value yet. |
Sequential: reads the always-valid p-value and whether the confidence sequence excludes zero:
| Status | Condition |
|---|---|
| Green | The confidence sequence excludes zero: safe to stop. |
| Yellow | Always-valid p < 2α: not yet safe to stop. |
| Red | Neither, or no always-valid p-value yet. |
Low confidence is normal early in an experiment. It does not mean something is wrong: it means keep running.
Traffic Health
Checks the sample size of your smallest variation. Small samples make conversion rates volatile: a single conversion can move the rate by percentage points.
Naming note: the notification for this check is labelled "Traffic guardrail breached (safety)". It is about traffic delivery, not metric movement: the separate guardrail metrics feature tests whether the metrics you can't afford to hurt stayed inside their safety margin, and has its own alert.
| Status | Smallest variation has |
|---|---|
| Green | 1,000+ visitors: sufficient sample size. |
| Yellow | 100–999 visitors: low, results may be noisy. |
| Red | Under 100 visitors: insufficient traffic. |
Unlike the SRM check, this one always counts all-time visitors, not just those since your last split change.
Green here means "enough data to not be pure noise", not "enough data to decide". The sample size you actually need depends on the effect you are trying to detect; use the sample-size calculator to plan it, and the Days Remaining card to track progress.
Green, yellow, red: what to do
All green: everything looks healthy. Read the results.
Yellow: the experiment is in a transitional state. Keep running and check back in a few days.
Red SRM: pause and investigate the cause. Do not ship a result based on mismatched data. The exception is the "no visitors yet" red, which clears itself the moment traffic arrives.
Red confidence or traffic health: keep running. Both resolve naturally as traffic accumulates. There is nothing to fix unless traffic is unusually low, in which case check that your experiment is actually running and your targeting is not too narrow.
The peek counter
Below the health guardrails, running experiments show a peek counter: how many times your team has checked these results before the experiment reached its planned finish.
This is not a telling-off. Checking is normal and often necessary. The counter exists because, under Frequentist analysis, how often you looked changes what your results actually mean, and that is a fact you deserve to see rather than discover later.
Why looking changes the answer
A Frequentist p-value assumes one single, pre-planned look at your planned sample size. Every extra look is another chance for random noise to cross the significance line. Check often enough and stop at the first "win", and something will eventually look significant even when both variations are truly identical.
So the counter shows your real false-alarm rate given the number of looks: not the 5% you configured. It can only ever rise as you check more. If a number here ever went down as you kept looking, it would be lying to you.
What counts as a peek
A peek is recorded when all of these are true:
- You opened the Results page in the dashboard and it showed real numbers.
- The experiment was running at the time.
- It had not yet reached its planned sample size or its scheduled end date.
Looking under a different engine (via Explore under) still counts: you saw the numbers either way.
What does not count
- Refreshing the page. Looks are counted once per person per day (UTC). Refreshing fifty times in a minute is one decision opportunity, not fifty. Three teammates each looking on the same day is three.
- The experiments list and dashboard tiles. Those show cached summaries, not fresh results.
- The compare-engines modal. It is part of the same page view you were already counted for.
- API reads. Results fetched with a service token or a personal access token are not counted: there is no person behind them, and an integration polling on a schedule is not somebody making repeated decisions.
- The daily digest email or Slack message. That is pushed to you; you did not go looking.
- Anything from before this feature existed. Counting starts at the first recorded look. Earlier looks were never recorded anywhere, and A vs B will not invent them.
The number includes the look you are taking
The count always includes your current view. Open a running experiment for the first time today and it reads 1 look, before anything has been written down. That is deliberate: the number on screen is the number your next decision would be based on.
Viewing results never changes your numbers: A vs B does not quietly recalculate anything because you looked. But your checks are recorded, in their own audit log, and they feed your experiment's trust grade. We would rather tell you that plainly than count you in silence.
About the risk figure
The inflated false-alarm percentage is labelled an illustrative estimate, and it means it. It comes from published tables that assume your looks were evenly spaced (which real checking never quite is), and those tables are anchored at a 5% base rate. See Statistical Methodology for the exact model and its limits. Treat it as a directionally honest warning, not a precise measurement.
What to do about it
- Under Sequential analysis, nothing. Sequential stays valid at every look: that is the entire point of it. The counter still appears, but only as a record for your trust grade.
- Under Bayesian analysis, peek inflation does not apply the same way; a Bayesian threshold is not a hard error-rate guarantee either way. The count still feeds your trust grade.
- Under Frequentist analysis, the honest options are to let the experiment reach its planned sample before acting, or to switch future experiments to Sequential. If you do stop early, do it knowingly: the stop dialog asks you to log why.