Guardrail Metrics
A guardrail metric is a metric you aren't trying to improve, but can't afford to make worse: a number you protect, not one you're chasing. Revenue while you test a new navigation. Page speed while you add a widget. Bounce rate while you tighten a signup form.
A vs B runs a proper "is this safe?" test on every guardrail metric, not just an "is this different?" test. That distinction matters in both directions:
- A tiny dip that is statistically significant (unlikely to be random chance, though still small in size) can stay inside your margin and count as Safe. You are not over-alarmed by noise-level movement on a high-traffic metric.
- "No news" is never blessed as safe. Until the data actually rules out harm bigger than your margin, the metric stays Not yet established safe. A wide, uninformative result cannot pass the check by default.
A worked example
Say Revenue per visitor is a guardrail, with the default 2% margin, and Control converts at $50 per visitor on average. Your allowed margin is 2% of $50, or $1. A vs B will call the challenger Safe only once the data rules out a drop bigger than that $1, at your experiment's confidence level. A challenger that is down $0.40 on average, with an interval that stays above −$1, is Safe: that dip is well inside what you said you'd tolerate. A challenger whose worst plausible outcome is still a $1.50 drop is Breached, even if the point estimate looks small, because the interval never rules out real harm.
Setting a metric as a guardrail
Give a metric the Guardrail role when you attach it to an experiment. This is the same field as listing it under Guardrail metrics in the analysis plan: the plan for what you'll measure and how you'll decide, agreed before launch. The two views read the same underlying data, so they can never disagree. A metric can never be both primary and guardrail.
The margin
Every guardrail has a safety margin: the relative worsening you are willing to tolerate. The default is 2% of the control's value. You can change it per project in settings, or per metric on the experiment.
When your analysis plan is sealed at launch, the margins are frozen into the plan. Editing a margin afterward never rewrites how a running experiment is judged. That's the same guarantee that protects your significance level (how strict your test is about calling a result real).
The three states
For each guardrail metric and each variation (each version you're testing), A vs B computes a confidence bound: a range for the true difference from control, built so "worse" always means worse for that metric's own goal. An increase counts as worse for bounce rate; a decrease counts as worse for revenue. A vs B compares that range against the margin:
- Safe: the data rules out, at your experiment's confidence level, any harm worse than the margin. Formally, the lower bound of the range sits above −margin.
- Breached: the data establishes a regression bigger than the margin. Even the most optimistic end of the range is worse than −margin. You'll also get a safety alert (see below).
- Not yet established safe (inconclusive): the range still straddles the margin. Keep collecting. This state also covers metrics with too little data, and metrics whose control value is zero (a relative margin is undefined there).
With more than one non-control variation, a guardrail metric's overall status is the worst status across all of them: one Breached variation is enough to mark the whole metric Breached, even if the others are Safe.
Under the Sequential engine (the A vs B stats engine built to stay valid even when you check results daily and stop early), the bound stays valid no matter how often you check. So guardrail statuses are peek-safe: checking them daily never inflates your false-alarm rate.
Why guardrails are one-sided and uncorrected
Guardrails are deliberately excluded from multiple-comparison correction (an adjustment that keeps your false-positive rate honest when you're checking more than one metric at once). Correction protects you from cheering a false positive. On a safety test it would do the opposite: make real harm harder to detect. Each guardrail is tested one-sided against its own margin, at your experiment's configured confidence level. Its p-value (the odds of seeing a result this extreme if nothing actually changed) never enters the correction family used for secondary metrics (metrics tracked for context, without deciding the result on their own).
Effect on the verdict
If any guardrail is Breached, the experiment's verdict can never be "Ship," whatever your primary metric (the one metric an experiment is actually judged on) says. The Decision Header reflects this, and the guardrail panel underneath names the metric and the bound that established the breach.
Alerts
When a guardrail metric first flips to Breached, A vs B sends the "Guardrail metric breached (safety)" notification. It goes to your configured destinations: Slack, Teams, Jira, or a webhook (an automatic message A vs B sends to your own server). Safety events skip your organization's keyword gate. Each experiment alerts at most once per 24 hours, so a refresh cadence can't spam a channel.
This is a different alert from "Traffic guardrail breached (safety)", which fires on traffic-health problems (a starved variation or a dead snippet), not on metric movement.
When a guardrail is the wrong tool
Guardrails are a floor, not a decision. Don't put every secondary metric under Guardrail just to watch it: each one is a live operational check with its own alert, and too many will train your team to ignore them. Reach for a guardrail specifically when hurting that metric would be bad enough that you'd want to know within a day, not at the end of the experiment. A metric you're actively trying to improve belongs as your primary metric, not a guardrail: a metric can't be both, and a guardrail's test is built to detect harm, not to declare a win.