Analysis Defaults
The Analysis tab in Project Settings controls how A vs B analyses every A/B test in this project. That test might be a Web Experimentation experiment, or a feature-flag rule of type A/B Test. The tab holds two cards. The first is "Analysis defaults." It sets your defaults for every engine: which one runs, and whether variance reduction (a technique that shrinks a metric's natural noise, using each visitor's own pre-experiment data) applies. The card also holds each engine's own dials, and a safety margin for guardrail (a metric you don't want a change to make worse) metrics. The second card is Pre-registration (the requirement to declare your analysis plan before you can launch), which you turn on or off.
| Setting | Controls |
|---|---|
| Stats Engine | Which statistical method analyses your results. |
| Variance Reduction | Whether CUPED (the technique behind variance reduction) strips predictable noise before the engine runs. |
| Bayesian settings | The decision threshold, the ROPE band, and the effect prior (your starting assumption about how big a real effect usually is). Used only on the Bayesian engine (the engine that reports a probability of beating Control). |
| Frequentist settings | The confidence level (α) and the multiple-comparison correction (an adjustment that keeps your false-positive rate honest across several tests at once). Used only on the Frequentist engine (the classical significance-testing engine described below). |
| Sequential settings | The planned horizon. Used only on the Sequential engine (the engine built to stay valid even when you check results daily). |
| Guardrail defaults | The safety margin for guardrail (a metric you are not trying to improve, but do not want to make worse) metrics. Applies no matter which engine a test uses. |
| Pre-registration | Whether experiments must declare an analysis plan before launch. |
Every default here pre-fills every new experiment and flag rule. You can override any of them on the individual experiment or rule before it goes live. Once the experiment is running, or the rule is enabled, they lock.
- Stats Engine: choose Bayesian, Frequentist, or Sequential. Bayesian carries a "Default" badge.
- Variance Reduction: choose Auto or Off. Auto carries a "Recommended" badge.
Accessing the Analysis tab
Open your project, click the gear icon in the left sidebar to open Project Settings, then click the Analysis tab. The tab is available in both Web Experimentation and Feature Flag (a project type that manages flags instead of running on-site experiments) projects.
Stats Engine
The Stats Engine is the statistical method A vs B uses to decide whether a result is real. Three engines are available:
- Bayesian (default): reports the probability each variation (one specific version being tested, Control or a challenger) beats Control, along with credible intervals. Easy to read. Forgiving of mid-experiment peeks.
- Frequentist: classical p-values (the chance of seeing a result this extreme if the variation actually changed nothing) and confidence intervals. Declares a winner only when the significance level you set (α) is met. Requires sticking to your pre-declared sample size.
- Sequential: always-valid inference. Peek as often as you like, stop as early as the evidence is in. Slightly wider intervals in exchange for that flexibility.
See Choosing a Stats Engine for a deeper comparison and guidance on when to pick each.
Variance Reduction
Variance reduction is a pre-processing step that runs before the engine. It subtracts each visitor's pre-experiment behaviour from their during-experiment value, which shrinks the noise in your results. On revenue and engagement metrics this typically reaches significance 30 to 50% sooner. There are two options:
- Auto (recommended): A vs B applies CUPED variance reduction whenever the metric has a meaningful pre-experiment signal. When the data isn't there, A vs B silently skips the adjustment, so the numbers match the raw values exactly.
- Off: A vs B always reports the raw observed numbers. Useful for auditing, debugging, or regulated workflows where every step of the calculation must be reproducible by hand.
See Variance Reduction (CUPED) for a deeper explanation of what the adjustment does, when Auto applies it, and what you'll see on the Results page.
Bayesian settings
These three controls are the project defaults for the Bayesian engine. A Frequentist or Sequential test ignores them.
- Bayesian settings: the decision threshold, plus the optional ROPE band and effect prior.
- Frequentist settings: the confidence level and the multiple-comparison correction.
- Sequential settings: the planned horizon.
- Guardrail defaults: the safety margin, which applies no matter which engine a test uses.
Default decision threshold
How sure A vs B must be before calling a winner: the chance a variation has to beat Control before it counts as significant. The default is 0.95, a 95% probability to beat Control. You can set anything from 0.50 to 0.999. Leave it empty to use the platform default. This single number powers the winner verdict, the results table, and the A/A guardrail.
The Bayesian decision threshold and the Frequentist confidence level are two different controls. They share the same 0.95 default, but changing one does not change the other.
Default ROPE band
Off by default. A ROPE (region of practical equivalence) is a band around zero. It is measured in the units of your primary metric (the one metric an experiment is actually judged on). Inside that band, a difference is too small to matter. Turn it on to declare a band you would not act on, even if the difference were real. A vs B seeds it at ±0.005 when you switch it on: a 0.5 percentage-point band, right for a conversion-rate metric. The band must bracket zero. A revenue metric usually needs a much wider one.
Leave it off and every difference counts, however small.
Default effect prior
Off by default. An effect prior tells the Bayesian engine what your tests usually produce. It pulls an early, noisy result toward what's realistic instead of taking it at face value. Turn it on and A vs B seeds two numbers: a mean of 0, meaning no assumed direction, and a spread of 0.30. That spread means most real effects land within about ±30%. Both numbers move together. You cannot set one without the other.
A badly chosen prior can hide a real win, so the Results page always says when a prior moved the answer. See Effect Priors for the full explanation, including what "shrinkage" means and why it protects small samples.
Frequentist settings: confidence and corrections
These two controls are the project defaults for the Frequentist engine. A Bayesian or Sequential test ignores them.
Default confidence level
Your confidence level determines the significance level: α = 1 − confidence. The default is 0.95 (α = 0.05), the industry standard. You can set anything from 0.80 to 0.99 in steps of 0.01, so α is bounded to the 0.01 to 0.20 range.
Raising confidence to 0.99 makes a winner harder to declare and needs more data. Dropping to 0.90 does the opposite. Pick it for statistical reasons before launch, never to rescue a result afterwards.
Default multiple-comparison correction
Testing several variations at once means running several tests. Each extra test is another chance at a false positive, so the correction method controls how A vs B compensates:
- Tiered (recommended): the default. Leaves your primary metric's comparison across variations uncorrected. Applies the Benjamini-Hochberg false-discovery-rate correction across your secondary metrics (metrics tracked for context, which do not decide the result on their own). Choose it when your primary was genuinely pre-declared, so you get full sensitivity there without a long secondary list manufacturing false positives.
- Bonferroni (most conservative): multiplies each p-value by the number of tests.
- Holm-Bonferroni: the same family, slightly less punishing.
- Benjamini-Hochberg (FDR): controls the false-discovery rate across every metric. The better pick when you have many metrics or variations and no single pre-declared primary.
- None: raw, uncorrected p-values. Only if you know exactly why.
This picker always shows on the Frequentist settings card. The correction only changes anything once a Frequentist experiment runs more than two variations (more than one challenger against Control). With a single challenger there is only one test, so there is nothing to correct for.
See the Frequentist engine for what each method does to your p-values.
Sequential settings
Sequential is the only engine with one further dial: the default planned horizon, your expected total number of visitors. A vs B uses it to size the "honest peek" overlay, the display that tells you whether it's still safe to stop early. Leave it empty and there is no project default. Each experiment can still declare its own horizon.
This setting has no effect on Bayesian or Frequentist tests. See the Sequential engine for how always-valid inference works.
Guardrail defaults
The default safety margin is shown as a percentage. It is how much worse than Control a guardrail metric may get before A vs B calls it breached. The default is 2%, and you can set anything above 0% and below 50%. Unlike the other defaults on this page, this one is never empty. Every project always has a concrete margin.
This default applies no matter which engine a test uses. Each experiment can still override it per metric. See Guardrail Metrics for what Safe, Breached, and Inconclusive mean.
Pre-registration
Turn on Require pre-registration before launch and every experiment in the project must declare and complete an analysis plan before it can launch. The plan locks at launch. Any later change requires a recorded amendment.
This is the setting that makes "we predicted this in advance" auditable rather than remembered. It matters most alongside the Tiered correction above, whose whole justification is that the primary metric was chosen before anyone saw the data.
Planning sample size
Open the Sample-Size Calculator from its own entry in the sidebar (Web Experimentation projects), or follow the callout near the bottom of the Analysis tab, to plan your experiment before launch. Opened from an experiment, it prefills daily traffic and baseline rate from that experiment's last 30 days of data, labelled as measured, so you start from real numbers instead of guesses. It has three modes:
- Fixed-horizon: required sample size per variation and estimated duration.
- Power calculator: achieved statistical power (the chance of detecting a real effect, if one truly exists) at a given sample size.
- Duration estimator: how long until the experiment reaches its target.
Each mode is engine-aware. Picking Bayesian, Frequentist, or Sequential changes both the inputs and the meaning of the output. Once you have a number you're happy with, click Use as target to write it straight onto the experiment's analysis plan. See the calculator guide for details on each mode.
Changing the project-level defaults only affects experiments and feature-flag A/B test rules created after the change. An already-running experiment or a launched rule keeps the analysis configuration it was launched with: switching mid-flight would invalidate the result.
Who can change defaults
Only team members with the Edit Project Settings permission can change the analysis defaults. By default this includes the Owner, Admin, and Developer roles.