Variance Reduction (CUPED)
Variance reduction, also called CUPED, uses each visitor's behaviour before the experiment started to shrink noise in your results. It's a step A vs B runs before the statistics engine. On revenue, pageview, and engagement metrics, it can cut the traffic you need for a significant result by 30 to 50%. That means faster tests and tighter intervals, without changing the answer.
A vs B lets you run an A/B test two ways. You can build an experiment in the Web Experimentation builder. Or you can use a feature-flag rule of type A/B Test, served through the @avsbhq/browser and @avsbhq/node SDKs. Variance reduction works the same way in both. The same Auto/Off setting and the same CUPED math apply. The same adjustment label shows up on the experiment Results page and on the Flag Rule Results modal.
What CUPED actually does
Most of the noise in an experiment comes from visitors just being different from each other. Some buy a lot, some never buy. Some browse for an hour, some leave in 30 seconds.
CUPED uses each visitor's behaviour before the experiment as a baseline, and subtracts that baseline from what you observe during the experiment. What's left is closer to the true effect of your variation, with most of the visitor-to-visitor noise removed.
The math is a single linear adjustment per visitor:
adjusted = during − θ × (pre − mean_pre)Here θ (theta) is the optimal weight for the pre-period value. A vs B derives it from the covariance between your visitors' pre-period and during-period values. It computes θ automatically whenever a metric is eligible.
Auto vs Off
Variance reduction is one setting with two options. New projects default to Auto.
Auto (recommended)
A vs B applies CUPED only when it will actually help. For each metric on the experiment, A vs B checks the pre-experiment data and decides whether the adjustment is worth running. When the data isn't there, A vs B skips the adjustment and reports the raw numbers instead: nothing is distorted.
Auto looks at four things before applying CUPED to a metric:
- Metric type. Continuous and count metrics (pageviews, custom numeric events, revenue) are eligible. Binary conversion-rate metrics (clicks) are out of scope for V1.
- Pre-experiment history. The project needs at least 14 days of tracking data before the experiment started.
- Sample size. At least 500 visitors with a pre-experiment value for the metric.
- Correlation. The pre-period values must be at least weakly correlated with the during-period values for the adjustment to do anything useful.
A composite metric combines several other metrics into one, weighted number. Under Auto, A vs B treats that combined number as a single per-visitor value and runs the same CUPED adjustment on it, using each visitor's combined value from the 30 days before the experiment as the baseline. The same eligibility checks apply: enough pre-experiment history, enough visitors with a pre-experiment value, and a real correlation. When the pre-experiment data isn't there, the composite simply shows its raw numbers, exactly like any other skipped metric.
Ratio and percentile metrics are different: variance reduction does not apply to them at all. Their rows say so directly, in a note on hover, so a skipped-for-data metric and a does-not-apply metric can't be confused.
Off
A vs B always reports the raw observed numbers with no adjustment. Use this when you want a baseline you can reproduce by hand. It suits auditing, debugging, and regulated workflows, where every step of the calculation must stay visible.
Auto only reduces variance: it never inflates it. When the data isn't there to support the adjustment, Auto declines and you get the same numbers as Off. There is no scenario where Auto produces a worse answer than Off.
What you'll see on the Results page
Every Results page shows a small label near the top summarising the analysis configuration. The label combines the stats engine and the variance-reduction decision into one line, formatted as <Engine> · <variance-reduction state>:
- The engine is Bayesian, Frequentist, or Sequential, depending on what the experiment was launched with.
- The variance-reduction state is one of:
- Off: variance reduction is turned off; numbers are raw.
- Auto (CUPED applied): CUPED ran on every metric; numbers are variance-reduced.
- Auto (no adjustment: reason): CUPED was eligible but the data didn't support it, or the metric can't support CUPED at all (ratio and percentile metrics land here). The reason is shown in plain English.
- Auto (CUPED applied to some metrics): mixed. CUPED ran on some metrics and was skipped on others.
Examples: Bayesian · Off, Frequentist · Auto (CUPED applied), Sequential · Auto (no adjustment: pre-period correlation too low).
On a mixed result, check each metric's own row rather than the top label. Hovering a metric's lift number shows a tooltip when CUPED actually ran on it, and ratio and percentile metrics show a note explaining that variance reduction does not apply to them. A metric that was skipped for lack of data still shows just the plain lift number; the top label carries its reason.
Where to configure variance reduction
Variance reduction can be configured in three places:
- Project level: set the project-wide default in Project Settings → Analysis. Every new experiment and every new feature-flag A/B test rule in the project picks up that default.
- Experiment level: override the project default for a single experiment on the Analysis step before launching. Once the experiment is running, the choice is locked: mixing analyses mid-flight invalidates the result.
- Flag rule level: override the project default on a single feature-flag A/B test rule from the rule's Analysis section. Once the rule has ever been enabled, the choice is locked for the same reason.
What's not in V1
- Ratio and percentile metrics. Variance reduction does not apply to these metric types. Their rows disclose it directly. See the callout above.
- Binary conversion-rate metrics. CUPED applies to continuous and count metrics in V1. Click-conversion metrics decline with reason this metric type is not yet supported.
- Multi-covariate adjustment (ANCOVA). V1 uses a single covariate: the pre-experiment value of the same metric.
- Custom pre-period windows. V1 uses a fixed 30-day pre-experiment window.
- Per-metric variance-reduction toggles. The setting applies experiment-wide; Auto then makes a per-metric decision automatically.