Winsorization

Picture a checkout experiment. Revenue per visitor over the last 30 days looks like this:

Plain text
Mean:    $42Median:  $28p95:     $180p99:     $410Max:     $5,820  (one purchase from a corporate restock)
Plain text5 lines

Most visitors spent nothing, or a normal amount. One visitor placed a single $5,820 order. That one value pulls the mean up. It also makes the spread of the values, called variance, far bigger than a typical visitor's spending would suggest. Now cap that one value down to $410, the 99th percentile, before averaging:

Plain text
Mean:    $39  (down $3, about 7%)Max:     $410 (was $5,820)Variance: down about 22%
Plain text3 lines

Nobody was removed from the test. That visitor is still counted. Only the value used in the maths changed. Capping the most extreme values at a percentile, before they're used in the maths, is called winsorization. It's the standard fix for long-tailed metrics like revenue, time-on-site, and session duration. On these metrics, one huge value can swing the result far more than its frequency deserves.

When to use it

Reach for winsorization whenever a metric has a long right tail. That's a sign a few extreme values are pushing your variance higher than they should:

  • Revenue per visitor (B2C). Most purchases are $20 to $100, but the occasional $5,000 purchase happens. Left uncapped, a few high-value orders drive up the variance. Your confidence interval (the range your true result is likely to sit inside) ends up wider than it needs to be.
  • Revenue per visitor (B2B). Even more extreme: most contracts are $10k to $50k, but one $1M deal can shift the mean tenfold or more. Capping at p95 or p99 keeps that one deal from taking over the test.
  • Session duration. A visitor who leaves a tab open overnight pushes the tail into hundreds of minutes. Capping at p99 typically brings these down to around 30 minutes, closer to what real engagement looks like.
  • Items in cart. Most people add 1 to 3 items. A few add 47, often automation or a restock. Winsorization keeps those carts in the test without letting them define the average.
When NOT to winsorize

If your metric is binary (click rate, conversion rate, signup rate), there's nothing extreme to cap. Winsorization doesn't apply. Sometimes a metric's extreme values are exactly what you care about, for example revenue from your biggest-spending customers in a high-end product test. Capping those away would defeat the purpose. Rule of thumb: winsorize when you suspect the outliers are noise. Don't, when they're the signal you're measuring.

Which measures support it

Winsorization is available on every measure that produces a continuous or numeric value: Total value per visitor, Total value, Rate, Percentile, and Composite (weighted). Conversion-rate measures (Unique conversions per visitor, Total events, and Unique visitors who fired) are binary or count-based. There's no continuous value to cap, so the option doesn't appear for them. On a Composite (weighted) metric, each component is winsorized on its own, before the weighted sum is computed.

You configure winsorization on a metric's per-measure settings inside the experiment builder. You don't set it when you first create the metric. That means the same metric can be winsorized in one experiment and left alone in another.

Choosing where the cap comes from

Once winsorization is on, you also choose where the cap is computed:

  • Shared cap across all variations (recommended, the default). A vs B computes one cap from all of the experiment's traffic and applies it to every variation. Every variation is capped at the same cutoff, so the comparison between them stays fair. Your chosen confidence level keeps its usual false-alarm guarantee.
  • Per-variation caps (expert). Each variation gets its own cap, computed only from its own visitors. This adapts to each variation's own spread of values. But the variations end up capped at different cutoffs. On a skewed metric, that difference alone can look like a lift, and it raises your false-positive rate. Use this only when you have a specific reason to.

The percentiles and the cap scope you choose are saved exactly as you entered them, and reopening the measure settings shows you the same values back. If you set a lower percentile that sits at or above the upper percentile, the save is refused and the message names the rule: the lower percentile has to be strictly below the upper one.

A vs B computes every cap from the metric's non-zero values only. A visitor who never fired the event has a value of zero. Zeros don't drag the cap down, and they never get clamped themselves: a zero means "nothing happened," not a small observation. That also means a mostly-zero metric can't be flattened by its own zeros. If a variation, or the whole experiment, has fewer than two non-zero values, A vs B computes no cap and clamps nothing.

Winsorization is set per experiment, not on the metric itself. The same metric can be winsorized in one experiment and left alone in another.
  1. The Enable winsorization toggle.
  2. The cap scope choice between a shared cap and per-variation caps.
  3. The Lower percentile and Upper percentile inputs.

How A vs B computes the cap

A vs B sorts the metric's non-zero per-visitor values from that experiment's own traffic. It never reaches for historical data. Using older data, or data from outside the experiment, would leak information the statistics depend on. It then finds the percentile you chose, estimating between the two closest points. This is the same method NumPy's default quantile function uses. Every value above that cap is brought down to it, before A vs B computes the mean, variance, lift, and confidence interval. Results also record the cap and how many values it changed, so nothing about the adjustment is hidden.

Winsorization always runs before CUPED, a technique that uses a visitor's behaviour from before the experiment to reduce noise in the result. Because winsorization runs first, CUPED's own noise reduction works on the already-capped values, not the raw ones:

Plain text
raw values  →  winsorize  →  CUPED  →  variance / lift / CI
Plain text1 line

What you'll see in results

Winsorization's effect shows up once the experiment has real data, on the Results page. A variation with at least one capped value gets a neutral Outliers capped badge next to its measure. A note underneath the numbers spells out exactly how many values were capped in each variation, like "3 outlier value(s) capped in Variation B". There's no before-you-save preview while you're still deciding whether to turn winsorization on. The fastest way to see the effect is to turn it on, then read the capped count on the next results refresh.

Winsorization's effect becomes visible on the Results page once the experiment has data: A vs B does not show a before-you-save preview.

Sealed-plan amendments

A sealed analysis plan is an experiment's locked-in statistical plan: once sealed, it changes only through a logged amendment. Turning winsorization on or off, or changing the cap percentile, on a metric a sealed plan references is refused outright: the save is blocked, and the error names the experiments whose plans reference the metric. Amend each plan first, through the experiment's plan card, then make the metric change. The refusal itself is recorded in the audit log.

FAQ

What percentile should I use?

p99 is the most common default: it caps the worst 1% of values. p95 cuts more, and is sometimes the right call for a very long-tailed metric. Go much below p95, and you change so much of the data that the mean stops being a true picture of it. Turn winsorization on, then check the capped-values note after your next results refresh. If only a handful of values get capped and the mean barely moves, the cap is doing its job without hiding real signal.

Does winsorization interact with CUPED?

Yes. Winsorization runs first, then CUPED applies to the already-winsorized values. This is the standard order across the experimentation industry. Cleaning up outliers before reducing noise from pre-experiment behaviour means the noise-reduction step works on a clean distribution, not a raw one.

Can I change winsorization on a running experiment?

Only if no sealed plan references the metric. Once a plan referencing it is sealed (which happens at launch), the change is refused until you amend that plan through the experiment's plan card, as described above. On a metric no sealed plan references, the change saves normally.

When should I use the lower cap?

Rarely. The lower cap is for a metric with a negative tail, like a net-revenue metric where refunds push some visitor values below zero. Most metrics you'll test never go negative, so they only need the upper cap.

Was this helpful?