Statistical Methodology

A vs B ships three statistical engines. You pick one per experiment (or set a project default), and it decides how your results are computed and what the Results page shows you:

EngineThe question it answersWhat you read
Bayesian (default)"How likely is it that B is better than control?"A probability to beat control, plus a credible interval.
Frequentist"If there were no real difference, how surprising is this data?"A p-value and a Significant badge, plus a confidence interval.
Sequential"Can I stop now without fooling myself?"An always-valid p-value and a confidence sequence: Safe to stop, or Inconclusive.

None of them is universally right. See Choosing a stats engine for how to pick, and Comparing engines for running the same data through all three side by side.

Everything below the engine sections (winsorization, CUPED, the delta method, bootstrap intervals) applies to whichever engine you choose. The engine decides how uncertainty is expressed. It does not change how the underlying numbers are prepared.

The Bayesian engine

Bayesian statistics answers one question: "Given the data collected so far, how confident am I that Variation B beats Control?" The answer is a probability, for example 92%. A probability like that is easy to act on.

An analogy: flipping a coin

Imagine you suspect a coin is biased toward heads. You flip it 20 times and get 13 heads. Is the coin biased?

A frequentist asks a narrower question: how surprising would 13 heads be, if the coin were actually fair? There is about a 13% chance of seeing 13 or more heads out of 20 fair flips. So at α = 0.05 you would not reject the null. "Not proven biased" is not the same thing as "fair".

A Bayesian asks how likely it is that the coin favours heads, given what you saw. Start from a flat prior (Beta(1, 1), "I know nothing"). Add 13 heads and 7 tails, and you get a Beta(14, 8) posterior for the coin's true heads-rate. Integrate that posterior above 0.5:

Plain text
P(heads-rate > 0.5)  =  ∫₀.₅¹ Beta(p; 14, 8) dp  ≈  90.5%
Plain text1 line

So there is about a 90.5% probability the coin favours heads. The two answers do not contradict each other. They answer different questions. The 13% describes the data, assuming the coin is fair. The 90.5% describes the coin, given the data you actually saw.

90.5% is not a licence to call it

A 90.5% probability still leaves nearly a 1-in-10 chance the coin is fair or tails-biased. The Bayesian number is easier to read, not automatically more certain.

Winning probability

The probability to beat control shown for each variation is the posterior probability that this variation has a higher true conversion rate than control. A vs B computes it with a Beta-Binomial model, the standard Bayesian approach for conversion-rate experiments.

As data arrives, the probability updates continuously. With very little data it hovers near 50% (no evidence either way). As one variation consistently converts better, it climbs toward 90%, 95%, 99%.

This number is computed exactly, not estimated by random simulation. Many tools approximate it by drawing thousands of random samples. That means the same data can show a slightly different probability every time the page loads. A vs B solves the underlying integral directly instead, using deterministic numerical integration (Gauss-Legendre quadrature) rather than random draws. The same data always produces the same number. Results never jitter between refreshes, and a verdict sitting right at a threshold cannot flicker on and off.

Credible intervals

A credible interval is the Bayesian counterpart of a confidence interval, with a more direct reading. A 95% credible interval of [+5%, +15%] means: "there is a 95% probability the true improvement lies between 5% and 15%." You can read it literally. A frequentist confidence interval cannot be read that way. It is a statement about the long-run behaviour of the procedure, not about this one result.

Narrower intervals mean more certainty. A wide interval (say [−5%, +25%]) means you do not yet have enough data to know the effect size, whatever the point estimate says.

Like the winning probability, credible intervals for conversion rates are exact calculations from the posterior. They are deterministic, with no random sampling involved.

The prior

Bayesian statistics needs a starting assumption, called a prior. By default A vs B uses a non-informative prior (a flat Beta distribution). The starting assumption is "we know nothing", so results are driven entirely by your data, with no thumb on the scale.

You can optionally replace that with an effect prior. This tells the analysis what effects your tests usually produce (for example "most real effects land within ±30%"), so early flukes get pulled toward reality. When an effect prior is active, the Bayesian probability, the lift estimate, and its interval all come from a posterior on the relative lift instead of the flat Beta model. The results page always says so. See Effect priors for what it does, when to use it, and the honest warning that comes with it.

The Frequentist engine

The Frequentist engine reports a p-value: the probability of seeing a result at least this extreme if there were no real difference. Below your significance level α (0.05 by default) the result is marked Significant.

A p-value is not the probability that your variation works. That is a common and costly misreading. It is a statement about the data under an assumption of no effect.

Pick Frequentist when you need results in the vocabulary your organization already audits in, or when a stakeholder expects "p < 0.05". See the Frequentist engine, including how A vs B corrects for testing multiple variations and metrics.

The Sequential engine

The Sequential engine reports an always-valid p-value and a confidence sequence. A confidence sequence is an interval that is valid simultaneously at every sample size, not just at one planned endpoint. That is what makes it safe to look whenever you like: the guarantee holds at every look, by construction.

The cost is width. A confidence sequence is deliberately wider than a fixed-horizon interval at the same sample size. That extra width is what pays for unlimited peeking. Sequential needs more data to call the same effect.

See the Sequential engine.

The interval on relative lift

Every comparison also carries an interval on the relative lift itself: the range of plausible values for "how much better (or worse) than control, in percent". Each engine computes it with the method that matches its own guarantees:

  • Bayesian: the posterior credible interval on the relative difference. For conversion rates it comes from the same Beta posteriors as the winning probability; for continuous metrics, from the normal posteriors of the two means. Deterministic either way, with no random sampling.
  • Frequentist: a delta-method interval while the control mean is precisely measured, switching automatically to Fieller's method when the control is noisy. A plain symmetric interval on a ratio becomes dishonestly narrow exactly when the denominator wobbles. Fieller widens and skews the interval the way the mathematics actually demands.
  • Sequential: the engine's own always-valid bound for the absolute difference, divided by the control mean. One honest caveat: the anytime-valid guarantee covers the absolute difference. Expressing it relative to the control treats the control mean as fixed. That is the same plug-in convention as the displayed lift percentage itself: the stop/continue decision always stays on the absolute boundary.

Where you meet it: the interval is quoted in the verdict sentence whenever a Bayesian or Frequentist result is callable (for example "…, +12.4% (95% credible interval +5.1% to +19.8%)"). It is also what a ROPE is tested against, if you have set one on the Analysis step. It is carried on the results API payload rather than shown as its own column.

When no honest interval exists, A vs B leaves it out rather than printing a number. That happens when a variation has no data yet, when the metric's variance is degenerate, or when the control's own measurement is so noisy that its interval includes zero. Dividing by a number that might be zero produces nonsense, and A vs B refuses to fabricate a confident-looking range from it.

For metrics where lower is better, the interval is normalized to goal direction. Its bounds are negated and swapped, so positive always means improvement, matching the goal-aligned lift percentage.

Peeking: what each engine actually does

"Peeking" means checking results as they accumulate and stopping when you like what you see. It is the most common way honest teams ship false winners, and the three engines handle it differently.

Frequentist: peeking genuinely breaks it. A p-value assumes one look, at a sample size fixed in advance. If you check every day and stop the moment p drops below 0.05, your real false-positive rate lands far above 5%. A vs B does not hide this. On a Frequentist experiment, the Results page shows a peek counter that estimates your inflated error rate from how many times you have looked. It also offers to switch future experiments to Sequential. You can still stop early: the decision just gets recorded honestly.

Sequential: peeking is safe by design. The always-valid guarantee holds at every sample size. This is the engine to choose if you know you will watch results daily.

Bayesian: you can read it any time, but stopping rules still matter. The posterior updates with each data point and is a valid summary of the evidence at any moment. There is no "you looked, now it is invalid" problem. But if your rule is "stop the moment probability-to-beat crosses 95%", you will still stop on favourable noise more often than the number suggests, especially at small samples. Reading the posterior is free. Stopping on it is a decision, and the decision carries risk.

No engine rescues you from too little data

With 50 visitors per arm, a 90% winning probability can flip completely as data arrives. Whichever engine you choose, wait for a sample size you planned in advance (use the sample-size calculator) before making a final call.

How the peek counter estimates your real false-alarm rate

The counter is keyed on the number of looks, and nothing else, not on how far through your planned sample you are. Progress is not what breaks a p-value. Repeatedly looking and stopping at the first crossing is. A model keyed on progress would show the most risk early and fall back towards "no inflation" exactly as your peeking piled up, which is backwards.

The multipliers come from the classical repeated-significance-testing results of Armitage, McPherson & Rowe (1969), "Repeated significance tests on accumulating data", Journal of the Royal Statistical Society A, 132(2). That paper gives the cumulative type-I error for k equally-informative looks:

Looks (k)Cumulative false-alarm rate at α = 0.05As a multiple of α
15.0%1.00×
28.3%1.66×
310.7%2.14×
412.6%2.52×
514.2%2.84×
1019.3%3.86×
2024.8%4.96×
5032.0%6.40×

Between those points A vs B interpolates in log(k); beyond 50 looks it continues the same slope. The reported rate is capped at 50%. Past a coin flip, the exact figure has stopped meaning anything actionable.

A look is counted once per person per UTC day; see the peek counter for exactly what does and does not count.

The limits of this estimate: stated plainly

Every number the counter shows is labelled an illustrative estimate, because of two real simplifications:

  1. The table assumes evenly spaced looks. Real checking is bursty: three times on launch day, then nothing for a week. The true inflation depends on when you looked, not only how often.
  2. The table is anchored at α = 0.05. The true multipliers depend on α. At α = 0.01 the k = 5 factor is nearer 3.3× than the 2.84× above. A vs B applies the α = 0.05-anchored factors to whatever α you configured, which understates the inflation for a stricter α.

The second limitation could be removed with a group-sequential boundary solver. We have deliberately not built one. The counter's job is to be a directionally honest warning, not a design boundary, and a warning that is roughly right and clearly labelled beats a precise number nobody reads. If you want a method with no peeking penalty at all, that is exactly what the Sequential engine is for. Its boundaries are valid at every look by construction, so there is no inflation to estimate.

Money metrics (revenue / profit per visitor)

Revenue per visitor and profit per visitor are continuous per-visitor tests, not conversion tests. Every exposed visitor contributes one value (the amount they spent, or the profit they generated, 0 for non-buyers). The engine analyses the distribution of those values:

  1. Winsorization first. Per-variation values above the configured percentile cap (99th by default on the money metrics) are clamped down. That way a single whale order cannot dominate the mean or blow up the variance. The results row reports how many values were capped.
  2. CUPED second (when enabled and the auto-gate passes). Each visitor's pre-experiment spend is used as a covariate to strip predictable noise from their in-experiment value. The covariate is a strict pre-exposure attribute. It is measured for every exposed visitor over the window before the experiment started (0 for visitors with no pre-period activity), and it is never conditioned on what the visitor did during the experiment. A covariate restricted to in-experiment buyers would quietly bias the lift toward zero, and an A/A test would not catch it. So the platform's covariate queries are anchored on exposures by construction and guarded by tests.
  3. The continuous engine last. The adjusted per-visitor values feed the engine you selected: Bayesian normal posterior, frequentist Welch-style z-test, or the sequential always-valid bound. That produces a mean per visitor, a 95% interval, and the money lift versus control.

The key property: significance and lift are computed on the money value itself, never on the purchase-conversion rate. A test whose variation converts 10% fewer visitors but earns 26% more per visitor reports a positive, money-denominated lift. The conversion drop is still visible, on the separate purchase-conversion metric, where it belongs. Money amounts are analysed in integer minor units (cents) and formatted at display time.

The delta method (ratio metrics)

Ratio metrics (average order value, items per order, refund rate) divide one random variable by another. The standard t-test produces wrong confidence intervals on that shape. A vs B uses the delta method, a first-order Taylor expansion that approximates the variance of X/Y from the variances and covariance of X and Y:

Plain text
Var(X/Y) ≈ (μ_X / μ_Y)² · [ Var(X)/μ_X² − 2·Cov(X,Y)/(μ_X·μ_Y) + Var(Y)/μ_Y² ]
Plain text1 line

This is the standard variance estimator across the industry; Eppo and Statsig both document this exact form for ratio metrics like revenue per visitor. For the frequentist engine the delta-method variance feeds a z-test on the ratio. For the Bayesian engine it feeds a normal-approximation posterior. For the sequential engine it feeds an always-valid bound that is wider than the fixed-horizon CI by design. See Ratio metrics for the user-facing explanation.

Bootstrap CIs (quantile metrics)

Quantile metrics (p50 / p90 / p95 / p99 of a continuous value) do not have a closed-form sampling distribution. This is the one place A vs B does resample. It uses a BCa (bias-corrected and accelerated) bootstrap (Efron 1987). The method resamples the visitor values with replacement (1000 resamples by default), recomputes the quantile on each resample, and takes percentiles of the resulting distribution as the 95% CI. The bias-correction and jackknife acceleration terms adjust for bias and skew in the bootstrap distribution, which matter at extreme percentiles and small samples. When the acceleration term cannot be estimated, it falls back to the plain percentile method.

Resampling here is still deterministic: the draws come from a seeded generator, so the same data always yields the same interval. As with the Bayesian numbers, an interval never shifts just because you reloaded the page.

For very large experiments (N > 100k per variant), the bootstrap falls back to a t-digest-sampled subset returned by ClickHouse rather than the raw rows. The results row carries a Sampled bootstrap (N > 100k) badge so you know the interval came from a summary, not the raw data. If the t-digest estimate and an exact ClickHouse recompute disagree by more than 5%, a second badge tells you to verify the number against the exact recompute before you trust it. See Quantile metrics.

Weighted-sum statistics (composite metrics)

Composite metrics combine multiple component metrics into a single weighted decision signal. A vs B builds one per-visitor composite series. Every exposed visitor contributes a single number (the weighted sum of their component values, counting 0 for components they never triggered), and the engine reads the mean and variance straight off that series.

Plain text
y_v = Σᵢ wᵢ · xᵢ,v        for each exposed visitor vVar(composite) = Var(y)   the sample variance of that one series
Plain text2 lines

This is why no covariance term appears. Because each visitor reduces to a single number, any correlation between components is already inside Var(y). It is captured exactly, rather than estimated pairwise. There is no co-observation count to fall below, and no widened "conservative" interval to fall back to.

The series is computed once per result run (a single query across every exposed visitor) and reused by the three engines: weighted z-test for frequentist, weighted Normal posterior for Bayesian, and weighted always-valid bound for sequential. See Composite metrics.

Composite metrics skip CUPED

CUPED (variance reduction) does not apply to composite metrics today. It only runs on ordinary continuous metrics. A composite metric's results are always the raw weighted series, whatever your project's CUPED setting says.

Guardrail non-inferiority test

Guardrail metrics get a one-sided non-inferiority test rather than a difference test. Take the goal-aligned difference Δ̂ (positive = good for that metric's own goal), the Welch standard error SE, the control mean μ̂_C, and a relative margin δ_rel. The absolute margin is δ_abs = δ_rel · |μ̂_C|, and the bounds are

Plain text
LB = Δ̂ − q·SE        UB = Δ̂ + q·SE
Plain text1 line
  • Safe when LB > −δ_abs: harm worse than the margin is ruled out at α.
  • Breached when UB < −δ_abs: a margin-exceeding regression is established at α.
  • Inconclusive otherwise: the interval straddles the margin.

The multiplier q is engine-matched. It is the one-sided normal critical value Φ⁻¹(1 − α) under Frequentist and Bayesian. Under Sequential it is the always-valid confidence-sequence multiplier evaluated at 2α: a two-sided sequence at 2α gives a valid one-sided time-uniform bound at α, so guardrail statuses are peek-safe. The reported non-inferiority p-value is p_NI = Φ(−(Δ̂ + δ_abs)/SE). Under Sequential it is a snapshot quantity; the status, not the p-value, carries the anytime-valid guarantee.

Guardrail p-values never enter the multiplicity-correction family: correcting a safety test would reduce power to detect harm, which is the anti-safe direction. Margins are frozen into the sealed analysis plan at launch, so a later margin edit never rewrites a running experiment's judgement.

Expected loss (risk of shipping)

Every non-control arm reports a pair of numbers that turn a close call into a priced decision:

  • Expected loss of shipping: if you ship this variation and it is secretly worse, this is the average amount you lose per visitor.
  • Expected foregone gain of keeping control: if you keep control and the variation is secretly better, this is the average amount you leave behind per visitor.

Both are averages over the posterior's losing scenarios, not worst cases. They are reported in the metric's own units per visitor: rate points for conversion metrics, minor currency units (cents) for money metrics.

For continuous and ratio metrics the difference between arms is approximately normal: Δ ~ N(m, σ²), with m the goal-aligned mean difference and σ the Welch standard error of the difference. The pair has a closed form (the truncated-normal partial expectation, z = m/σ):

Plain text
EL(ship) = σ·φ(z) − m·Φ(−z)EL(keep) = σ·φ(z) + m·Φ(z)
Plain text2 lines

σ is always the engines' own plain sampling standard error. Under the sequential engine it is deliberately not the width of the always-valid interval, which is several times wider by design. Deriving σ from that interval would silently inflate the sequential risk figure by the same factor. Using the sampling SE is what makes the number comparable across engines.

The two are linked by an exact identity (EL(keep) − EL(ship) = m) which the test suite pins. When σ = 0 the pair collapses to the point mass max(0, ∓m).

For binary conversion metrics the engine uses the exact flat-prior Beta posteriors (Beta(1 + conversions, 1 + visitors − conversions) per arm) and evaluates the loss integral deterministically, with no random draws:

Plain text
EL(ship) = E[max(θ_control − θ_variation, 0)] = ∫₀¹ F_variation(t) · S_control(t) dt
Plain text1 line

where F is a posterior CDF and S a posterior survival function. This single-integral form is exactly equal to the textbook two-term loss identity, but it never subtracts two nearly-equal numbers. Near-tie experiments keep full precision as a result. The integral runs on the same mass-based quadrature panels as the winning probability. It stays accurate even when one arm has millions of visitors and the other a handful, which is the shape a ramping experiment takes.

The monthly money figure. When the primary metric is a money metric, the results page projects the per-visitor risk to a monthly amount: observed visitors per day (distinct exposure days, across all arms), times 30, times the per-visitor expected loss, in the project currency. The projection only appears when three things are all true: the money read succeeded, the leading variation's per-visitor risk is computable, and at least three distinct days of traffic were observed. Otherwise the card falls back to native units rather than fabricate a monthly figure.

Honesty notes. The number takes the statistical model at face value and assumes next month's traffic looks like the experiment's. Under the frequentist engine it is the same computation, framed as an estimated risk. Under the sequential engine it is a snapshot estimate: unlike the sequential verdict, it is not peek-proof.

Honest lift: the winner's-curse correction

Conditioning on success biases the estimate. You only act on a result because it crossed the decision boundary, and among all experiments that cross a boundary, over-estimates cross more often than under-estimates. The "honest lift" caption on the Results page corrects for exactly that selection event.

The estimator

Work in z-space, where the observed effect is a = |z| and the engine's decision boundary is c (both in standard-error units). The selection event is two-sided: a flagged winner or a flagged significant loser both condition the data. So for a true standardized effect μ, the probability of being selected is Φ(μ − c) + Φ(−μ − c). The estimator is the conditional maximum-likelihood estimate given selection (in the tradition of Zöllner & Pritchard's winner's-curse correction). It maximizes

Plain text
ℓ(μ) = −(a − μ)²/2 − log( Φ(μ − c) + Φ(−μ − c) ),   μ ∈ [0, a]
Plain text1 line

The reported estimate shrinks the observed lift by μ̂ / a. Because a z-statistic is proportional to the effect at a fixed standard error, shrinking z shrinks the relative lift by the same factor. No unit conversions and no reconstructed standard errors are needed. Worked examples at the default boundary c = 1.96: an observation at z = 3.0 keeps ~84% of its size; at z = 2.5 it keeps ~46%; a result sitting exactly on the boundary keeps ~25%. Far past the boundary (z ≥ 6) the correction vanishes, so well-powered results are barely touched.

We chose the conditional MLE over empirical-Bayes shrinkage deliberately. It needs no cross-experiment history of "typical" effect sizes (a brand-new account has none), and it is fully deterministic from the experiment's own numbers.

The boundary c, per engine

The correction is only honest if c is the boundary that actually selected your result:

  • Frequentist: c = z₁₋α/₂ at your frozen α, on the raw (pre-correction) z. When a multiple-comparison correction adjusted the family, the true selection boundary is higher than this c. By monotonicity, that means the true correction is at least as large as shown, so the caption discloses the shown adjustment as a lower bound. (A data-dependent Benjamini–Hochberg boundary is not cleanly invertible; a disclosed bound beats fake precision.)
  • Sequential: c is the real always-valid (AsympCS) boundary at the current sample size, the same multiplier the engine's confidence sequence uses. It is always wider than z₁₋α/₂. Using the fixed-horizon boundary here would understate the correction.
  • Bayesian: the winner probability is mapped through the normal approximation (a = Φ⁻¹(P)) and c = Φ⁻¹(t), where t is your resolved decision threshold. That is the same frozen/configured value the verdict itself judges with, so the selection event and the verdict can never disagree about the bar.

The wrong-sign number: plug-in Type-S

The caption's "≈X% of results selected like this one would point the wrong way" is the Gelman–Carlin Type-S error evaluated at the adjusted estimate, a plug-in:

Plain text
Type-S(μ̂) = Φ(−μ̂ − c) / ( Φ(μ̂ − c) + Φ(−μ̂ − c) )
Plain text1 line

This is the probability that a result selected through this boundary has the wrong sign, if the true effect equals μ̂. It lives in [0, 50%] and approaches 50% as the adjusted estimate approaches zero. It is deliberately not the unconditional Φ(−μ̂), which ignores selection and can be off by three orders of magnitude well past the boundary.

Honesty notes. The adjusted estimate is itself an estimate: a maximum-likelihood point, not a posterior mean, and it inherits the normal approximation. The Bayesian mapping through Φ⁻¹ is an approximation of a posterior tail by a z-value. Both approximations are disclosed here rather than hidden. The caption's job is to replace a number known to be biased upward with one that is not, and our seeded calibration suite holds it to that standard: across simulated selected experiments, the corrected estimate must remove more than half of the selection bias without losing on absolute error. That check is re-proved on every test run against the shipped code.

Was this helpful?