Bayesian Engine

The Bayesian engine is the default way A vs B judges an experiment. It answers one question directly: how likely is it that this variation beats Control? The answer is a probability, like 92%, plus a credible interval: a range that states, directly, how likely the true improvement is to fall inside it.

What Bayesian analysis actually does

Bayesian analysis asks one question: given the data so far, how likely is it that this variation is really better than Control? The answer is a probability, say 92%, about the thing you actually care about.

That is a different question from the one a p-value answers. A p-value is the chance of seeing data this extreme if the variation changed nothing at all. It is not the probability that the variation works. The Bayesian number is the probability that the variation is ahead. Both are legitimate statistics. Only one of them means what people usually assume it means.

For binary metrics, like a conversion rate, A vs B uses a Beta-Binomial model: the standard Bayesian approach for this kind of data. For continuous metrics, like revenue per visitor or time on page, it uses a normal approximation around the mean instead.

A worked example

Say Control converts 240 out of 5,000 visitors (4.8%), and Variant B converts 275 out of 5,000 (5.5%): a point lift of about +14.6%. Feed those into the formula below and A vs B works out that there is a 94.3% probability B really beats Control. That's good, but just under the 95% bar A vs B uses by default to call a winner (see below). The rest of this page explains how that number, and the interval next to it, get built.

Probability to beat control

Each variation's headline number is the posterior probability that its true conversion rate is higher than Control's. A posterior is just your updated belief after combining a starting assumption (the prior, covered further down) with the data you actually collected.

Starting from a flat prior, observing c conversions from n visitors gives a posterior of:

Plain text
Beta(1 + c, 1 + n − c)
Plain text1 line

For the example above, Control's posterior is Beta(241, 4761) (1 + 240, and 1 + 5,000 − 240), and B's is Beta(276, 4726).

A vs B then computes the probability that the variation's rate exceeds Control's directly from those two posteriors: an exact integral, not a simulation. That is the 94.3% from the worked example.

The same data always gives the same number

Many tools estimate this probability by drawing thousands of random samples from the posteriors. That makes the figure drift slightly on every recompute. A verdict resting right on the threshold can flicker on and off between page loads. A vs B computes it exactly instead, so the same inputs always give the same probability.

With very little data it hovers near 50%: no evidence either way. As one variation consistently converts better, it climbs toward 90%, 95%, 99%. 95% is the default threshold for calling a winner, and variations past it get a Significant badge.

That bar is its own setting: the Bayesian decision threshold field in the builder (the number box itself is labelled Chance to beat Control). You can change it per project or per experiment. It is a different setting from the confidence-level field described next. That field sets the width of the credible interval below, not this winner bar. The two default to the same number, 95%, which makes them easy to mix up. Changing one never moves the other.

Credible intervals

Alongside the probability, A vs B reports a 95% credible interval for the rate and the lift. That 95% comes from the same confidence-level setting every engine uses (default 95%); narrow or widen it there and this interval moves with it. Read the result literally:

A 95% credible interval of [+5%, +15%] means there is a 95% probability the true improvement is between 5% and 15%.

This is the interpretation people want from a confidence interval, the frequentist equivalent. They cannot have it: a confidence interval is a statement about the long-run behaviour of the procedure, not about this one result. A credible interval genuinely is a statement about this result.

For the worked example above, Control's 95% credible interval is [4.2%, 5.4%] and B's is [4.9%, 6.2%]. The relative-lift interval, how much better B is as a percentage of Control, is [−3.2%, +35.6%].

Two things to watch:

  • Width is the real signal. A narrow interval means you know the effect size. A wide one (say [−5%, +25%]) means you do not, whatever the point estimate says.
  • Overlapping zero means inconclusive. If the lift interval spans zero, you cannot rule out "no difference" (or a loss), no matter how good the headline looks. That is exactly the worked example above: despite a 94.3% probability, the lift interval crosses zero, so the honest read is "promising, not yet proven."

Like the probability, credible intervals for conversion rates are computed exactly from the posterior rather than sampled. The bounds do not shift on their own; they only move when the data does.

The prior

Bayesian statistics needs a starting assumption, called a prior: what you believe before any data arrives. A vs B starts from a flat, non-informative prior, Beta(1, 1), the uniform distribution. In plain words: before the data, every conversion rate is equally plausible.

In practice this means the results are driven entirely by your data. There is no thumb on the scale and no hidden assumption that your variation should win. You do not need to supply historical belief before you can run a test. It also means Bayesian and Frequentist results usually land in a similar place on the same data. The difference is what the number means, not a systematically friendlier answer.

The sample-size calculator can take an informative prior

The flat prior is fixed for results analysis. The sample-size calculator separately accepts an optional informative prior when planning a Bayesian experiment, which can reduce the sample you need.

Multiple variations

Multiple-comparison corrections are a Frequentist device. They exist to hold a fixed error rate across repeated significance tests. The Bayesian engine does not apply them. Each variation's probability is a direct posterior statement about that one variation, not a hypothesis test in a family of tests.

That does not make many-armed tests free. Adding arms splits your traffic, so every arm gets less data and every interval gets wider. If you scan five variations and act on whichever one crossed 95%, you are still selecting on noise. The engine is not the thing protecting you there. Decide up front which comparison matters. See Analysis plans.

Peeking

Peeking means checking results before an experiment finishes and reacting to what you see, which can make random noise look like a real effect. The Bayesian posterior updates with every new visitor, and it is a valid summary of the evidence at any moment. So there is no "you looked, now it is invalid" penalty: this is why Bayesian experiments show no peek-protection warning.

But reading and stopping are different acts. A rule like "stop the moment probability-to-beat crosses 95%" still stops on favourable noise more often than 5% of the time. That gets worse at small samples. The honest version: look as much as you like, and stop on a sample size you planned in advance.

Want a guarantee that survives continuous monitoring and early stopping too? Use the Sequential engine instead. It stays valid even when you check daily and stop early.

When to pick Bayesian

  • You want results that read as "87% probability B beats Control" rather than "p = 0.018".
  • Your stakeholders are non-technical and respond better to probability than to p-values.
  • Your traffic varies, so pre-committing to a precise sample size is hard.
  • You have no external requirement to report in p-values.

When not to pick Bayesian

  • A stakeholder or compliance reviewer expects p-values. Use Frequentist.
  • You are in a regulated setting that requires a fixed-α, pre-declared-sample-size design. Use Frequentist.
  • You know you will watch daily and stop early. Use Sequential, which is built to make that safe.

Where to configure

Bayesian is the default, so there is nothing to do unless you want to change it. To set an engine explicitly:

  • Project level. Project Settings → Analysis sets the default for every new experiment and feature-flag A/B test rule in the project.
  • Per experiment. In the experiment builder's Analysis card, before launch. Once the experiment is running the engine is locked: mixing engines mid-flight would invalidate the result.
  • Per flag rule. In the feature-flag A/B Test rule editor's Analysis section, before the rule is enabled. Same lock-on-launch behaviour.

Further reading

Was this helpful?