Winning Probability

Winning probability is the single most important number on the Results page. It tells you how sure the statistical model is that a variation is really better than control, not just better by random chance. Here is how to read it, and when to act on it.

What it means

When the Results page shows a variation with a winning probability of 95%, here is what that means. Based on all the data collected so far, there is a 95% probability that this variation truly has a higher conversion rate than control.

There is still a 5% chance the observed difference is a fluke, meaning control is actually just as good or better. But the evidence strongly points toward the variation winning. A 95% probability is a high bar. Most business decisions can be made confidently at this level.

Conversely, a probability of 60% means: there is a 60% chance the variation is better, but also a 40% chance it is not. That is barely better than a coin flip. You need more data.

You will never see a flat 100%

A vs B caps the displayed probability at "greater than 99.9%" (shown as "> 99.9%", or "> 99%" where a surface displays whole numbers). No experiment can prove a winner with absolute certainty, so the display never claims one. Read a capped value as extremely strong evidence, not as a guarantee.

The 95% threshold

By default, A vs B uses 95% as the threshold for calling a winner. When a variation crosses this threshold, it is marked with a Significant badge in the results table. The Winning Probability summary card also highlights the leading variation.

You can adjust this threshold under Settings → Analysis → Bayesian settings, where it becomes the default for every new experiment and flag rule in the project. Some teams prefer a lower threshold, such as 90%, when moving fast and the cost of a wrong decision is low. Others use a higher threshold, such as 99%, for high-risk changes, or changes they cannot undo.

Whatever you set is the bar every surface judges against. Five places all read that same resolved threshold. They are the Significant badge, the winner tick in the results table, the decision header's verdict, the compare-engines modal, and the A/A guardrail. (An A/A guardrail catches a broken experiment. It splits traffic into two identical groups. Then it watches for a difference that should not be there.) Every surface reads the same number. So they can never disagree about who won.

The Significant badge

The Significant badge appears next to a variation in the results table once its probability to beat control has crossed your threshold. This badge is your at-a-glance signal that there is enough evidence to declare a winner.

The badge does not automatically stop the experiment. You decide when to end it. Stop it yourself, or let it keep gathering data.

When to call a winner

The winning probability crossing 95% is a necessary condition for calling a winner, but not the only one. Before making a final decision, also verify:

  • The experiment has run long enough. Results collected in the first few days of an experiment can be misleading. Two common causes are novelty effects (visitors clicking on something simply because it is different) and day-of-week bias. Try to run experiments for at least one full week, and ideally two, so you cover a complete business cycle.
  • No health guardrails are red. Check the health guardrails section for warnings. Watch especially for a sample ratio mismatch: a red flag that fires when the traffic split between variations does not match what you configured. That usually signals a bug, not a real result. A significant result on corrupted data is not a reliable result.
  • Secondary metrics look sensible. Your primary metric is the one metric an experiment is actually judged on. A secondary metric is tracked alongside it for context, without deciding the result on its own. If your primary metric improved but a related secondary metric got worse, investigate before shipping. For example, if purchases went up but so did refunds, that warrants a closer look.
  • The effect size is meaningful. A 0.1% improvement in conversion rate with 99% winning probability might not be worth the engineering effort to ship. Consider whether the size of the lift justifies making the change.

Why you should not stop too early

Peeking means checking results before an experiment is finished, and reacting to what you see. A Bayesian model handles peeking better than a classic frequentist test does. But it is not immune to stopping too early. Stop the moment your winning probability first crosses 95%, and you can inflate your false positive rate. That rate is how often a test calls a result a win when nothing actually changed. Early on, random variation can push a probability above 95% for a moment. It can then settle back down as more data arrives.

A practical rule: once the winning probability crosses 95%, wait at least 24 to 48 more hours to confirm it stays above the threshold. If it drops back below, the experiment needs more data. If it remains above, you can be more confident in the result.

Experiment runtime matters

Run your experiment for at least one full business cycle: 7 to 14 days. That is the single most useful thing you can do for reliable results, whatever the winning probability says. Traffic patterns vary by day of the week. An experiment run only on weekdays can look very different from one that spans a full week.

No winner does not mean the experiment failed

Sometimes an experiment concludes without a winner, meaning the winning probability never reached the threshold. That is a valid result too. It means the two variations performed similarly. Knowing that a change does not improve, or hurt, conversion is sometimes exactly the information you need to prioritize other work.

Was this helpful?