Results Look Wrong
Some numbers on the Results page look surprising the first time you meet them. This page explains the most common concerns.
Before you read on, check which stats engine your experiment uses: the engine pill sits next to the results title. A vs B ships three (Bayesian, Frequentist, Sequential), and several of the symptoms below are specific to one of them. If a section below names an engine you are not using, it does not apply to you.
"The conversion rate seems too low"
Conversion rates in A vs B are calculated per unique visitor, not per pageview. The formula is:
Conversion rate = Unique visitors who converted ÷ Total unique visitors in the variationA visitor who lands on the page three times only counts once in the denominator. A visitor who converts on their second visit is still only counted once as converted. This means conversion rates are often lower than what you might expect from pageview-based analytics.
This is the statistically correct way to measure conversion rates for A/B testing. It ensures frequent visitors do not have outsized influence on the results. If your site has visitors who return many times, your per-visitor conversion rate will be lower than your per-session or per-pageview rate.
"The results keep changing"
If you check the results every few hours and see the winning probability shifting back and forth, this is completely normal and expected. A vs B uses a Bayesian model that updates continuously as new data arrives. When you have a small number of visitors, a few conversions in one direction can swing the probability significantly. As more data accumulates, the estimate stabilises.
Think of it like an election. Early returns from a small number of precincts can point strongly to one candidate. But as more votes are counted, the picture becomes clearer and more stable. The same applies to your experiment results.
The right approach is to set a target number of visitors before you run the experiment. Run it until you hit that target, then read the results. Do not check every few hours and stop the moment things happen to look good.
"Peeking" means checking results early and stopping an experiment the moment you see a positive one. It is one of the most common sources of false positives in A/B testing. At any given moment, a result might look good by chance. Only stop an experiment when you planned to stop it, or when you clearly have enough data.
"Control is winning: is that bad?"
No. A control winning is a completely valid and valuable result. It means your variant (the version being tested against control) did not improve the metric compared to the original. Now you know that, rather than guessing. This is one of the most important outcomes in A/B testing: learning what does not work.
A control winning should lead you to ask:
- Was the hypothesis wrong? (The change I thought would help actually does not.)
- Was the implementation right? (Did the variant look and work as intended?)
- Was I measuring the right thing? (Is this metric actually linked to the business goal?)
Answer these questions, form a new hypothesis, and design the next experiment. A control winning is learning, not failure.
"Confidence is stuck at 50%" (Bayesian)
A probability of around 50% means A vs B has no evidence that either variation is better than the other. Both are equally likely to be the winner, given the data seen so far. This is the Bayesian starting point before any data is collected, and it persists when there is not enough data to distinguish between the two.
This is not a bug. It means you need more data. The probability will move away from 50% as more visitors are counted and the two conversion rates diverge, if they ever do. If after thousands of visitors the probability stays near 50%, it suggests the true effect size is very small. The two variations may be genuinely equivalent.
"The p-value is stuck / it never says Significant" (Frequentist)
A p-value that hovers around 0.4 or 0.8 and refuses to drop is not a malfunction. It is the test telling you that data like yours is entirely unremarkable if the variation changed nothing.
Three things worth separating:
- You need more data. p-values are noisy at small samples and drift a lot. Check the Traffic Health guardrail (a check on whether your sample is even big enough to trust) and your planned sample size before concluding anything.
- The effect is smaller than you powered for. If you are well past your planned sample size and p is still high, the honest reading is usually this: the effect, if any, is too small for this experiment to detect. That is a real finding.
- There is genuinely no effect. Also a real finding, and a common one.
A p-value near 1.0 does not mean "the variation is definitely identical to control": it means the data is unsurprising under that assumption. Frequentist tests can fail to reject; they never prove the null.
A p-value assumes a single look at a sample size fixed in advance. Watching daily and stopping the first time it crosses 0.05 is the fastest way to ship a false winner. Your true error rate ends up far above 5%. The Results page shows a peek-protection warning with an estimate of the inflation. If you know you will want to watch continuously, use the Sequential engine instead: the engine built to stay valid under exactly that kind of daily checking.
"Sequential says Inconclusive but the lift looks big" (Sequential)
This is the Sequential engine working as designed, not ignoring your data.
Sequential buys you the right to look as often as you like, and it pays for that with wider intervals. Its confidence sequence is valid at every sample size simultaneously, a much stronger guarantee than a fixed-horizon p-value. The price is that it needs more data to reach the same verdict. A lift that a Frequentist test would already call significant can still read Inconclusive under Sequential.
So an Inconclusive verdict means: not yet enough evidence to stop safely. Your options are to keep collecting, or to accept that you chose the conservative engine and the answer is not in yet.
If you want to see how the same data reads under the other engines, use Compare engines on the Results page. Treat that as context for a decision, not as licence to shop for the engine that agrees with you. The engine you chose up front is the one whose guarantee you actually hold.
"The revenue card shows a huge number"
Usually because it is measuring something bigger than you think it is, or a different metric than you expect.
The Revenue observed card is a gross total of revenue seen on the non-control variations during the experiment window. It is not an incremental figure, not a projection, and not a per-month or per-year rate. Nothing is subtracted for what control earned, and nothing is extrapolated forward.
It also does not always come from your primary metric (the one metric an experiment is actually judged on). A vs B searches every metric attached to the experiment for the first one that carries a real revenue total. So a click-conversion primary with a revenue metric attached as a secondary still shows a real number. The card names which metric the total came from, in a line under the figure. If that metric tracks profit rather than revenue, the card reads Profit observed instead. Check the label before comparing it with a number from elsewhere.
- This is a gross total, not a lift: read the Revenue per visitor row for the number to act on.
- The card names its source metric here, since it may not be your primary metric.
So on a high-traffic experiment the number is large simply because a lot of money flowed through the variants. Most of that you would have earned anyway. It answers "how much revenue passed through the variants while this ran?". That is rarely the question you actually want answered.
For the decision, read the revenue metric rows instead, not this card. The Revenue panel's Revenue per visitor row gives you the mean per visitor, an interval, and a money lift versus control. That is an actual like-for-like comparison, and the lift is the number to act on.
Because the card sums observed revenue rather than modelling anything, a large figure carries no evidence that the variation caused revenue. Pair it with the money lift on the Revenue per visitor row and the health guardrails before drawing conclusions.
"Results differ across segments"
Seeing different results for different segments is extremely common and often the most valuable insight from an experiment. Mobile and desktop users almost always behave differently. New visitors and returning visitors respond to changes differently. Premium users and free users have different motivations.
If your overall result is inconclusive but one segment shows a strong positive effect, consider running a follow-up experiment targeted specifically at that segment. Conversely, if one segment shows a negative effect while another shows positive, a one-size-fits-all variant may not be the right approach.
Segment analysis is exploratory: use it to generate hypotheses, not to declare a winner within a subgroup. Segments only filter and describe results after the fact; they are not something you can target visitors with directly. The experiment was not powered to detect effects within individual segments. So confidence levels for segment-level results will be lower than for the overall result.
"How do I know when I have enough data?"
Decide the answer before you start, not by watching the Results page.
Use the sample-size calculator to work out how many visitors you need for the smallest effect worth detecting. Do this at your chosen engine and confidence level. Record that target, then run to it. The calculator is where the sample-size and duration math lives; the Results page reports what happened, it does not tell you when to stop.
As a rough sanity check, you generally need at least 100 conversions per variation before results stabilise enough to act on. For low-conversion-rate experiments (below 1%), that can mean thousands of visitors per variation.
If you reach your planned sample size and the result is still inconclusive, the true effect is probably smaller than the one you powered for. The experiment ran long enough, and the answer is genuinely "no meaningful difference." That is a result, not a failure.
While the experiment runs, the Traffic Health guardrail tells you whether your smallest variation has a workable sample at all (green at 1,000+ visitors). See Health Guardrails.
Decide your stopping criterion before the experiment starts. Use either a target number of visitors, a target number of conversions, or a calendar date. Then stop on that criterion, not because the result happens to look good or bad on any given day.