Reading Your Results

The Results page gives you a full picture of how your experiment is doing. A variation is one version being tested: control, or a challenger. The page has a summary section at the top, a detailed variation table in the middle, and metric breakdowns below. Here is a walk-through of every part, and what it tells you.

Every number on this page comes only from visitors who allowed measurement. A visitor who refuses analytics in your cookie banner may still see the experiment, but adds no data to it. See Consent Mode for how that works.

The top of the Results page: summary cards and the conversion-rate time series.

Filtering the view

Above the numbers sits a filter bar with a date range, a segment picker, and a baseline picker. Every chart, card and table on the page answers the filtered question together, so what you see is always one consistent view.

  • Date range works in UTC calendar days. On plans that keep a limited reporting window, the control names the real window it is showing, with a "plan limit" note, instead of claiming "All time" when the window was shortened. A notice above the results explains the plan window in full. Both appear only when the window really cut off some of this experiment's data: an experiment that started inside your plan's window shows neither.
  • Segments narrow the view to matching visitors, such as one device type or one country. A segment you pick after seeing the results is an exploratory look rather than a verdict, so while one is applied the filter bar tags its winning probability Exploratory slice. Read a tagged number as a lead worth testing properly, not as a result to ship on. Shared links keep your date range and segments, so a teammate opening your link sees the same view you were looking at.
  • Baseline changes which variation the others are compared against. It changes presentation only, never the underlying data.

If the filtered view has no data, the page says so and offers the way forward: Clear filters removes the date range and segments in one click, and Refresh results re-reads the numbers.

Keeping the numbers current

Opening the Results page reads your numbers fresh, and Refresh results recalculates them on the spot whenever you want the very latest.

Under the Decision Header there is also an Auto-refresh every 5 minutes tick box, so you do not have to keep pressing the button while a test is live. It is off by default, remembered per experiment, and it only runs while the experiment is running and the tab is in view.

An automatic update shows the most recent saved calculation rather than recalculating from scratch. A vs B recalculates that saved copy in the background roughly every 15 minutes while your experiment is collecting data, so an auto-refreshed page is usually a few minutes behind the very latest events. You never have to guess by how much: the "as of" time beside the numbers always names the moment they were calculated, and pressing Refresh results brings both the numbers and that time right up to date.

What an automatic update leaves out

Three parts of the page are worked out only on a full read, so an automatic update does not refresh them:

  • Segment Lift, the per-segment breakdown.
  • Recommendation performance, if your experiment uses recommendations.
  • The visitor count in the confirmation when you turn a variation off.

They do not disappear. Each keeps its place and says it was not worked out this time, and pressing Refresh results brings it straight back. Nothing you were reading is lost or wrong; it is simply waiting for a full read.

Warning

Before turning a variation off, press Refresh results first. The confirmation tells you how many visitors have already been counted in that variation, and it can only show that number after a full read. Turning a variation off cannot be undone for the visitors already measured, so it is worth seeing the figure before you decide.

Analysis label

Near the top of every Results page, a small pill shows how A vs B analysed this experiment. For example: Bayesian · Auto (CUPED applied) or Bayesian · Off. CUPED is a technique that uses a visitor's pre-experiment behaviour to cut noise from the result. Hover the pill for a plain-English explanation, and, when CUPED ran, how much noise it removed.

When CUPED runs on a specific metric, that metric's lift in the Secondary Metrics list also gets a hover tooltip. It shows the θ (theta) value and the percentage of variance CUPED removed for that metric. See Variance Reduction (CUPED) for a deeper explanation.

CUPED never applies to a composite metric

A composite metric is one built by combining several underlying metrics into one score. Turning on variance reduction tightens standard metrics and money metrics, but it never touches a composite metric. A composite row's lift is always computed from the raw numbers, whatever your project's CUPED setting says. So a composite metric's hover tooltip never shows a θ value or a variance-reduction percentage: there is nothing to disclose. This is current behaviour, not a display bug. Do not read a silent composite row as "CUPED failed". It was simply never in scope for that row.

Guardrail panel

A guardrail metric is one you protect, not one you try to move: things like page load time or refund rate. If your experiment has guardrail metrics, a panel directly under the Decision Header shows one row per guardrail:

  • Safe: not worse than X%. The data has ruled out harm bigger than the margin.
  • Breached: worse than X% at Y% confidence. A margin-exceeding regression is established. The verdict can never be "Ship" while any guardrail is breached.
  • Not yet established safe: the data hasn't resolved it either way. A wide result is never called safe by default.

Each row also shows the worst (or best) plausible effect against the allowed margin, as a percent of the control's value. A vs B tests guardrails one-sided against their margins. It never applies multiple-comparison correction to them, on purpose: correcting a safety test would make real harm harder to catch.

Decision risk

Say your primary metric is a money metric, and the experiment has run long enough to project traffic (at least three days with visitors). A Decision risk panel then appears alongside the guardrail panel. It puts a monthly figure on the decision itself:

  • Risk of shipping ≈ $X/mo: say the leading variation is secretly worse than the control. This is roughly what that mistake costs per month at your current traffic. It is an expected loss: an average across the scenarios where the variation really is worse, weighted by how likely each one is. It is not a worst case.
  • Upside ≈ $Y/mo: the other side of the same coin. Roughly how much you would leave on the table each month by keeping the control, if the variation is genuinely better.

A line under the figures names the traffic the projection is scaled from. For example: "Projected from ≈24,000 visitors/month, this experiment's observed traffic." The number is never a black box.

The panel is plain about its assumptions on purpose. It assumes next month's traffic looks like this experiment's, and it takes the statistical model at face value. On the Sequential engine it adds one more caveat: the money figure is a snapshot at this look. Unlike the Sequential verdict itself, it is not peek-proof.

The card only appears for money primary metrics with enough traffic to project. For any other primary metric, or before three days of visitors, A vs B leaves the card out entirely rather than show you a guess.

Summary cards

Four summary cards sit at the top of the Results page. Each one answers a key question about your experiment at a glance.

The first card depends on your engine

A vs B ships three stats engines. The leftmost card changes to match the one your experiment uses. The engine pill next to the results title tells you which is in play. If you are unsure which you are reading, start with Choosing a stats engine.

Winning Probability (Bayesian)

On a Bayesian experiment (A vs B's default engine), the first card shows the probability that the leading variation is genuinely better than the control. 95% or higher is the threshold for calling a winner. A variation that crosses it gets a Significant badge. A value below 80% is generally inconclusive: you need more data. See Winning Probability for a full explanation.

Statistical Significance (Frequentist)

On a Frequentist experiment, the same card reads Statistical Significance and shows the leading variation's p-value instead of a probability. The Significant badge appears when the p-value falls below your significance level, α (0.05 by default).

Read the p-value for what it actually is. It is the odds of seeing data at least this extreme if the variation made no difference at all. It is not the odds that your variation works. That is the single most common misreading, and it flatters results. A p-value of 0.03 does not mean "97% chance B wins".

Frequentist experiments also show a peek-protection warning when you look before reaching your planned sample size. That warning is not decoration. A p-value assumes one look at a pre-planned sample size. Stopping the moment it dips below 0.05 pushes your real false-positive rate well past 5%. The banner estimates by how much. The warning is always judged against your experiment's whole visitor count, never a date-filtered or segment-filtered view, so it reads the same whichever filters you happen to have open. See the Frequentist engine.

Sequential experiments

On a Sequential experiment, you read an always-valid p-value and a confidence sequence instead of a fixed-horizon result. The guarantee holds at every look, so checking daily never weakens it. There is no peek-protection warning to worry about. The trade is width: a confidence sequence is wider than a fixed-horizon interval at the same sample size, on purpose. So Sequential needs more data to call the same effect. Flag-rule results state the verdict directly as Safe to stop or Inconclusive. See the Sequential engine.

Observed Lift

Observed Lift is the relative gain in conversion rate between the best-performing variation and the control. Say the control converts at 3% and a variation converts at 3.6%. The observed lift is then +20%. This tells you the size of the difference, not just whether one variation is better.

Under the lift, the card shows the interval of the lift: the range the true lift plausibly sits in. While the control has no conversions yet, a relative lift can't be worked out (there is nothing to divide by), so the card shows a dash and says so, with the interval of the lift below it as the range the data allows so far.

Revenue Impact

Revenue Impact appears when your experiment includes a metric with revenue data. It estimates how much extra revenue the winning variation would bring in per month, compared to the control. The estimate is based on your current traffic and the observed revenue-per-visitor gap. It assumes your traffic stays steady and the observed lift holds.

Days Remaining

Days Remaining estimates how much longer the experiment needs to run before reaching a meaningful result. It is based on your current traffic, the observed effect size, and your target winning-probability threshold. If the experiment already passed that threshold, this card just says it is ready to call.

Variation performance table

The variation table shows one row per variation (one specific version being tested: control, or a challenger), including the control. Each row has these columns.

Visitors

The total number of unique visitors in this variation. A vs B counts each visitor once, no matter how many sessions they had or pages they visited.

Conversions

The number of visitors in this variation who completed the primary metric. Each visitor can only convert once.

Conversion Rate

Conversions divided by Visitors, as a percentage. This is the core performance number for each variation.

Improvement vs Control

The relative change in conversion rate compared to the control. A positive number means the variation converts better than the control. A negative number means it converts worse. The control row always shows 0%, since it is the baseline everything else is measured against.

Credible Interval / Confidence Interval

The interval column shows the range this variation's true effect most likely falls in. Narrower means more certainty. Wider means you need more data. Whatever your engine, the same rule of thumb applies: if a variation's interval overlaps zero, the difference from control is not yet conclusive.

What the interval actually means depends on your engine:

  • Bayesian shows a credible interval (a range that states, directly, how likely it is that the true result falls inside it). Read it literally: "there is a 95% probability the true improvement is in this range."
  • Frequentist shows a confidence interval (a range built to hold the true result a set share of the time, over many repeats). That is a statement about the procedure's long-run behaviour, not about this one result. Tempting as it is, you cannot read it as "95% probability the truth is in here".
  • Sequential shows a confidence sequence: valid at every look, and wider than the other two at the same sample size, by design.

Probability to Beat Control / p-value

The decision column tracks your engine:

  • Bayesian: Probability to Beat Control. The probability this specific variation genuinely beats the control. Above 95% is strong evidence.
  • Frequentist: the variation's p-value, significant below α (0.05 by default).
  • Sequential: the always-valid p-value, significant when the confidence sequence excludes zero.
Significant badge

The Significant badge tracks your engine. On Bayesian it marks a variation past the 95% probability threshold. On Frequentist it marks a variation whose p-value fell below α. It means "you have evidence". It does not mean "ship it". Check the health guardrails and your planned sample size before you decide.

Metric sections

Below the variation table sit collapsible sections for each metric attached to the experiment: the primary metric first, then each secondary metric in order. Each section shows the same variation table breakdown, scoped to that one metric.

Opening a secondary metric section gives you supporting context. Say your primary metric is purchases, and a secondary metric is add-to-cart clicks. If purchases went up and add-to-cart also went up by a similar amount, that supports the result. If purchases went up but add-to-cart went down, something unexpected may have changed in the visitor's journey.

Info

The health guardrails section appears above or alongside the results table when there is a possible data quality issue. Always check for health warnings before you decide on your experiment. See Health Guardrails for details.

Measure badge

Every metric row, primary and secondary, shows a small measure pill next to the metric name. It reads in plain English: Unique conversions per visitor, Total events, Value per visitor, Rate, or Composite (weighted). You can scan what kind of analysis is running without opening the metric's configuration. A percentile measure with a set percentile renders as p95 of value. A winsorized row (one with outlier capping switched on) adds a small · Capped suffix.

Reading ratio metric rows

A ratio metric row divides one quantity by another: average order value (revenue divided by orders) or units per order are two examples. Point estimates render in the metric's own units, for example $42.10 / order for control and $45.80 / order for a variant. Lift is a percentage, with a confidence interval built using the delta method on Frequentist and a normal-approximation posterior on Bayesian. A small badge on the row, Ratio (delta method), shows at a glance which method is in play. Revenue per visitor is not a ratio row: it is a continuous per-visitor money metric, covered in the Revenue panel section below.

A non-zero dropped count in the diagnostic means visitors with a zero denominator. They were excluded from the variant because they had no exposure to the ratio, for example average order value for a visitor with no orders. The count is shown so you can sanity-check the exclusion. See Ratio metrics for the underlying math.

Reading quantile metric rows

A quantile metric row shows the percentile in question, such as p90 of page-load time. It also shows control and variant point estimates in the metric's own units. Lift is a percentage. The confidence interval comes from a bias-corrected percentile bootstrap (1000 resamples by default). The row badge reads Quantile (bootstrap CI).

A Show both CIs toggle on quantile rows displays the bootstrap interval next to a normal-approximation one. This is useful for checking whether the two methods agree at extreme percentiles. Above p95, the row also shows the gap between two ClickHouse estimation methods: quantileTDigest (used for the point estimate) and quantileExact. A large gap warns that the faster method may be less exact at small samples.

For very large experiments, over 100,000 visitors per variant, the bootstrap falls back to a sampled subset, and the row notes approximate (sampled bootstrap). From the row you can request a high-precision recompute (5000 resamples) that runs in the background. See Quantile metrics for details.

Revenue panel

When an experiment has at least one purchase metric, a Revenue panel appears below the secondary metrics. It shows five sub-metrics: Revenue per visitor (the headline), Average order value, Purchase conversion rate, Revenue per paying visitor, and Units per order. Each sub-metric card can be opened or closed.

Reading the money rows (mean per visitor + CI)

Revenue per visitor and Profit per visitor are continuous per-visitor tests. Their cards show, per variation:

  • Mean per visitor: the average amount earned across every exposed visitor, in your project currency. Non-buyers count as 0.
  • 95% CI: the range the true mean per visitor most likely falls in. A missing interval (—) means there isn't enough data yet for that arm.
  • Lift: the relative difference in mean per visitor versus the control. This money lift is the number that decides revenue metrics. It can point the opposite way from the conversion-rate lift, when a variation drives fewer but bigger orders.
  • Significance (headed Prob. to Beat Control on Bayesian): the engine's verdict. A significance call on Frequentist and Sequential, a probability-to-beat on Bayesian.

When the default outlier cap clipped values, the card says how many extreme values it capped, and in which variation. Capping is why one huge order can't flip the verdict on its own.

"No data yet" vs "couldn't load"

The panel tells apart two states that both used to just show $0.00:

  • No revenue data yet: the query worked, but nothing is attributed yet (no exposures, or no purchases after exposure). The card explains what will appear, and points you at purchase-tracking setup. This is a normal state early in an experiment.
  • We couldn't load the revenue data: the analytics database did not respond. This is a read failure, not missing data. Your orders are safe. The card shows a Retry button instead of numbers, so a broken read is never mistaken for "the variation earned nothing".

A freshness line, "Money data updated …", appears under the panel only when the server reports a real computation time. Profit recomputation timestamps aren't available yet, so that line stays hidden for profit rather than show a made-up time.

Gross vs net toggle

The Revenue panel has a Net / Gross toggle in its header:

  • Net (default): revenue after refunds.
  • Gross: raw order totals, with refunds ignored.

Switching the toggle fetches fresh data, so every revenue figure updates together.

Covered only vs Include all (profit)

The Profit section has its own toggle:

  • Covered only (default): profit counts only orders where every item has cost data. Orders with incomplete costs are excluded, and a note says how many.
  • Include all: every order counts. Items with no cost data contribute zero cost, which can overstate profit.

A one-line caption under each toggle states which basis the numbers on screen are using.

Excluded orders note

Say an order was recorded in a currency that does not match your project currency. A vs B excludes it from statistics, and a note appears on the affected metric card:

N orders in other currencies excluded

See Revenue Deduplication and Refunds for why A vs B excludes mismatched-currency orders instead of converting them.

Recommendation performance

If a variation's code uses the recommendations API, a Recommendation performance table appears lower on the page, per variation and recipe. Impressions counts every time the variation asked for recommendations, even when none could be shown. Served counts only the times products actually came back and rendered. Click-through rate divides clicks by served impressions, and Added to cart counts distinct visitors who added a product after clicking a recommendation.

Reading composite metric rows

A composite metric row shows the weighted composite point estimate per variation, and lift as a percentage. It also shows a confidence interval built from the combined weighted variance. That variance calculation accounts for how the components move together, so correlated components don't get double-counted. The row badge reads Composite (weighted).

Info

As covered under Analysis label above, CUPED never applies to a composite row. A composite row's numbers are always the raw, unadjusted ones, whatever your project's variance-reduction setting says.

Click the Decomposition disclosure on a composite row to open a per-component breakdown. It shows each component metric's own lift, its confidence interval, and its weighted contribution to the composite. The decomposition answers why a composite moved. If the composite is up but one component is flat or negative, Decomposition makes that obvious instead of hiding it inside the headline number.

Say a component metric gets paused or archived mid-experiment. The composite row then flags component unavailable. A vs B then asks you to amend the analysis plan, through the sealed-plan amendment flow, before results recompute. The composite is never silently re-weighted onto the remaining components. See Composite metrics for the underlying math.

The decision story

Below the headline verdict, open Full decision story for a plain-English, point-by-point account of how A vs B read your experiment. It is a short numbered list covering:

  • The result and the observed lift.
  • The pre-registration status and sample progress.
  • How long the experiment ran, and whether you checked results early.
  • The traffic-split health, and the recommendation.

Every sentence ties back to a real number from your experiment.

The story has one unusual property, and it is the whole point: it cannot make anything up. Every sentence is a fixed template, filled in with a value A vs B actually computed. There is no language model anywhere in the loop. Say a fact isn't available, like a money projection before you've attached a revenue metric. That sentence is simply left out, never guessed. Each line carries a small source note naming the exact inputs it read. A sceptic can trace any sentence back to the number behind it.

Deterministic by design

The decision story, the peek counter, and every caption on this page are built the same way. Each is a fixed template, filled in with a computed value. Nothing on this trust path is written by an AI, so nothing on it can invent a number or a claim.

A few things worth knowing:

  • It narrates the engine of record. While you're exploring under a different engine, the story stays hidden. It only ever describes the official analysis, so it can never be mistaken for a second opinion.
  • A broken traffic split jumps to the top. Say the sample-ratio-mismatch check is red. The traffic-split sentence then moves to the first line. The recommendation becomes "fix the traffic split before reading anything else", matching the order the rest of the page follows.
  • Copy story puts the whole thing on your clipboard as plain text, ready to paste into a decision log or a ticket.

Your trust grade

Every experiment's official results carry a trust grade from A to F. It is a single letter answering "how much should we trust the way this experiment was run?" It grades the process, not the outcome. A test that loses cleanly can grade A. A test that wins sloppily can grade D.

The grade is a 100-point score, built from five parts:

ComponentPointsWhat earns them
Pre-registration & adherence25A plan sealed before launch earns 15 (sealed after launch: 8; no sealed plan: 0). The remaining 10 are for sticking to it: each amendment after sealing costs 5, down to 0.
Sample vs plan20Reaching your planned sample earns 20; ≥75% earns 12; ≥50% earns 8; less earns 4. No planned sample set: 5 (we can't verify a plan that was never stated).
Peeking discipline20On the Sequential engine every look is safe, so it always earns 20. Otherwise the peek counter decides: at most 1 look earns 20; 2–3 looks 14; 4–6 looks 8; 7 or more 2. A recorded early stop under Frequentist earns 0: that is the realized harm the count only estimates.
Traffic integrity25The SRM check: green 25, yellow 12, red 0.
Runtime & sample health105 for running at least 7 full days (a whole weekly cycle; 2+ days earns 2), plus 5 when the smallest arm has healthy traffic (low: 2).

Letters band by score: A ≥ 90 · B ≥ 75 · C ≥ 60 · D ≥ 40 · F below 40.

Two hard caps sit above the score:

  • A red SRM caps the grade at D, whatever the score. A broken traffic split poisons every number downstream. Acknowledging the warning reveals the verdict, but it does not lift the cap: the mismatch is still a fact.
  • A "too good to be true" lift caps the grade at B. Say the leading variation's observed lift on the primary metric is ±25% or more. The grade is then capped at B and flagged, per Twyman's law. In real experiments, a move that large on a primary metric usually means a tracking or bucketing problem, not a miracle. The flag cites the observed lift on purpose: a statistical adjustment cannot launder a tracking bug. Verify your tracking before you ship.

When both caps apply, the lower one (D) wins.

There is no manual override. The grade comes from recorded facts only: plan timestamps, the peek ledger, the SRM test, visitor counts. To dispute a grade, fix the underlying fact instead: re-run with a sealed plan, fix the split, reach the target. An amendment still deducts points even with a good reason, but the reason is shown alongside so a reader can judge it. That is what makes the letter worth trusting.

The grade appears as a chip in the Decision Header. Click it for the full component breakdown. It only appears on official-engine views: exploring under another engine never produces a grade.

Honest lift

Selected winners look bigger than they really are. This is the winner's curse: you only ship a result because it crossed your significance bar. Among everything that crosses a bar, the lucky overestimates are over-represented. The smaller your sample, the worse the exaggeration. A barely-significant result can easily show double its true effect, or more.

When your result crosses its engine's decision boundary, A vs B adds an honest lift caption under the headline:

Observed +12.0% is likely optimistic. Selected winners overstate their true effect. Post-selection estimate ≈ +3.0%; ≈8.6% of results selected like this one would point the wrong way.

  • The post-selection estimate is a winner's-curse-corrected estimate of the true lift, using a method called a conditional maximum-likelihood estimate. See the methodology page for the exact math. It is an estimate, not a guarantee, but it is a far fairer number to put in a forecast than the observed one.
  • The wrong-way number uses a method called the plug-in Gelman-Carlin Type-S error, calculated at the adjusted estimate. Among results selected exactly like yours, this is the share that would point in the wrong direction if the true effect equals the adjusted estimate. Near the significance boundary it approaches ten percent, which is worth knowing before you ship.
  • On a well-powered result, the caption flips to reassurance: the observed lift is a fair estimate, with the (barely different) adjusted number shown alongside.

The caption appears only on official-engine views, and only when the result actually crossed the boundary: an unselected result has no selection to correct. If a multiple-comparison correction applied to your metric family, the caption notes the adjustment is at least this large. The exact corrected boundary cannot always be inverted cleanly, so A vs B discloses a lower bound rather than fake precision.

Hidden when an informative prior is active

Say you have set an informative prior for this metric on the Bayesian engine. The honest lift caption does not appear at all. A vs B's prior already pulls an early, lucky-looking result back toward reality. Adding this correction on top would discount the same exaggeration twice.

The same two numbers ride the CSV export as honest_lift_pct and wrong_sign_risk, on the leading variation's primary-metric row.

Was this helpful?