Segment Lift
A variation is one specific version being tested: control, or a challenger. The Segment Lift section on the Results page answers a question the overall lift number cannot. For whom did your variation help? For whom did it hurt?
Segment Lift breaks your primary metric (the one metric your experiment is judged on) down across every segment you have. It shows the lift inside each segment value. It also corrects for a hidden fact: you are really running many small tests at once, not one. A top-movers card names the segments that helped most and hurt most.
What is Segment Lift?
A new checkout flow might lift overall conversion by +3%. Inside that one number, mobile users could be up +15% while desktop users are down −6%. The single number hides both stories. The choice is no longer "ship or kill". It becomes "ship to mobile, keep the old flow on desktop, and find out why".
The stats term for this is a heterogeneous treatment effect (a real difference between groups, hidden inside one average lift). Segment Lift makes those differences visible. You don't have to click through every segment by hand. A vs B computes every one in a single pass.
How A vs B computes per-segment lift
For each available segment on the experiment, such as device, country, browser, or any custom segment you send, A vs B runs one ClickHouse query. That query groups exposures (visitors actually counted in the experiment) and conversions by segment value and variation. Each segment-value-and-variation cell then goes to the same stats engine that powers your primary metric card. So the Segment Lift numbers are computed exactly the way the headline lift is, just scoped to a smaller group of visitors.
- Per-segment query: one ClickHouse query per segment key, run in parallel.
- Per-cell engine call: your engine (Bayesian, Frequentist, or Sequential) scores each cell. It reports lift, significance, and a confidence interval (a range likely to hold the truth).
- Interaction test: a chi-square test checks whether the differences between segment values are real, not noise.
- Correction: every p-value (the odds of seeing a result this extreme if nothing really changed) in the panel is corrected together. The method is called Benjamini-Hochberg, and it keeps the significance flags trustworthy.
If you already have a date range or segment filter active on the Results page, the panel's header says "Computed within your current filters." That line means exactly what it says: every cell in the panel is computed from the same narrowed group of visitors you're currently looking at, not from the whole experiment. Clear your filters to see Segment Lift computed against everyone.
Segment Lift needs your primary metric to be a conversion metric: one that counts whether each visitor converted, yes or no. It does not run when the primary metric measures revenue or another value, counts total events, or is a ratio, percentile, or composite metric (several metrics combined into one score). Per-segment conversion rates cannot describe those metrics, so instead of numbers the panel shows a short note saying the metric is not supported. Switch your primary metric to a conversion metric, or read the segment breakdown from the segment filter instead.
Why multiple-testing correction matters
Say you run 20 segment-by-segment tests at a 5% significance level. You should expect one false positive by chance alone, even if no segment really differs from another. This is the segment-shopping problem: dig through enough segments and you will always find a "significant" one. Without correction, the segment dropdown becomes a tool for fooling yourself.
A vs B applies the Benjamini-Hochberg procedure across every cell in the panel. It controls the false-discovery rate: of the cells flagged significant, the share that are false positives is held at your chosen significance level.
The Segment Lift table shows this adjusted p-value (the odds of seeing a result this extreme if the segment actually changed nothing) as p_adj. That is the number to read. A cell that looks significant before correction will often disappear after it. That is correct behaviour, not a bug.
When the top-movers card says no segment cleared the reliability bar, it means nothing survived this screening at your significance cutoff, called alpha (0.05 unless your plan set another value). Apparent per-segment differences below that bar are within the range chance alone produces.
Reading the interaction p-value
Each segment block carries an interaction badge ("Strong interaction", "Possible interaction", or "No interaction") with its own p-value. This answers a different question from per-cell significance:
- Per-cell significance asks: "Did variation B beat control inside this one segment value?"
- Interaction significance asks: "Did the treatment effect differ across segment values? Did mobile really respond differently from desktop, or could that gap just be noise?"
Check interaction before you personalize on a segment. A "strong interaction" badge means the gap between segment values is unlikely to be chance. A "no interaction" badge means the segment slices look much the same, even if one cell happened to clear the bar on its own.
Sometimes the test cannot run at all: a segment with too little data, or one whose values saw no conversions. The badge then reads "Not testable" with no p-value. That is different from "no interaction". "No interaction" is a tested answer; "not testable" means the question could not be asked.
Sample-size policy
Tiny cells produce wild lift numbers. A vs B applies two limits to keep the panel honest:
- Below 30 visitors per arm: the cell is dropped from analysis. Per-cell lift shows as
—. - Between 30 and 1,000 visitors per arm: the cell renders with a Low N badge. You still see the estimate, but treat it as directional, not conclusive.
Engine-specific behaviour
The Segment Lift table follows whichever engine you have selected for the main Results page.
- Frequentist and Sequential: the significance column shows the adjusted p-value,
p_adj. A cell is flagged significant whenp_adjis below your significance level. - Bayesian: the significance column shows the probability to beat control. A cell is flagged significant when that probability is above 85%, or below 15% for a strong loser.
That 85% bar is fixed. It does not move with your experiment's confidence level, on purpose. Segment Lift cells are exploratory leads, not a second decision next to your primary verdict. So tightening your primary confidence level never quietly hides a segment finding.
Frequentist and Sequential keep their Benjamini-Hochberg correction. Bayesian cells skip it: a Bayesian probability does not need the same false-discovery-rate fix.
What this is, and isn't, for
Segment Lift is a discovery signal, not a deployment decision. A flagged cell is a hypothesis worth retesting. It is not a green light to ship a per-segment rollout from the panel alone. The right follow-up to an interesting finding is usually a new experiment aimed at that segment. It could also be a follow-on rollout, decided by a person, outside the panel.
The segment filter is the right surface for drilling into one slice once the panel has flagged it. Use Segment Lift to find candidates. Then filter to that slice to see its full breakdown: charts, secondary metrics (tracked for context, not for the decision), and the time series.
The 100-cell cap
Per-segment fan-out can grow large fast. A project with hundreds of country values could take over the panel. To keep it readable, and to keep the query cost in check, A vs B caps each experiment at 100 cells.
When the cap kicks in, the panel footer says so. A vs B then keeps segments in order of total visitor count. The most-trafficked segment values survive.