Frequentist Engine
The Frequentist engine is the classical approach to A/B test statistics: a p-value, a 95% confidence interval, and a plain yes-or-no answer on whether a result is statistically significant (real, not just noise). It's the format most stakeholders and compliance reviewers expect to see.
What Frequentist analysis actually does
Frequentist analysis asks one question: if there were really no difference between the variations (one specific version being tested, like Control or Variant A), how likely is it we'd see a result at least this extreme? That probability is the p-value. A small p-value means the difference you saw would be surprising if nothing had really changed, so you can treat the result as significant.
In A vs B, the Frequentist engine uses:
- A two-proportion z-test for binary metrics (conversion rates). It uses a pooled standard error against Control to compute the p-value, and unpooled Wald intervals for the confidence interval you see on screen.
- A Welch two-sample z-test for continuous metrics (revenue, pageviews, time on page). It uses unpooled variance and a normal-approximation p-value.
Reading a p-value
By default A vs B calls a result significant when p < 0.05. That threshold is the significance level, written as α = 0.05. You can change this default for the whole project, and override it per experiment, in Project Settings and the experiment builder's Analysis step. See Analysis Defaults.
A few common values in plain English:
p = 0.03→ a 3% chance you'd see this much difference by pure luck. Significant at α = 0.05.p = 0.12→ a 12% chance under pure luck. Not significant at α = 0.05.p < 0.001→ extremely unlikely to be luck. Very strong evidence of a real difference.
Reading the 95% confidence interval
Alongside every p-value, A vs B reports a 95% confidence interval for each variation's observed rate or mean. A confidence interval is a range built so that, if you repeated the experiment many times, the true result would land inside it roughly 95% of the time.
On the results page, the confidence interval shows as a horizontal bar. The wider the bar, the less certain the estimate. As more visitors are exposed (counted in the experiment, usually because they saw the part being tested), the interval narrows.
Multiple variations and corrections
When an experiment has more than two variations (say Control, Variant A, and Variant B), A vs B runs more than one test. Each extra comparison raises the chance that one of them looks significant purely by luck. A vs B corrects for this automatically, using the multiple-comparison correction (MCC) method set on the experiment or project. Tiered is the platform default:
- Tiered (default, recommended): your primary metric, the one metric an experiment is actually judged on, is corrected across variations using Bonferroni. Every secondary metric, tracked for context but never the deciding factor, is corrected together using Benjamini-Hochberg, which controls the false-discovery rate across the whole secondary set. The idea: protect your one pre-declared hypothesis at close to full strength, while still guarding against a long list of secondaries manufacturing a false win.
- Bonferroni: simple and conservative. Multiplies each p-value by the number of comparisons.
- Holm-Bonferroni: slightly less conservative than plain Bonferroni.
- Benjamini-Hochberg: controls the false-discovery rate across every comparison. A good pick when you're tracking many metrics or variations and Tiered's split isn't what you want.
- None: reports raw p-values with no correction. Only use this if you understand the tradeoff.
The p-values shown on the results page are the corrected ones, and the significance badge reflects the correction.
Some older experiments were set to Tiered Bonferroni, a different method from today's Tiered default. Tiered Bonferroni leaves the primary metric completely uncorrected and gives secondaries full Bonferroni. It's no longer offered when you pick a new method, but a stored experiment that already uses it keeps working exactly as before.
Correcting your primary lightly only makes sense when you picked it before you saw the data. If you choose your primary after looking at which metric moved, you've undone the protection. Analysis plans exist to make that pre-declaration real rather than remembered.
Why peeking matters
A Frequentist p-value is only valid once, at the sample size you planned in advance. If you check the numbers every day and stop the moment the result looks good, the real false-positive rate climbs, often from 5% to 15 to 20%. Stopping on the day that happens to cross the threshold is a form of p-hacking (checking your data early and often, then acting on whichever moment looks best), even when you don't mean to do it.
Use the sample-size calculator to decide how many visitors each variation needs before you launch. Don't call the experiment early just because the p-value dipped below α for a moment. A vs B warns you if you try: see Early stopping & peek protection.
If you need to check results safely and often, pick the Sequential engine instead. It's built for exactly that.
Peek protection
While a Frequentist experiment is still collecting its planned sample, A vs B shows a banner at the top of the results page: "Day 5 of 14 · 3,400 of 10,000 visitors · not yet valid for stopping decisions." It disappears the moment you reach your target sample size or your scheduled end date.
If you try to pause or stop a Frequentist experiment before it reaches its target, a modal interrupts the action and offers three choices:
- Let it run: dismiss the modal and keep collecting visitors.
- Stop (or Pause) anyway and log: go ahead with the action. A vs B stamps the experiment with an audit-visible badge reading "Early-stopped · validity reduced," logs an
EARLY_STOPPEDentry to the audit log, and adds an "Early stop" section (with your reason, if you gave one) to any CSV export of the results. - Switch future experiments to Sequential: changes your project default. The current experiment continues unchanged; new experiments will default to Sequential. This choice is disabled when your project default is already Sequential, or on a screen where Sequential isn't offered.
The same peek check applies whether you Pause, Stop, or use Archive from the experiments list to end a running experiment in one step: all three are gated the same way before your target sample size is reached. See Early stopping & peek protection for the full reasoning and what the audit trail looks like.
When to pick Frequentist over Bayesian
- A stakeholder wants p-values. Many product, marketing, and research teams speak in p-values. If your org already reports results this way, Frequentist is the simplest fit.
- Regulatory or compliance reporting. Pharma, finance, and healthcare workflows often need a fixed-α, pre-declared-sample-size design. Frequentist maps directly onto that.
- You can commit to a sample size up front. If your traffic is predictable and you're willing to plan the experiment's length before launch, Frequentist is efficient and easy to explain.
Where to configure
Set the engine in one of three places:
- Project level. Project Settings → Analysis sets the default for every new experiment and feature-flag A/B test rule in the project.
- Per experiment. In the experiment builder's Analysis step, before launch. Once the experiment is running, the engine is locked: switching engines mid-flight would invalidate the result.
- Per flag rule. In the feature-flag A/B Test rule editor's Analysis section, before the rule is turned on. Same lock-on-launch behaviour as experiments.
Your significance level (α, above) only governs Frequentist and Sequential results. It's stored separately from, and never used to calculate, the Bayesian engine's own decision threshold (the probability-to-beat-control an arm must clear to be called a winner). Changing one never silently moves the other.