Comparing Engines

A vs B locks the official stats engine at launch: that's what shows up in audit logs, exports, and the experiments-list summary. But on the results page itself, you can re-render the analysis under any engine, any time, without changing the official record.

What "Explore under" does

On every experiment's results page, the engine badge is followed by an Explore under dropdown listing all three engines (Bayesian, Frequentist, Sequential). The currently-official option is marked (official) and is selected by default.

Picking a different engine refetches the results from raw exposures and metric events, runs the analysis under the engine you picked, and re-renders the entire results surface. The engine badge updates to reflect what you're looking at, and a yellow Exploratory view banner appears below the header so it's never ambiguous which engine the numbers came from.

Click Reset to official on the banner (or pick the official engine in the dropdown) to go back to the locked-in view.

The official engine never moves

Switching the dropdown does not change anything in the database. The experiment's official engine, audit log, exports, alerts, and summary numbers all continue to reflect the engine that was set at launch. Explore-under is purely a viewing affordance.

When to use it

  • Cross-checking a result. A Bayesian "87% probability variant beats control" is reassuring; seeing the same data under Frequentist with p < 0.05 can confirm the result is robust to the choice of engine.
  • Talking to stakeholders. A statistically-curious stakeholder asks "but what would the p-value be?" The Frequentist view answers without re-running the experiment. A risk-averse stakeholder asks "is this safe to stop early?" The Sequential view answers that one.
  • Demonstrating engine differences. Pick an experiment where the methods might disagree (small effect, modest sample size) and walk through the three views side by side. Useful for onboarding and for justifying engine choices on future experiments.

Caveats and exploratory caveats

Explore-under is exactly that: exploratory. The recomputed numbers are mathematically valid for the engine you picked, but a few subtleties are worth keeping in mind:

  • Engine-specific defaults apply. When you explore under a different engine, the engine-specific configuration from your project settings applies: alpha defaults to 0.05 unless your project sets otherwise. Per-experiment alpha and MCC overrides also flow through if they're set.
  • Sequential explored retrospectively is not the same as Sequential by design. The always-valid guarantee is mathematical and requires the analyst to commit to Sequential before the experiment starts. Re-rendering a Bayesian-launched experiment under Sequential is a useful comparison, but the "safe to peek" property is a property of the original analysis plan, not a re-labelling.
  • The summary numbers in the experiments list don't change. The lift-percentage and conversion totals shown on the experiments list are mirrored from the official engine's analysis. Exploratory views do not overwrite them.
  • The results page shows a note in place of the verdict while you're exploring. The verdict lives on the official view only, so it never gets replaced or half-shown under a different engine. One control on the note takes you straight back to the official view where the verdict reappears.

Compare engines, side by side

Explore-under shows one engine at a time. The Compare engines button (also in the results page header, right next to the Explore-under dropdown) opens a panel that renders Bayesian, Frequentist, and Sequential results in three columns at once, computed on the same underlying data.

Each column speaks its engine's native vocabulary:

  • Bayesian: the chance the arm is genuinely better, plus its credible interval.
  • Frequentist: classical p-value, 95% confidence interval, and a significance tag. The column always names the correction it used for extra comparisons. The default is Tiered: Bonferroni on the primary metric, Benjamini-Hochberg on the rest. Your project or experiment can pick Bonferroni, Holm, Benjamini-Hochberg, or none instead.
  • Sequential: always-valid p-value, 95% confidence sequence, and a "safe to stop" or "not yet conclusive" label.

The interval in each column is that arm's own value (where its conversion rate or point estimate sits), not the size of the gap to control. The columns label it as such.

Percentile (p95-style) metrics get an estimate and an interval only: no engine produces a p-value or a win probability for a percentile, so those cells read "—" rather than showing a number that doesn't mean anything.

Lift is always goal-aligned: positive means better, on every metric. If your metric's goal is to go down (bounce rate, load time, refunds), an arm that cut it by 20% reads +20%, and the column header says so. This matches the results page exactly.

The official engine's column is anchored first and given the visual weight, labelled Engine of record. The other two are styled as muted cross-checks. Engine colours are identity only: no engine is coloured as a pass or a warning.

Every column also carries an i button next to its name. Click it for a plain-English description of that method and what it's good for.

The agreement banner

At the top of the panel, A vs B computes whether the methods actually agree and says so in a sentence: "All three methods agree Variant B beats control", "2 of 3 methods call B a winner; the other is not yet conclusive", or "Methods disagree on which variation is winning: treat with caution."

Two details worth knowing:

  • It only counts engines that had data. If one engine had no traffic to analyse, or its analysis failed, the banner says "Both methods", never "All three": an engine that never ran agreed to nothing. A failed column says so plainly rather than showing zeros.
  • A winner has to be a winner at your metric's goal. On a lower-is-better metric (bounce rate, latency, refunds), an arm can be statistically significant because it made things worse. The banner never calls that a winner.
Switching engines to get a friendlier answer is p-hacking

Your decision stands on the official method. If the official engine says "not yet conclusive" and a cross-check says "significant", the honest reading is not yet conclusive. Picking whichever of three methods happens to agree with you inflates your false-positive rate: you get roughly three chances to find a winner that isn't there. The cross-checks are there to tell you how robust a result is, not to give you a second opinion to act on.

Why do the methods differ?

When the methods don't fully agree, A vs B adds a short, plain-English explanation of why (something only possible because all three engines are computed on the same screen at once). Like the decision story, these are fixed templates keyed to real numbers, never generated prose: at most three are shown, in priority order, and when no specific cause can be identified A vs B says so rather than inventing one.

The causes it can name:

  • Peeking vs. a fixed horizon. The Frequentist p-value assumes a single look at the planned sample. If you've checked early, or you're not yet at the planned sample, Sequential's wider "not yet" is the more honest read, and this note says so.
  • One-sided vs. two-sided. Bayesian's "probability B is better" answers a one-sided question; a p-value answers the two-sided "is B different?". Near the decision boundary those two framings genuinely disagree, and the note shows the two-sided p next to its one-sided equivalent.
  • Multiple-comparison correction. With more than two variations, the Frequentist p-value is corrected for the extra comparisons (Bonferroni, Holm, or Benjamini-Hochberg); Bayesian and Sequential apply no such correction here. When a raw-significant result was corrected away, the correction itself is the disagreement.
  • Small sample, flat prior. At small samples the Bayesian probability moves faster than a p-value on the same data. The note names the smallest arm's size and reminds you these tend to converge as data accrues.
  • Sequential is wider by design. Sequential's always-valid interval is deliberately wider early (the price of being safe to read at every look), so it typically confirms later than the fixed-horizon methods on identical data.

If none of these apply, the note says plainly that no single cause was detected and points you back to the engine of record. Each explanation carries a small source affix naming the numbers it read, the same "cannot make things up" guarantee as the decision story.

Which verdict should I trust?

The panel works this out from how you actually ran the test, and says so under the banner:

SituationWhat A vs B tells you
Traffic split looks broken (the sample-ratio check on the results page is red)Trust nothing yet. The agreement verdict is greyed out: the methods may be agreeing about invalid data. Fix the split first.
Fewer than 100 visitors in the smallest variation (the split itself is fine)Trust no engine yet. A caution, not a block: keep the test running until each variation has more visitors. The banner never calls this a broken split.
Frequentist official, still running, planned sample not reachedTrust Sequential for an early read. A fixed-horizon p-value read early is optimistic.
Frequentist official, planned sample reached (or the scheduled end date passed)Trust Frequentist. It's at its horizon and is the result of record.
Frequentist official, stopped before the planned sampleFrequentist is still of record, but its p-value is slightly optimistic; cross-check Sequential before acting.
Sequential officialTrust Sequential. It's safe to read at any time. The Frequentist column here is retrospective and slightly optimistic.
Bayesian officialTrust Bayesian, against its decision threshold. The p-value columns are cross-checks.

A/A tests and still-running tests

On an A/A test the reading is inverted: the methods finding no winner is the pass, and a flagged difference usually points at a setup or splitting problem rather than a real effect. The banner says which of the two you're looking at.

While an experiment is running, the panel notes that the readings can still change. Once it's stopped or completed, the numbers are final and the note disappears.

Provenance

A line above the columns states when the numbers were computed, shown as "Data as of … (same window and filters as the results page)", so the panel is never quietly older than the page behind it. The columns honour the same date range, segment filters, and revenue settings as the results page.

Compare engines runs the analysis pipeline once, on a single fetch of the underlying raw data, and dispatches all three engines on it, so opening the panel does not triple your ClickHouse load compared to a normal results-page load. Like Explore-under, it never changes the official engine of the experiment.

Where to find it

Open any experiment's results page. The Explore under dropdown sits next to the engine badge in the header, with the Compare engines button right after it. Both are available for any experiment status (draft, running, paused, completed, archived) as long as you have viewReports permission for the org.

Was this helpful?