Effect priors
An effect prior is what you believed before the experiment ran.
Most A/B tests do not move the needle much. Say you have run many tests and almost all of them landed between −5% and +5%. A result claiming "+80%!" after fifty visitors is almost certainly noise, not a miracle. An effect prior lets the maths know that. Early flukes get pulled toward reality instead of being announced as wins.
The Bayesian engine is the only engine that uses it. It is off by default: if you never touch it, every number stays exactly what it would have been without this feature.
The two settings
| Setting | Plain meaning | Suggested starting point |
|---|---|---|
| Expected effect | "On average, I expect tests to do about this" | 0% |
| Effect spread | "Most real effects land within about ± this much" | ±30% |
Both settings work on the relative lift (the percentage change), so one setting genuinely serves every metric. A ±30% spread means the same thing for a signup rate, revenue per visitor, or load time.
They travel together: setting one without the other gets rejected. A prior needs both a location ("what I expect") and a scale ("how sure I am").
You can set them:
- Project-wide: Settings → Analysis defaults → Bayesian settings. Pre-fills every new experiment and flag rule.
- Per experiment: the builder's Analysis step.
- Per flag rule: the rule's analysis overrides.
Like every analysis setting, the prior is frozen at launch. Changing the project default later never re-judges a running or finished test.
What changes when it is on
With a prior active, the Bayesian engine judges the contrast (the lift) under a posterior (the belief you get once your guess meets the data). It blends your stated expectation with the data:
- The winning probability becomes the posterior probability under the prior. A "+80% at 96% probability" from tiny traffic reads as far more modest, because effects that large are rare under your stated belief.
- The headline lift becomes the posterior estimate, shown as "Estimated Lift (prior applied)" with the observed lift disclosed alongside it.
- The lift interval becomes the posterior credible interval (a range that says how likely the true result is to fall inside it).
- Per-arm conversion rates and their intervals stay data-only. The prior is a belief about the difference, not about either arm's absolute level.
Disclosure is unconditional. The "Judged by" line names the prior whenever it is active (e.g. Bayesian · win prob ≥ 95% · prior N(0%, 30%)). The decision header also says in words when the prior moved the headline. For example, it might read: "the observed +80.0% reads as +6.2% once weighed against that expectation."
A prior that changed the answer without saying so would be worse than no prior at all.
One thing worth knowing: with a prior active, the results page no longer shows the separate post-selection ("winner's curse") estimate. The two are alternative cures for the same problem (exaggerated early winners), and using both would shrink the number twice. The prior version is stronger: it corrects the headline itself, not just a caption under it.
Enough data always wins
The prior's weight fades as evidence builds up. The posterior estimate is a weighted average of what you expected and what you observed, and the data's weight grows with sample size. A genuinely large effect still gets found; it just has to earn that with evidence, not a lucky first week. This washout property is what makes the prior safe: it dampens noise, not signal.
Our calibration suite re-proves this on every test run. Under the prior's own assumptions, the 95% interval covers the truth about 95% of the time. With enough traffic, even a wrongly placed prior returns to near-nominal coverage.
The honest warning
A badly chosen prior can hide a real win. Say you assert "effects are within ±2%" but reality delivers +15%. The posterior will hug zero until a lot of data piles up. Our calibration measures interval coverage collapsing to roughly 1–2% under that kind of dogmatic misspecification. This is not a flaw to fix; it is what "I am very sure effects are tiny" means.
Practical guidance:
- Prefer a wide spread (±20–30%) unless you have real history saying otherwise. A wide prior tames wild early swings while barely touching honest results.
- 0% expected effect is almost always right. If you knew the direction, you would not be testing.
- Do not tighten the spread just to make a result look good. The prior is frozen at launch specifically so nobody can steer a live test with it.
- If a prior-on result surprises you, the observed lift is always shown next to the posterior. The raw data never disappears.
When not to use a prior
Leave it off if:
- You are testing something you have never tried before, like a brand-new feature or a new market. A prior needs real history behind it; without one, an "expected effect" is just a guess.
- You need the raw, unshrunk number for an outside report that expects the flat-prior calculation instead.
- You are tempted to narrow the spread just to make a live result look more convincing. That is steering the test, not stating a belief, which is exactly what the launch freeze exists to stop.
- Expected effect is what you believe a typical test does to this metric, on average. Leave it at 0% unless you have real history saying otherwise.
- Effect spread is how wide a range you think most real effects fall within. A wide spread (±20–30%) is the safer starting point.