Analysis plans & pre-registration
A pre-registered analysis plan is a short declaration you write before an experiment starts. It names the primary metric, the guardrail metrics, the engine, the win bar, and how big the experiment is meant to get. A vs B locks the plan the moment the experiment launches, and records every change made after that. So the gap between what you declared up front and what you decided later is always visible. You can see it on the results page and in the audit log.
Why it matters
The danger most teams underestimate in experimentation is researcher degrees of freedom: the small after-the-fact choices that quietly inflate false-positive rates. Adding a metric after the fact because it's the only one moving. Picking a different segment because it tells a better story. Switching engines because the official one doesn't reach significance.
Each of these is reasonable on its own. Together, they erode the credibility of every result.
Pre-registration is the cheapest, most effective safeguard against this. Ron Kohavi's Trustworthy Online Controlled Experiments is the standard reference for the field, and it makes a strong claim. Locking the analysis plan before data starts arriving, it argues, is the single most impactful thing a team can do to keep results trustworthy. A vs B builds pre-registration into the experiment lifecycle itself. It is not a side convention teams have to remember on their own.
Pre-registration is a per-project toggle. Open the project's Analysis tab in Project Settings. Turn on Require pre-registration before launch when you want every experiment in the project to launch with a sealed plan. Leave it off and individual experiments can opt in selectively.
Declaring a plan
The experiment builder has a dedicated Analysis step (between Metrics and Review). The step holds two cards, and the split matters:
- The Analysis card is where you choose how the experiment is judged: engine, win bar, correction, variance reduction. Each field inherits from your project defaults, so you only have to think about what differs from the norm.
- The Analysis plan card is what you pre-register: which metrics count, how big the experiment should get, and how long it should run.
Each setting is editable in exactly one place. The Analysis card is the only place the engine and win bar are set. The plan records the choice, rather than offering a second copy of the control.
On the Analysis card:
- Stats engine. Bayesian, Frequentist, or Sequential. Defaults to the project's configured engine.
- Confidence level. Sets α for Frequentist and the always-valid α for Sequential. For a Bayesian experiment it instead sets the width of the credible interval; the win bar itself lives in a separate field, described next.
- Bayesian decision threshold. Shown only when the engine is Bayesian. The chance-to-beat-control an arm must clear before it's called a winner (95% by default, labelled Chance to beat Control on the field itself). This is a different setting from confidence level above: the two happen to default to the same number, but changing one never moves the other.
- Multiple-comparisons correction. Shown only when the engine is Frequentist. Tiered (the recommended default), Bonferroni, Holm-Bonferroni, Benjamini-Hochberg, or none. It only changes anything once an experiment has more than two variations; with just Control and one variant there is nothing to correct for.
- Variance reduction. Auto applies CUPED when your metric has a pre-experiment signal to borrow from, and skips it when that wouldn't help. Off always reports the raw number.
- Effect prior. Shown only when the engine is Bayesian. An optional informative prior on the relative lift: what you expect before any data arrives (for example, "expect 0%, with most real effects within ±30%"). Off by default, which is a flat, uninformative prior.
On the Analysis plan card:
- Primary metric. The single metric the experiment is meant to move. Required.
- Secondary metrics. Anything else you want to track but don't want to gate the decision on.
- Guardrail metrics. Metrics that must not get worse: typically conversion rate, revenue, error rate, or load time. Optional but strongly recommended. A metric is either a secondary or a guardrail, never both: picking one moves it out of the other.
Initialize plan starts from the roles you set on the Metrics step: the main goal becomes the primary, metrics marked as guardrails become guardrails, and the rest become secondaries. The card always shows the metrics attached on the Metrics step, even if you came straight from there. If a value is refused, for example a target sample size below 1, the card names the field and the rule.
- Target sample size. How many visitors, in total, the experiment should reach. Type it in yourself. Or follow Estimate with the sample-size calculator, work out the number from a baseline rate, MDE, and power, then click Use as target to write it onto the plan for you. No copying numbers back by hand.
- Target duration (days). The expected runtime. Used for the day-counter in the results header.
The Review step then shows the whole thing back to you as one line: engine, win bar, correction. Next to it sits the primary metric, the target, and the plan's state: Sealed, Drafted (it seals when you publish or schedule), or Not required. It's your last chance to catch a setting that isn't what you meant.
See Choosing a stats engine for the engine decision and Early stopping for what the target sample size means in practice.
Sealing
The plan is sealed the moment the experiment launches. From that point on:
- The plan's fields become read-only in the builder.
- The results page surfaces the sealed plan in a dedicated card.
- Each metric on the results page is tagged Pre-registered or Exploratory, depending on whether it was on the plan when sealing happened.
- The official stats engine, the confidence level, and (for Bayesian) the decision threshold are locked. The Explore-under and Compare engines views still let you re-render under a different engine. Those views are clearly labelled exploratory, though, and they never overwrite the sealed result.
The sealed plan is what drives the numbers. At launch, the plan's engine, confidence level, multiple-comparisons correction, and variance-reduction setting are copied onto the experiment itself. For a Bayesian experiment, the decision threshold and the effect prior are copied too. The sealed plan card on the results page shows each of these back to you.
So the bar you committed to calling a winner at (for example, "95% probability to win") is on the record. So is the prior you started from, not just in your head. Every result you read afterwards is computed from that copy. Not from the plan record as it stands today, and not from your project defaults as they stand today. That is what makes the guarantee real rather than cosmetic. Once a test is running, editing a project default cannot retroactively change how an already-launched experiment is judged.
While the experiment is still a draft, the plan follows what you choose. Change the engine on the Analysis card after you've declared a plan, and the plan updates with it. That way, whatever gets sealed at launch is the last choice you actually made. It is never a stale value you'd already replaced. The plan stops following the moment it seals.
A draft experiment can be edited freely. Once you launch, the plan is sealed. That's intentional, so that "adjusting the plan" mid-flight requires going through the amendment workflow rather than quietly rewriting history.
Amendments
Sometimes a sealed plan has to change. A guardrail you forgot to declare turns out to matter. The target sample size needs to grow because traffic is lower than estimated. A new variation gets added. A vs B handles this through an explicit amendment workflow, instead of letting fields silently change.
To amend a sealed plan:
Open the plan card
Open the experiment's Analysis plan card on the results page.
Propose amendment
Click Propose amendment. A modal opens with the current value of every field; you change only what you mean to change. Metric fields are picked by name from a checklist, so there are no identifiers to copy.
Add a reason
Add a short reason: one line is fine. The audit log captures who, when, what changed, and why.
Submit
Submit. The amendment is recorded straight away and becomes visible in the plan card and in the experiment's audit history.
Amending a running experiment's plan does not change its results. The numbers stay pinned to the analysis the test launched with. This is the entire point of sealing: if loosening a setting after seeing the data could move the verdict, the seal would protect nothing. An amendment records that your intent changed and why. It does not re-judge data that has already been collected under the original plan.
So the plan card on the results page keeps showing the launch values. Any field whose current intent has since diverged gets an Amended badge. The amended values themselves live in the amendment history, one click away. If you genuinely need a different analysis to drive the decision, run a new experiment.
Every amendment is preserved indefinitely. The full chain is exported as part of the experiment's record: the original sealed plan, plus every amendment. Each one carries its timestamp, actor, before/after values, and reason. Download it from the audit log or the per-experiment export.
A metric referenced by a sealed plan cannot be deleted or archived: the action is refused and the error names the blocking experiments. That protection exists because the detachment would bypass the amendment workflow above, and unarchiving would not re-attach the usage. Amend each blocking plan first, then delete or archive. Changing a sealed-plan metric's definition (its measure, direction, weights, or winsorization) is refused the same way until the plan is amended.
Pre-registered vs exploratory
On the results page, every metric is tagged with one of two pills:
- Pre-registered: the metric was on the sealed plan at launch.
- Exploratory: the metric was not. Useful for poking around your data, but not appropriate for drawing conclusions you'd defend in a stakeholder review.
A metric added by amendment after launch stays Exploratory. That's for the same reason amendments don't move the numbers: it wasn't declared before the data arrived. Treating it as pre-registered would launder an after-the-fact choice into a pre-registered one. The amendment is still recorded, and the metric is still measured. It just doesn't get to claim it was predicted.
The visual distinction is deliberately subtle in the UI (pills, not warnings), but the information is always there for anyone reviewing the result. When in doubt, pre-register.
The project-level toggle
Open the project's Analysis tab in Project Settings. The Require pre-registration before launch toggle sits in the Analysis defaults section. Turn it on and every experiment in the project must complete the analysis-plan card before it can launch. With it off, declaring a plan is opt-in per experiment.
Most teams start with the toggle off. They treat pre-registration as a habit on important experiments, then turn it on once the team is fluent with the workflow. Regulated environments (pharma, finance, compliance-sensitive product areas) typically turn it on from day one.
When you'd want this
- Multi-stakeholder reviews. A pre-registered plan eliminates the "but didn't we say we cared about X?" debate after results land.
- Regulated environments. An auditable record of what was decided in advance is non-negotiable for some compliance regimes. The amendment log makes this straightforward.
- Experiments where the team has a strong prior. The temptation to rationalise a non-result is strongest when you really wanted the variation to win. Pre-registering keeps you honest with yourself.
- Reusable templates. Once a team agrees on its standard analysis plan for a given surface (checkout, signup, onboarding), pre-registration becomes a checklist rather than a debate.