Composite metrics
A composite metric is a score built by combining several underlying metrics with weights. It gives an experiment with more than one goal a single number to decide the winner. Use one when two or three goals might disagree, like more revenue but shorter sessions. Then you don't have to guess by eye which result matters most.
What is it
You build a composite inside the experiment builder's Metrics step, not when you first create a metric. Pick the Composite (weighted) measure on a binding. A binding is a saved way of measuring something, such as conversion rate or revenue per visitor. Then add other metric bindings as its components: any binding already attached to this experiment, or one saved to your project's library. A composite needs at least 2 components, and allows at most 10.
To reuse an attached binding in other experiments' composites, click Save to library under its name in the Metrics step.
Each component gets a weight. Type any non-negative numbers, like 60 and 40, or 0.6 and 0.4. A vs B rescales them to sum to exactly 1.0 when you save.
- Pick Composite (weighted) as the measure to start building one.
- Add each component from this experiment's metrics or your project's saved metric library, then give it a weight.
- A vs B rescales your weights to sum to 1.0 automatically when you save.
- This warning appears only when one component would carry more than 85% of the total weight.
The composite's score for each variation is the weighted sum of its components' averages:
composite = Σᵢ wᵢ · μᵢWhen to use
Use a composite when an experiment must balance competing goals, and deciding by eye would be a guess:
- Engagement vs revenue: a change that raises session length but lowers revenue per visitor. Composite = 0.5 × session-length + 0.5 × revenue-per-visitor lets the team pick that trade-off up front.
- Conversion vs retention: an onboarding flow that drives more signups but also more 7-day churn. Composite = 0.4 × signup-rate + 0.6 × 30-day-retention favours stickier users, matching what the team agreed.
- Satisfaction vs activation: a tutorial step that improves NPS but slows activation. Composite = 0.7 × activation-rate + 0.3 × NPS keeps activation the main driver while still rewarding satisfaction wins.
If one component carries more than 85% of the total weight, the composite builder shows a warning. It says the composite is effectively a single metric, with only a token contribution from the rest. A 0.95 × revenue + 0.05 × NPS composite isn't multi-objective: it's revenue with cosmetic NPS. Saving isn't blocked, since sometimes that's what you want, but the warning makes the trade-off deliberate.
Example
You're testing a new onboarding flow. The product team wants both higher activation and higher 30-day retention, and has agreed in advance that activation matters twice as much as retention.
- Create an Activation Custom Metric (fires on first meaningful action) and save a binding for it with the Unique conversions per visitor measure.
- Create a 30-day retention Custom Metric (fires when a 30-day-active session is recorded) and save a binding for it with the same measure.
- In the experiment builder's Metrics step, add a third binding and pick the Composite (weighted) measure. Add the two bindings above as components with weights 0.66 / 0.33 (auto-normalised to 0.667 / 0.333).
- Mark the composite binding as the primary metric (the one metric the experiment is actually judged on).
After the experiment has data:
Composite (Onboarding success): Control: 0.241 (weighted) Variant: 0.268 (weighted) Lift: +11.2% (95% CI: +3.4% to +19.5%) Badge: Composite (weighted)Show breakdown: Component Mean (control → variant) Weight Contribution Activation 0.250 → 0.287 67% +0.025 30-day retention 0.220 → 0.229 33% +0.003 Total lift +0.027The breakdown shows that activation drove most of the win. Retention nudged the same way. If retention had gone flat or negative instead, this view would show that right away, instead of hiding it inside the composite.
- The composite's own score and lift appear in the results table like any other primary metric.
- Click Show breakdown to see which components drove the result.
- Contribution is each component's weighted share of the difference from control, in the composite's own units, not a percentage.
How A vs B computes it
Every time results load, A vs B runs one query that builds a single weighted number for each visitor:
y = w₁ · x₁ + w₂ · x₂ + ...using that visitor's own value for each component. A vs B averages y across every visitor exposed (they saw the variation and were counted in the experiment). It also measures how much y varies across those visitors.
Every visitor's components are combined before anything gets averaged. So components that move together, like revenue and order count, are already accounted for. No separate step has to estimate how they correlate.
The result is mathematically the same as this two-step version:
composite = Σᵢ wᵢ · μᵢVar(composite) = Σᵢ wᵢ² · Var(Xᵢ) + 2 · ΣᵢΣⱼ>ᵢ wᵢ · wⱼ · Cov(Xᵢ, Xⱼ)A vs B never actually estimates that covariance term on its own. Building y per visitor first bakes the correlation into the one number being measured.
Three stats engines read the combined average and variance:
- Frequentist: a z-test against control, using the combined variance.
- Bayesian: a Normal posterior (your updated best guess for the true value). It combines a starting assumption with the data collected, using the combined average and variance.
- Sequential: an always-valid bound, using the same combined variance. It's wider than the interval above on purpose, so you can check results early without inflating your false-positive rate.
CUPED (less noise from history) and winsorization (capping outliers)
Winsorization (capping extreme values, like one giant order, before averaging) applies per component, before the weighted score. Each component is capped at its own configured percentile first. Then A vs B combines the capped values.
CUPED (a technique that uses a visitor's pre-experiment behaviour to shrink noise in the result) applies to composite metrics the same way it applies to standard and money metrics. When your project's variance-reduction setting is Auto, A vs B builds each visitor's combined score for the 30 days before the experiment, using the same components and the same weights, and uses it as that visitor's baseline. If enough visitors have a pre-experiment score and it genuinely predicts their in-experiment score, the composite's result is adjusted and its interval tightens; the metric's row then shows the usual CUPED tooltip. If the pre-experiment data isn't there, the composite simply reports its raw numbers. The decomposition panel always explains the raw movement per component, since the adjustment applies to the combined score, not to each component separately. See Variance Reduction (CUPED) for the eligibility checks.
FAQ
What if I don't know what weights to use?
Start with equal weights, like 0.5 / 0.5 for two components, and iterate from there. Many teams use composite metrics to force the up-front decision about what matters more. That beats arguing afterwards about whether a secondary metric (tracked for context, not for deciding the winner) should disqualify a winning primary metric.
Can a composite reference another composite?
No. A vs B blocks this when you save, using a cycle check across every composite binding in your organization. The math would still work, but a composite of composites gets confusing fast, and no comparable analytics tool supports it either.
What happens if a component metric is paused mid-experiment?
Archiving (pausing) a metric only stops it from tracking new events going forward. The composite keeps computing with whatever data already exists for that component, and there's no separate warning shown for a paused component today.
Does changing a weight require an amendment?
An analysis plan (the metrics, engine, and stopping rule you commit to before an experiment starts) locks in your decisions. They can't be second-guessed by hindsight. If the composite is referenced by a sealed plan, A vs B refuses a weight change outright: the save is blocked until the plan is amended first through the experiment's plan card. The same block applies to a change in measure, direction, or winsorization. Name-only edits (renaming, updating the description) are always allowed.