Composite metrics

A composite metric is a score built by combining several underlying metrics with weights. It gives an experiment with more than one goal a single number to decide the winner. Use one when two or three goals might disagree, like more revenue but shorter sessions. Then you don't have to guess by eye which result matters most.

What is it

You build a composite inside the experiment builder's Metrics step, not when you first create a metric. Pick the Composite (weighted) measure on a binding. A binding is a saved way of measuring something, such as conversion rate or revenue per visitor. Then add other metric bindings as its components: any binding already attached to this experiment, or one saved to your project's library. A composite needs at least 2 components, and allows at most 10.

To reuse an attached binding in other experiments' composites, click Save to library under its name in the Metrics step.

Each component gets a weight. Type any non-negative numbers, like 60 and 40, or 0.6 and 0.4. A vs B rescales them to sum to exactly 1.0 when you save.

A composite is built in the experiment builder's Metrics step, not when the underlying metrics are first created.
  1. Pick Composite (weighted) as the measure to start building one.
  2. Add each component from this experiment's metrics or your project's saved metric library, then give it a weight.
  3. A vs B rescales your weights to sum to 1.0 automatically when you save.
  4. This warning appears only when one component would carry more than 85% of the total weight.

The composite's score for each variation is the weighted sum of its components' averages:

Plain text
composite = Σᵢ wᵢ · μᵢ
Plain text1 line

When to use

Use a composite when an experiment must balance competing goals, and deciding by eye would be a guess:

  • Engagement vs revenue: a change that raises session length but lowers revenue per visitor. Composite = 0.5 × session-length + 0.5 × revenue-per-visitor lets the team pick that trade-off up front.
  • Conversion vs retention: an onboarding flow that drives more signups but also more 7-day churn. Composite = 0.4 × signup-rate + 0.6 × 30-day-retention favours stickier users, matching what the team agreed.
  • Satisfaction vs activation: a tutorial step that improves NPS but slows activation. Composite = 0.7 × activation-rate + 0.3 × NPS keeps activation the main driver while still rewarding satisfaction wins.
Effective-primary warning

If one component carries more than 85% of the total weight, the composite builder shows a warning. It says the composite is effectively a single metric, with only a token contribution from the rest. A 0.95 × revenue + 0.05 × NPS composite isn't multi-objective: it's revenue with cosmetic NPS. Saving isn't blocked, since sometimes that's what you want, but the warning makes the trade-off deliberate.

Example

You're testing a new onboarding flow. The product team wants both higher activation and higher 30-day retention, and has agreed in advance that activation matters twice as much as retention.

  1. Create an Activation Custom Metric (fires on first meaningful action) and save a binding for it with the Unique conversions per visitor measure.
  2. Create a 30-day retention Custom Metric (fires when a 30-day-active session is recorded) and save a binding for it with the same measure.
  3. In the experiment builder's Metrics step, add a third binding and pick the Composite (weighted) measure. Add the two bindings above as components with weights 0.66 / 0.33 (auto-normalised to 0.667 / 0.333).
  4. Mark the composite binding as the primary metric (the one metric the experiment is actually judged on).

After the experiment has data:

Plain text
Composite (Onboarding success):  Control:  0.241  (weighted)  Variant:  0.268  (weighted)  Lift:     +11.2%  (95% CI: +3.4% to +19.5%)  Badge:    Composite (weighted)Show breakdown:  Component          Mean (control → variant)   Weight   Contribution  Activation          0.250 → 0.287              67%      +0.025  30-day retention     0.220 → 0.229              33%      +0.003  Total lift                                              +0.027
Plain text11 lines

The breakdown shows that activation drove most of the win. Retention nudged the same way. If retention had gone flat or negative instead, this view would show that right away, instead of hiding it inside the composite.

The Decomposition panel only appears once the experiment has data: there's no preview of it before the experiment runs.
  1. The composite's own score and lift appear in the results table like any other primary metric.
  2. Click Show breakdown to see which components drove the result.
  3. Contribution is each component's weighted share of the difference from control, in the composite's own units, not a percentage.

How A vs B computes it

Every time results load, A vs B runs one query that builds a single weighted number for each visitor:

Plain text
y = w₁ · x₁ + w₂ · x₂ + ...
Plain text1 line

using that visitor's own value for each component. A vs B averages y across every visitor exposed (they saw the variation and were counted in the experiment). It also measures how much y varies across those visitors.

Every visitor's components are combined before anything gets averaged. So components that move together, like revenue and order count, are already accounted for. No separate step has to estimate how they correlate.

The result is mathematically the same as this two-step version:

Plain text
composite = Σᵢ wᵢ · μᵢVar(composite) = Σᵢ wᵢ² · Var(Xᵢ) + 2 · ΣᵢΣⱼ>ᵢ wᵢ · wⱼ · Cov(Xᵢ, Xⱼ)
Plain text2 lines

A vs B never actually estimates that covariance term on its own. Building y per visitor first bakes the correlation into the one number being measured.

Three stats engines read the combined average and variance:

  • Frequentist: a z-test against control, using the combined variance.
  • Bayesian: a Normal posterior (your updated best guess for the true value). It combines a starting assumption with the data collected, using the combined average and variance.
  • Sequential: an always-valid bound, using the same combined variance. It's wider than the interval above on purpose, so you can check results early without inflating your false-positive rate.

CUPED (less noise from history) and winsorization (capping outliers)

Winsorization (capping extreme values, like one giant order, before averaging) applies per component, before the weighted score. Each component is capped at its own configured percentile first. Then A vs B combines the capped values.

CUPED (a technique that uses a visitor's pre-experiment behaviour to shrink noise in the result) applies to composite metrics the same way it applies to standard and money metrics. When your project's variance-reduction setting is Auto, A vs B builds each visitor's combined score for the 30 days before the experiment, using the same components and the same weights, and uses it as that visitor's baseline. If enough visitors have a pre-experiment score and it genuinely predicts their in-experiment score, the composite's result is adjusted and its interval tightens; the metric's row then shows the usual CUPED tooltip. If the pre-experiment data isn't there, the composite simply reports its raw numbers. The decomposition panel always explains the raw movement per component, since the adjustment applies to the combined score, not to each component separately. See Variance Reduction (CUPED) for the eligibility checks.

FAQ

What if I don't know what weights to use?

Start with equal weights, like 0.5 / 0.5 for two components, and iterate from there. Many teams use composite metrics to force the up-front decision about what matters more. That beats arguing afterwards about whether a secondary metric (tracked for context, not for deciding the winner) should disqualify a winning primary metric.

Can a composite reference another composite?

No. A vs B blocks this when you save, using a cycle check across every composite binding in your organization. The math would still work, but a composite of composites gets confusing fast, and no comparable analytics tool supports it either.

What happens if a component metric is paused mid-experiment?

Archiving (pausing) a metric only stops it from tracking new events going forward. The composite keeps computing with whatever data already exists for that component, and there's no separate warning shown for a paused component today.

Does changing a weight require an amendment?

An analysis plan (the metrics, engine, and stopping rule you commit to before an experiment starts) locks in your decisions. They can't be second-guessed by hindsight. If the composite is referenced by a sealed plan, A vs B refuses a weight change outright: the save is blocked until the plan is amended first through the experiment's plan card. The same block applies to a change in measure, direction, or winsorization. Name-only edits (renaming, updating the description) are always allowed.

Was this helpful?