Holdouts

Holdouts are not yet configurable in the dashboard.

The SDK can read a holdout assignment. But there is currently no way to create one. There is no holdout authoring UI and no API. The flag datafile (the small file your SDK downloads that lists every live flag and rule) does not carry holdout configuration yet either. In practice, a holdout can never actually fire today, and the code samples below have nothing to read.

This page documents the concept and the SDK contract so the shape is clear. Treat it as a design reference, not a feature you can turn on today.

A holdout is a fixed slice of your users that you deliberately keep out of every experiment. Comparing that group against everyone else shows you the combined effect of everything your team shipped, not just one experiment at a time.

What it is

A holdout is a named object that lists one or more flags. It uses the same hashing method as a normal rule: MurmurHash3, a deterministic algorithm, meaning the same user always produces the same result. When a user hashes into the holdout, every flag on its list returns that flag's default variation (usually control) for that user. This happens no matter what a normal rule would otherwise have served. The SDK marks the decision source: 'holdout', so you can count exactly how many reads were held out.

A held-out read fires its own event: flag_holdout_exposure. This is a different event name from flag_exposure, the event a normal rule fires. It carries the holdout's own id and key. Point your analytics pipeline at the event name. That keeps held-out users out of a per-experiment result. It still counts them in a holdout-level report.

A worked example

Say you hold out 5% of your users: 5,000 people out of 100,000. Over one quarter, your team ships 12 experiments to the other 95%. At quarter end, the held-out group's average revenue per user is $42. Everyone else's is $45.

That $3 gap is the combined effect of all 12 experiments together. No single experiment's own result can show you this. Each one only measures its own variations against each other, on the users it targeted. The holdout is the only place that adds everything up.

When to use it

Holdouts are most valuable when your team runs many experiments at once. Typical cases:

  • Measuring programme ROI. Compare the held-out group (no experiments) to everyone else. That shows you the net effect of your whole experimentation pipeline on a metric like retention or revenue.
  • A long-running baseline. Keep a 2 to 5% holdout for months. Then you always have a clean comparison point when leadership asks what your metrics would look like with no product changes at all.
  • Interaction detection. Two experiments might ship at the same time and affect each other. The holdout group, in neither one, acts as a neutral baseline.
Warning

A holdout reduces the traffic available for experiments. A 5% holdout means every experiment runs on only 95% of your users. Plan holdout size with that in mind. This matters most for an experiment that needs strong statistical power: the chance an experiment detects a real difference, if one truly exists, rather than missing it in the noise.

When NOT to use it

  • You run one or two experiments a quarter. A holdout costs you real traffic for a signal that stays close to zero at that volume. Save it for a team shipping experiments continuously.
  • You need an answer fast. A holdout is a slow instrument. Separating a real revenue difference from ordinary noise takes months, unlike a single experiment's own result.
  • Nobody is asking the programme-level question. If a single big win already justified the quarter, a holdout answers a question nobody needs proved.

How it works

The datafile carries a top-level holdouts array. Each entry lists the flags it participates in. It also carries a traffic allocation (the share of visitors held out, from 0 to 1), a hash attribute ('user.key' by default), and a map of the default variation to serve on each held-out flag:

TypeScript
// Excerpt from a v2 datafile, illustrativeconst datafile = {  holdouts: [    {      id: "hld_abc",      key: "global-holdout-2024",      trafficAllocation: 0.05,       // 5% of users      hashAttribute: "user.key",      flagIds: ["checkout_v2", "pricing_experiment", "nav_redesign"],      defaultVariationByFlag: {        "checkout_v2": "control",        "pricing_experiment": "control",        "nav_redesign": "control",      },    },  ],}
TypeScript17 lines

Before evaluating any flag's rules, the SDK checks every holdout that flag participates in. If the context hashes into the holdout bucket, the flag returns the holdout's default variation. Evaluation stops there. A user's hash never changes, so they keep the same holdout membership unless trafficAllocation shrinks and pushes them back out.

Every flag in a holdout shares the same hashAttribute and allocation. So the same 5% of users are held out across all of that holdout's flags at once. This is the mutual exclusion guarantee (no held-out user ever enters a conflicting experiment): a held-out user never enters any experiment in the set.

Per-SDK usage

Holdouts are evaluated transparently: no SDK call changes to support them. The SDK reads the holdout configuration from the datafile and applies it during every flag read. You can tell a decision was held out by checking flag.source:

For the @avsbhq/browser SDK:

TypeScript
import { AvsbClient } from '@avsbhq/browser'const client = new AvsbClient({ sdkKey: 'sdk_production_xxxxxxxxxxxxxxxx' })await client.onReady()// The browser SDK always collects reasons, so there is no option to pass.const flag = client.getFlag('checkout_v2', false)if (flag.source === 'holdout') {  // This user is in the holdout: they see the control variation  analytics.track('holdout_exposure', { flagKey: 'checkout_v2' })}
TypeScript12 lines

Server-side, the same operation (create the client, evaluate with reasons on, check the holdout source) in three languages:

import { AvsbServer } from '@avsbhq/node'import type { EvalContext } from '@avsbhq/node'import { DecideOption } from '@avsbhq/core'const server = new AvsbServer({ sdkKey: process.env.AVSB_SDK_KEY! })await server.onReady()const context: EvalContext = { kind: 'user', key: 'u_123' }const flag = server.getFlag('checkout_v2', false, context, {  decideOptions: [DecideOption.INCLUDE_REASONS],})// flag.source === 'holdout' | 'rule' | 'default' | ...// flag.reasons contains the holdout key when source is 'holdout'
TypeScript15 lines
Was this helpful?