Bandits

A bandit rule sends more traffic to your best-performing variation (one specific version being tested, control or a challenger) as data comes in. It does this on its own, instead of waiting for a fixed experiment to finish. You pick the algorithm and attach a reward metric (the event you want more of, like a click or a purchase). The SDK then serves each visitor the variation most likely to earn it.

What it is

A multi-armed bandit balances two competing goals. Exploration means trying variations you don't have enough data on yet. Exploitation means serving the variation that currently looks best. A standard A/B test splits traffic equally for its whole run. A bandit instead updates its model as data arrives and shifts allocations toward the winner.

A vs B supports three algorithms, all reading the same BanditConfig shape attached to a flag rule:

  • Epsilon-greedy: serves the current best variation with probability 1 - explorationRate, and a random variation with probability explorationRate (0.1 by default, if you don't set one). Simple, and easy to predict.
  • Thompson sampling: models each variation's reward as a Normal distribution built from its mean and variance. It draws one random sample from that distribution per decision (a seeded, repeatable draw, not true randomness), and whichever variation draws highest wins. It explores more while a variation's estimate is still uncertain, and settles down once a clear winner emerges.
  • UCB1: tries every variation once. After that, it always serves the one with the highest upper confidence bound: mean + sqrt(2 ln N / n_i). N is the total plays across every variation; n_i is plays of this one variation. No randomness after the first round: every choice follows the formula, which suits a setting where you want a repeatable decision.

When to use it

Reach for a bandit rule when:

  • You have several content variations (five hero images, three email subject lines) and just want the best one served fast. You don't need proof that holds up in a formal report.
  • Serving a losing variation costs real money or real users, and you'd rather minimize that cost than wait for a fixed experiment to finish.
  • You're personalizing content, and contextual attributes (a visitor's segment, or the page they're on) should influence which variation wins.

Skip bandits and run a standard, fixed A/B test instead when:

  • You need statistical significance (a result unlikely enough to have happened by chance, checked against an agreed threshold), not just a fast winner.
  • You want to read more than one metric at once, rather than optimizing toward a single reward.
  • The call is about your primary metric (the one metric an experiment is actually judged on), for a business decision that needs rigorous proof.

How it works

The bandit's configuration lives in the flag's rule. Its reward model is stored server-side. The datafile is the small JSON file your SDK downloads, listing every live experiment, flag, and rule for your project. Your SDK gets the latest bandit model each time that file refreshes, and reads it to pick a variation at evaluation time:

TypeScript
// BanditConfig shape (part of the flag rule in the datafile)interface BanditConfig {  algorithm: 'epsilon-greedy' | 'thompson-sampling' | 'ucb1'  explorationRate?: number     // only for epsilon-greedy (0–1), defaults to 0.1  rewardMetric: string         // event key tracked via client.track()  actions: Array<{    id: string    variationId: string    contextAttributes?: Record<string, number | string>  }>}// BanditModel snapshot, produced by offline training, read by SDKinterface BanditModel {  version: string  algorithm: 'epsilon-greedy' | 'thompson-sampling' | 'ucb1'  perAction: Record<string, { mean: number; variance: number; samples: number }>  // Optional: a per-context override, keyed by the sorted contextAttributes  // as "key=value;key=value". Used when actions declare contextAttributes.  contextualModels?: Record<string, Record<string, { mean: number; variance: number; samples: number }>>}
TypeScript21 lines

A worked UCB1 example

Say a flag has two actions: hero-a (mean reward 0.08, tried 120 times) and hero-b (mean reward 0.11, tried only 40 times). Total plays across both: 160.

Plain text
score(hero-a) = 0.08 + sqrt(2 * ln(160) / 120) = 0.08 + 0.29 = 0.37score(hero-b) = 0.11 + sqrt(2 * ln(160) / 40)  = 0.11 + 0.50 = 0.61
Plain text2 lines

UCB1 serves hero-b. It has both the better mean and fewer samples, so its "uncertainty bonus" is larger. Next round it might lose to hero-a again, if hero-a's own luck turns. Note also: any action tried zero times is served first, before this formula ever runs.

The decision is logged as a BanditDecisionLogEntry. This extends the standard DecisionLogEntry with five extra fields:

  • the bandit key.
  • the action ID.
  • the model version.
  • the probability assigned to that action.
  • the optimality gap (epsilon at decision time for epsilon-greedy; null for UCB1 and Thompson sampling).
Info

A bandit's action for one visitor comes from a deterministic hash of the bandit key and the visitor's id. This is a form of bucketing (sorting a visitor into a group by a consistent hash of their id). The same visitor always lands in the same group. So re-evaluating that visitor against the same model always serves the same action. It works differently from Sticky Bucketing's storage: only ab_test and targeted_delivery rules write an assignment to that store. A bandit's action can still change for the same visitor once the model retrains and a different action wins their hash bucket.

Per-SDK usage

Client-side, using @avsbhq/browser:

TypeScript
import { AvsbClient } from '@avsbhq/browser'const client = new AvsbClient({ sdkKey: 'sdk_production_xxxxxxxxxxxxxxxx' })await client.onReady()// Evaluation: same API as any flagconst flag = client.getFlag('hero_image_bandit', 'control')// flag.source === 'bandit', flag.variationKey === 'variant-b' (example)// Track the reward metric the bandit is optimizing ondocument.querySelector('.cta')?.addEventListener('click', () => {  client.track('hero_cta_click')})
TypeScript13 lines

Server-side, the same operation (create the client, evaluate the bandit flag with a visitor context, track the reward) in three languages:

import { AvsbServer } from '@avsbhq/node'import type { EvalContext } from '@avsbhq/node'const server = new AvsbServer({ sdkKey: process.env.AVSB_SDK_KEY! })await server.onReady()const userContext: EvalContext = { kind: 'user', key: 'u_123' }// Server-side evaluation with contextconst flag = server.getFlag('pricing_plan_bandit', 'starter', userContext)// Track reward event (e.g., after a plan upgrade)server.track('plan_upgrade', { value: 49, context: userContext })
TypeScript13 lines
Was this helpful?