Bandits
A bandit rule sends more traffic to your best-performing variation (one specific version being tested, control or a challenger) as data comes in. It does this on its own, instead of waiting for a fixed experiment to finish. You pick the algorithm and attach a reward metric (the event you want more of, like a click or a purchase). The SDK then serves each visitor the variation most likely to earn it.
What it is
A multi-armed bandit balances two competing goals. Exploration means trying variations you don't have enough data on yet. Exploitation means serving the variation that currently looks best. A standard A/B test splits traffic equally for its whole run. A bandit instead updates its model as data arrives and shifts allocations toward the winner.
A vs B supports three algorithms, all reading the same BanditConfig shape attached to a flag rule:
- Epsilon-greedy: serves the current best variation with probability
1 - explorationRate, and a random variation with probabilityexplorationRate(0.1 by default, if you don't set one). Simple, and easy to predict. - Thompson sampling: models each variation's reward as a Normal distribution built from its mean and variance. It draws one random sample from that distribution per decision (a seeded, repeatable draw, not true randomness), and whichever variation draws highest wins. It explores more while a variation's estimate is still uncertain, and settles down once a clear winner emerges.
- UCB1: tries every variation once. After that, it always serves the one with the highest upper confidence bound:
mean + sqrt(2 ln N / n_i).Nis the total plays across every variation;n_iis plays of this one variation. No randomness after the first round: every choice follows the formula, which suits a setting where you want a repeatable decision.
When to use it
Reach for a bandit rule when:
- You have several content variations (five hero images, three email subject lines) and just want the best one served fast. You don't need proof that holds up in a formal report.
- Serving a losing variation costs real money or real users, and you'd rather minimize that cost than wait for a fixed experiment to finish.
- You're personalizing content, and contextual attributes (a visitor's segment, or the page they're on) should influence which variation wins.
Skip bandits and run a standard, fixed A/B test instead when:
- You need statistical significance (a result unlikely enough to have happened by chance, checked against an agreed threshold), not just a fast winner.
- You want to read more than one metric at once, rather than optimizing toward a single reward.
- The call is about your primary metric (the one metric an experiment is actually judged on), for a business decision that needs rigorous proof.
How it works
The bandit's configuration lives in the flag's rule. Its reward model is stored server-side. The datafile is the small JSON file your SDK downloads, listing every live experiment, flag, and rule for your project. Your SDK gets the latest bandit model each time that file refreshes, and reads it to pick a variation at evaluation time:
// BanditConfig shape (part of the flag rule in the datafile)interface BanditConfig { algorithm: 'epsilon-greedy' | 'thompson-sampling' | 'ucb1' explorationRate?: number // only for epsilon-greedy (0–1), defaults to 0.1 rewardMetric: string // event key tracked via client.track() actions: Array<{ id: string variationId: string contextAttributes?: Record<string, number | string> }>}// BanditModel snapshot, produced by offline training, read by SDKinterface BanditModel { version: string algorithm: 'epsilon-greedy' | 'thompson-sampling' | 'ucb1' perAction: Record<string, { mean: number; variance: number; samples: number }> // Optional: a per-context override, keyed by the sorted contextAttributes // as "key=value;key=value". Used when actions declare contextAttributes. contextualModels?: Record<string, Record<string, { mean: number; variance: number; samples: number }>>}A worked UCB1 example
Say a flag has two actions: hero-a (mean reward 0.08, tried 120 times) and hero-b (mean reward 0.11, tried only 40 times). Total plays across both: 160.
score(hero-a) = 0.08 + sqrt(2 * ln(160) / 120) = 0.08 + 0.29 = 0.37score(hero-b) = 0.11 + sqrt(2 * ln(160) / 40) = 0.11 + 0.50 = 0.61UCB1 serves hero-b. It has both the better mean and fewer samples, so its "uncertainty bonus" is larger. Next round it might lose to hero-a again, if hero-a's own luck turns. Note also: any action tried zero times is served first, before this formula ever runs.
The decision is logged as a BanditDecisionLogEntry. This extends the standard DecisionLogEntry with five extra fields:
- the bandit key.
- the action ID.
- the model version.
- the probability assigned to that action.
- the optimality gap (epsilon at decision time for epsilon-greedy; null for UCB1 and Thompson sampling).
A bandit's action for one visitor comes from a deterministic hash of the bandit key and the visitor's id. This is a form of bucketing (sorting a visitor into a group by a consistent hash of their id). The same visitor always lands in the same group. So re-evaluating that visitor against the same model always serves the same action. It works differently from Sticky Bucketing's storage: only ab_test and targeted_delivery rules write an assignment to that store. A bandit's action can still change for the same visitor once the model retrains and a different action wins their hash bucket.
Per-SDK usage
Client-side, using @avsbhq/browser:
import { AvsbClient } from '@avsbhq/browser'const client = new AvsbClient({ sdkKey: 'sdk_production_xxxxxxxxxxxxxxxx' })await client.onReady()// Evaluation: same API as any flagconst flag = client.getFlag('hero_image_bandit', 'control')// flag.source === 'bandit', flag.variationKey === 'variant-b' (example)// Track the reward metric the bandit is optimizing ondocument.querySelector('.cta')?.addEventListener('click', () => { client.track('hero_cta_click')})Server-side, the same operation (create the client, evaluate the bandit flag with a visitor context, track the reward) in three languages:
import { AvsbServer } from '@avsbhq/node'import type { EvalContext } from '@avsbhq/node'const server = new AvsbServer({ sdkKey: process.env.AVSB_SDK_KEY! })await server.onReady()const userContext: EvalContext = { kind: 'user', key: 'u_123' }// Server-side evaluation with contextconst flag = server.getFlag('pricing_plan_bandit', 'starter', userContext)// Track reward event (e.g., after a plan upgrade)server.track('plan_upgrade', { value: 49, context: userContext })import osfrom avsb import AvsbServer, SingleContextserver = AvsbServer(sdk_key=os.environ["AVSB_SDK_KEY"])server.on_ready()context = SingleContext.user(key="u_123")# Evaluate: identical API to non-bandit flagsflag = server.get_string_flag("email_subject_bandit", "control", context)# Track rewardserver.track("email_open", context)import com.avsbhq.avsb.AvsbServer;import com.avsbhq.avsb.EvalContext;AvsbServer server = AvsbServer.builder(System.getenv("AVSB_SDK_KEY")).build();server.init();EvalContext ctx = EvalContext.user("u_123");var flag = server.getStringFlag("email_subject_bandit", "control", ctx);// Track reward metricserver.track("email_open", ctx);Related concepts
- A bandit and a holdout (a slice of visitors kept out of every experiment) work well together. The holdout still lets you measure the combined effect of everything you shipped, bandits included.
- Decision Logging: capturing bandit action probabilities
- Sticky Bucketing: how ab_test assignments persist, a different mechanism from a bandit's own hash-based consistency