What Are Datasets?

A dataset is a named, versioned, keyed lookup table served at the edge. You upload rows keyed by a string of your choosing, such as a user ID, a product SKU, or a postcode. A vs B indexes the rows, and every read is one fast lookup answered from the nearest edge location. That's fast enough to call from inside a variation (one specific version being tested, control or a challenger) without slowing the page down.

Datasets exist because experiments often need data that lives in your own systems. Maybe that's which customers are in the beta cohort. Maybe it's what copy to show for each product, or which items your recommendation pipeline picked tonight. Instead of hardcoding that data into a variation, or building your own lookup service, you push it to A vs B once. Then you read it back wherever your experiments run.

Three dataset types

All three types share the same underlying keyed store: the only difference is the shape of the value behind each key.

TypeQuestion it answersValue shapeTypical use
LIST"Is this key in the set?"Membership only: the key itself is the dataTarget lists: beta customers, suppressed accounts, store IDs in a rollout region
TABLE"What's the row for this key?"One JSON object per keyContent by product: per-SKU headlines, badges, pricing copy
FEED"What's the ordered list for this key?"An ordered items array per keyRecommendation feeds: "people also bought" rows from your own pipeline

Not sure which type to pick? A LIST carries no data beyond membership. A TABLE returns exactly one object. A FEED returns an ordered array, where position means rank.

The version lifecycle

A dataset never changes in place. Data flows through an explicit four-step lifecycle:

  1. Create the dataset: pick a slug, a name, a type, and (optionally) which field in your rows is the key. This is the stable container; you create it once.
  2. Create a version and upload: each refresh of the data is a new version. You upload rows as NDJSON (the dashboard also accepts CSV and converts it for you).
  3. Commit: A vs B validates every row, reports any failure with the offending line number and reason, and indexes the version for serving. A version that fails validation never serves: there is no partial state.
  4. Activate: one atomic flip makes the committed version live. Readers move from fully-old to fully-new; they never see a mix of two versions.

Rolling back is the same flip in reverse: re-activate any older committed version, and the previous data is live again within about a minute.

Versioning is the point

Because serving is always pinned to one specific version, a bad upload can never corrupt what visitors see. You activate only when you're ready, and you can roll back instantly if something goes wrong. Versioned serving is also built for experimentation. An experiment will be able to pin a different dataset version to each variation. That's how you'll A/B test two recommendation feeds against each other, in the upcoming recommendation-experiments release.

Retention

The last 5 non-live versions of each dataset are kept and remain available for instant rollback. Older versions are archived automatically and their data is removed from serving infrastructure.

Limits

LimitValue
Datasets per project50
Rows per version500,000
Size per row64 KB
Key length256 characters
Retained rollback versions5

The dataset's slug, type, and key field are immutable after creation. They define the serving namespace and how rows are read, so changing them would orphan existing versions. To change one, create a new dataset.

No PII in datasets: enforced by default

Datasets must not contain personally identifiable information: no names, email addresses, phone numbers, card numbers, or postal codes. Dataset rows are served from edge caches and read by client-side code. Key rows by opaque identifiers (user IDs, SKUs) and keep personal data in your own systems.

This is enforced. Every dataset ships on the Block policy. An upload whose rows contain PII-shaped values, such as emails, phone numbers, card-like numbers, or postal codes, fails instead of committing.

If you're certain a dataset is PII-free, or you're cleared to store the data, you can switch it to Warn in its settings. That switch needs an explicit acknowledgment. Once you've made it, Warn commits the upload and only flags the suspect rows. Numeric data like SKUs and costs never trips the check.

You can also run the upload self-test from the Datasets page. It confirms uploads can reach storage and the ingestion service, before you rely on them.

Reserved dataset slugs

The slug product-costs is reserved for profit metric calculation. You can create a TABLE dataset with this slug (keyed by sku, values shaped {"cost": <decimal>}) and A vs B's nightly profit job will use it automatically. The slug cannot be used for any other dataset type. See Profit Metrics for the full setup guide.

Where you can read a dataset

  • Browser snippet: avsb.dataset(slug) returns an async handle with get(key) and has(key). See Serving & Reading Datasets.
  • Node SDK: client.datasets.get(slug, key) from @avsbhq/node v1.3.0+.
  • REST API: manage datasets and versions programmatically; see the API & CLI reference.
  • CLI: avsb datasets push wraps the whole upload-commit-activate flow in one command, ideal for nightly jobs.
Was this helpful?