Uploading & Versioning
Every refresh of a dataset's data is a new version. This page covers the NDJSON row format, how to upload through the dashboard, what commit validation checks, and how activation and rollback work.
Each dataset numbers its own versions starting at v1: your first upload is always v1, the next v2, and so on. The v3 you see is that dataset's own count, not a shared global number, so two different datasets both have a v1.
The NDJSON format
Uploads are NDJSON: newline-delimited JSON, one complete JSON object per line. It streams, so a 500,000-row file never has to sit in memory at once.
Rules for every row:
- Each line is one valid JSON object.
- The key field (named
keyby default, configurable per dataset at creation) is required on every row, must be a string, and can be at most 256 characters. - The rest of the row is free-form JSON, up to 64 KB per line.
- At most 500,000 rows per version.
What a row looks like depends on the dataset type.
LIST: membership only. Rows need nothing beyond the key field:
{"key":"CUST-10381"}{"key":"CUST-20144"}{"key":"CUST-31907"}TABLE: one object of data per key:
{"key":"SHIRT123","headline":"Best seller","badge":"Low stock","priceCopy":"Now $29.99"}{"key":"SHOE456","headline":"New arrival","badge":null,"priceCopy":"From $59.99"}The recommendation engine does not read from a dataset. It reads title, description, category, stock and createdAt from your project's Live Catalog instead. Set that up once, and every recipe uses it: there is no dataset to pick. See Set up your product catalog and Recommendations setup.
FEED: by convention, an ordered items array per key. Store each item denormalised (everything the UI needs to render it), so displaying a feed is one read with no follow-up fetches:
{"key":"SHIRT123","items":[{"id":"SHOE456","title":"Runner","image":"https://cdn.example.com/shoe456.jpg","href":"/p/shoe456","price":29.99},{"id":"HAT789","title":"Cap","image":"https://cdn.example.com/hat789.jpg","href":"/p/hat789","price":14.99}]}Order in the items array implies position: the first item is your top recommendation.
If the same key appears on multiple lines in one upload, the last line wins and the key is counted once. The commit succeeds, but the result includes a duplicate-key warning with the number of duplicates found: worth checking, since duplicates usually mean an upstream export bug.
The dashboard's upload modal also accepts .csv files and converts them to NDJSON in your browser before upload: the first row is treated as the header and each subsequent row becomes one JSON object. The API and CLI take NDJSON only.
The modal detects the delimiter (comma, semicolon, tab, or pipe) from the header, shows a preview of the first rows with the resolved key column before you commit, and lets you override the delimiter if detection guessed wrong. Files over 128 MB are refused in the browser: the CLI has no such limit, since avsb datasets push runs outside a browser tab. A CSV over 500,000 rows is rejected client-side with the same limit the server enforces.
Uploading through the dashboard
Datasets live under Projects → Commerce → Datasets in the side navigation.
Create the dataset
Click Create dataset, choose a slug (kebab-case, immutable after creation), a name, the type (LIST, TABLE, or FEED), and optionally a custom key field. The key field defaults to key.
Upload a version
Open the dataset and click Upload version. Pick an .ndjson or .csv file and optionally give the version a label (for example 2026-06-12 nightly). The upload starts a new version: the current live version keeps serving untouched.
Watch the commit
After upload, the version validates and indexes. The version table shows its status moving from committing to committed, along with the row count. If any row is invalid, the version is marked failed and the table shows the offending line number and the reason (for example a missing key field or an oversized row). A failed version never serves; fix the file and upload again.
Check the data before activating
Click View data on the committed version to page through its actual rows and search within them, so you can confirm the file parsed the way you expected. Then use Test a key, pick that version from the version dropdown, and look up a single key against it, you verify the new version is right before switching any traffic to it.
Activate
Click Activate on the committed version. The flip is atomic: readers go from fully-old data to fully-new data, never a mix. Edge caches refresh within about a minute.
Verify the live version
Leave the version dropdown on the live version and run Test a key again to fetch any single key and see exactly what readers will get back.
Rollback
To roll back, activate an older committed version: it is the same atomic flip. The last 5 non-live versions are retained for exactly this purpose; older versions are archived automatically.
Viewing and downloading a version's data
Every committed or live version can be inspected and exported straight from the version table: you never have to guess what a version actually contains.
- View data opens a viewer that pages through the version's rows, with a search box that filters within the version by key or by any field's text. Each dataset shows its own columns, with the key field first. A row that somehow isn't valid JSON is shown verbatim rather than hidden, so nothing is silently dropped. Load more fetches the next page.
- NDJSON downloads the version exactly as stored: a byte-for-byte copy you can re-upload, diff, or archive.
- CSV downloads a spreadsheet-friendly conversion. Because empty fields are dropped on upload and unquoted numbers are stored as numbers, CSV is a best-effort reconstruction: for an exact copy, download NDJSON instead. CSV export is capped at 100,000 rows; for a larger version, download NDJSON.
Archived and failed versions no longer keep a stored copy of their rows, so View data and the download links appear only on committed and live versions. This is also why you can test a committed version's keys before activating it, but not an archived one.
Commit validation reference
A commit checks every row and fails the whole version on the first hard error:
| Check | Outcome when violated |
|---|---|
| Line is valid JSON | Version fails with the line number and reason |
| Key field present and a string | Version fails |
| Key at most 256 characters | Version fails |
| Row at most 64 KB | Version fails |
| At most 500,000 rows | Version fails |
| Duplicate keys in the file | Commit succeeds with a warning; last line wins |
| Rows that look like PII, on a Block dataset (default) | Version fails with pii_blocked and the offending line: remove the personal data, or switch the dataset to Warn |
| Rows that look like PII, on a Warn dataset | Commit succeeds with a PII warning (count and first line); activating requires an explicit acknowledgment |
The PII check flags emails, phone numbers, card-like numbers, and postal codes. Numeric data such as SKUs, prices, and costs never trips it. See What Are Datasets? for how to switch a dataset between Block and Warn.
Only committed versions can be activated. Activating a version that is still uploading, still committing, or failed is rejected. You can, however, activate one version while another version's commit is still in flight: versions are independent.
Discarding and recovering versions
Any version except the live (active) one can be discarded from the version table: an abandoned upload, a superseded committed version, or a failed one. Discarding deletes the version and reclaims its stored rows and uploaded file; the live version and the rest of the history are untouched.
A version that is still Processing (committing) cannot be discarded: the upload is mid-write, and pulling the version out from under it would strand the rows it is still saving. If one sits there for an unusually long time its background job was lost, and a recovery marks such versions failed automatically, after which the normal Discard action applies. Every recovery is recorded, so a stuck version is never silently abandoned.
Deleting a whole dataset reclaims all of its versions' stored rows and uploaded files too, no orphaned storage is left behind.