Datasets API & CLI Reference
Everything the dashboard does with datasets is available over the REST API, and the CLI wraps the common flows in single commands. Use the API for custom automation; use the CLI when a one-liner in a cron job is enough.
REST API
Dataset endpoints are part of the public REST API. Authenticate with a bearer service token, and use the standard envelope: successful responses return { "data": ... }, errors return { "error": ... }. A scope is a named permission on an API token that controls exactly what it is allowed to read or change. Reads on datasets require the datasets:read scope; writes require datasets:write.
All routes live under:
/api/v1/projects/{projectId}/datasetsEndpoints
| Method | Path | Purpose |
|---|---|---|
GET | /datasets | List datasets in the project |
POST | /datasets | Create a dataset |
GET | /datasets/{datasetId} | Get a dataset plus its version history |
PATCH | /datasets/{datasetId} | Update the name, description, or PII policy |
DELETE | /datasets/{datasetId} | Delete a dataset and all its versions (blocked for AvsB-managed datasets) |
GET | /datasets/{datasetId}/versions | List versions |
POST | /datasets/{datasetId}/versions | Create a version: returns an upload URL |
POST | /datasets/{datasetId}/versions/{versionId}/commit | Validate and index the uploaded file |
POST | /datasets/{datasetId}/versions/{versionId}/activate | Make a committed version live |
GET | /datasets/{datasetId}/lookup?key=... | Test a key against the live version, or an older committed one |
GET /datasets and GET .../versions return a page at a time. Pass limit and cursor the same way every other list endpoint does; see cursor pagination for the full pattern.
"AvsB-managed" datasets are ones A vs B creates for you, such as the product-costs table a Shopify sync keeps up to date. DELETE refuses to remove one, with 409 dataset_state_conflict, because a feature still depends on it.
Create a dataset
Only slug, name, and type are required. This is the smallest request that works:
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets \ -X POST \ -H "Authorization: Bearer $AVSB_TOKEN" \ -H "Content-Type: application/json" \ -d '{"slug": "pdp-related", "name": "PDP related products", "type": "FEED"}'const projectId = 'cm1a2b3c4d5e6f7g8h9i0j1k2'const res = await fetch(`https://app.avsb.cloud/api/v1/projects/${projectId}/datasets`, { method: 'POST', headers: { Authorization: `Bearer ${process.env.AVSB_SERVICE_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ slug: 'pdp-related', name: 'PDP related products', type: 'FEED' }),})const { data } = await res.json()console.log(data.id)import os, requestsproject_id = "cm1a2b3c4d5e6f7g8h9i0j1k2"res = requests.post( f"https://app.avsb.cloud/api/v1/projects/{project_id}/datasets", headers={"Authorization": f"Bearer {os.environ['AVSB_SERVICE_TOKEN']}"}, json={"slug": "pdp-related", "name": "PDP related products", "type": "FEED"},)data = res.json()["data"]slug must be kebab-case (lowercase letters, digits, hyphens; up to 64 characters), and type is LIST, TABLE, or FEED. Add the optional fields when you need them:
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets \ -X POST \ -H "Authorization: Bearer $AVSB_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "slug": "pdp-related", "name": "PDP related products", "description": "Nightly co-purchase feed from the warehouse pipeline", "type": "FEED", "keyField": "key", "piiPolicy": "BLOCK" }'description: free text shown in the dashboard's dataset list. Not read at serving time.keyField: which field in each uploaded row holds that row's key. Defaults tokey.piiPolicy:BLOCK(the default) fails an upload whose rows look like they contain personal data;WARNcommits the upload anyway and only flags the suspect rows.
slug, type, and keyField are immutable after creation: PATCH only accepts name, description, and piiPolicy.
{ "data": { "id": "<datasetId>", "shortId": 7, "projectId": "<projectId>", "slug": "pdp-related", "name": "PDP related products", "description": "Nightly co-purchase feed from the warehouse pipeline", "type": "FEED", "source": "USER", "keyField": "key", "piiPolicy": "BLOCK", "liveVersionId": null, "liveVersionShortId": null, "liveVersionNumber": null, "createdAt": "2026-06-12T09:00:00.000Z", "updatedAt": "2026-06-12T09:00:00.000Z" }}A brand-new dataset has no version yet, so the three liveVersion* fields start out null. Reusing a slug that already exists in the project fails instead of creating a duplicate:
{ "error": { "code": "validation_failed", "message": "A dataset with slug \"pdp-related\" already exists in this project" }}Requires the datasets:write scope, and counts against the write rate limit: 120 requests per minute for a scoped token, or the shared 600 a minute of an admin:* token.
Upload a version
Creating a version returns the version record plus a presigned upload URL:
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions \ -X POST \ -H "Authorization: Bearer $AVSB_TOKEN" \ -H "Content-Type: application/json" \ -d '{"label": "2026-06-12 nightly"}'{ "data": { "version": { "id": "...", "shortId": 12, "status": "UPLOADING", ... }, "uploadUrl": "https://..." } }PUT your NDJSON file to uploadUrl with the application/x-ndjson content type. The URL expires after 1 hour:
curl "$UPLOAD_URL" \ -X PUT \ -H "Content-Type: application/x-ndjson" \ --data-binary @feed.ndjsonCommit, poll, activate
POST .../commit starts validation, then indexing runs in the background. Poll the version (via the dataset's versions list or detail) until its status settles at COMMITTED or FAILED. A FAILED version's error field names the exact problem, for example { "line": 4012, "reason": "row exceeds 65536 bytes" }. Then POST .../activate to flip it live.
Activating a version whose status is not COMMITTED (UPLOADING, COMMITTING, FAILED, or ARCHIVED) is rejected with 409 dataset_state_conflict. Activation itself is atomic: edge readers converge on the new version within about a minute, the length of the read route's own cache.
The nightly-push recipe
The standard automation pattern: create a version, upload, commit, poll, activate:
#!/usr/bin/env bashset -euo pipefail# avsb:conformance fixture=list-dataset-versionsBASE="https://app.avsb.cloud/api/v1/projects/$PROJECT_ID/datasets/$DATASET_ID"AUTH="Authorization: Bearer $AVSB_TOKEN"# 1. Create a version (returns the version id + a presigned upload URL)CREATED=$(curl -sf "$BASE/versions" -X POST -H "$AUTH" \ -H "Content-Type: application/json" \ -d "{\"label\": \"$(date -u +%F) nightly\"}")VERSION_ID=$(echo "$CREATED" | jq -r '.data.version.id')UPLOAD_URL=$(echo "$CREATED" | jq -r '.data.uploadUrl')# 2. Upload the NDJSON filecurl -sf "$UPLOAD_URL" -X PUT \ -H "Content-Type: application/x-ndjson" \ --data-binary @feed.ndjson# 3. Commitcurl -sf "$BASE/versions/$VERSION_ID/commit" -X POST -H "$AUTH"# 4. Poll the versions list until COMMITTED (or fail on FAILED). The list# endpoint returns the frozen envelope { "data": [ ... ], "page": ... },# so the version rows are the top-level "data" array: select yours by id.# Bounded so a version that never resolves times out instead of looping# forever (60 attempts * 5s = 5 minutes).ATTEMPTS=0MAX_ATTEMPTS=60while [ "$ATTEMPTS" -lt "$MAX_ATTEMPTS" ]; do STATUS=$(curl -sf "$BASE/versions" -H "$AUTH" \ | jq -r ".data[] | select(.id == \"$VERSION_ID\") | .status") [ "$STATUS" = "COMMITTED" ] && break if [ "$STATUS" = "FAILED" ]; then echo "Commit failed: check the version error in the dashboard" >&2 exit 1 fi ATTEMPTS=$((ATTEMPTS + 1)) sleep 5doneif [ "$ATTEMPTS" -ge "$MAX_ATTEMPTS" ]; then echo "Timed out waiting for commit to finish after $MAX_ATTEMPTS attempts" >&2 exit 1fi# 5. Activatecurl -sf "$BASE/versions/$VERSION_ID/activate" -X POST -H "$AUTH"If you'd rather not maintain this script, avsb datasets push --activate below is the same flow in one command.
Test a key
The lookup endpoint answers against the live version by default: the same helper the dashboard's Test a key panel uses.
curl "https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/lookup?key=SHIRT123" \ -H "Authorization: Bearer $AVSB_TOKEN"const projectId = 'cm1a2b3c4d5e6f7g8h9i0j1k2'const datasetId = 'cm9z8y7x6w5v4u3t2s1r0q9p8'const res = await fetch( `https://app.avsb.cloud/api/v1/projects/${projectId}/datasets/${datasetId}/lookup?key=SHIRT123`, { headers: { Authorization: `Bearer ${process.env.AVSB_SERVICE_TOKEN}` } },)const { data } = await res.json()console.log(data.found)import os, requestsproject_id = "cm1a2b3c4d5e6f7g8h9i0j1k2"dataset_id = "cm9z8y7x6w5v4u3t2s1r0q9p8"res = requests.get( f"https://app.avsb.cloud/api/v1/projects/{project_id}/datasets/{dataset_id}/lookup", params={"key": "SHIRT123"}, headers={"Authorization": f"Bearer {os.environ['AVSB_SERVICE_TOKEN']}"},)data = res.json()["data"]{ "data": { "found": true, "value": { "sku": "SHIRT123", "headline": "Back in stock" }, "slug": "pdp-related", "version": 12, "versionOrdinal": 4 }}value is shaped by the dataset's type: the row object for TABLE, true for a LIST membership hit, or the row you uploaded for FEED. A miss returns { "found": false, "slug": ..., "version": ..., "versionOrdinal": ... } rather than a 404, since "not a member" is a valid answer for a LIST dataset.
Add &version=<versionOrdinal> to test a specific committed version instead of the live one, using the vN number shown in the dashboard's version history. Only COMMITTED or ACTIVE versions can be tested this way; an archived version's data has already been reclaimed.
CLI
The avsb datasets command group ships in avsb-cli; update to the latest version if avsb datasets is not recognized. All commands accept --project <project> (PRJ-42 or the numeric ID); without it, the project comes from your local manifest.
avsb datasets list
avsb datasets list --project PRJ-42avsb datasets list --jsonLists the project's datasets in a table with each dataset's slug, type, live version, and last-updated time. --json prints the platform's own rows instead, for scripts.
avsb datasets push
avsb datasets push pdp-related feed.ndjson \ --label "2026-06-12 nightly" \ --activate \ --timeout 300One command for the whole write flow: creates a version, uploads the file, commits, and polls until the commit finishes. On success it prints the row count; on a validation failure it prints the offending line, for example line 4012: row exceeds 65536 bytes (the 64 KB per-row cap).
| Flag | Meaning |
|---|---|
--label <text> | Free-form version label (e.g. "2026-06-12 nightly") |
--activate | Activate the version as soon as it commits |
--timeout <seconds> | How long to wait for the commit to finish |
Without --activate, the version is committed but not live: the command prints the exact avsb datasets activate invocation to run when you're ready.
avsb datasets activate
avsb datasets activate pdp-related v12Activates a specific committed version. The version argument accepts either the bare number (12) or the v-prefixed form (v12). Activating a version that is not committed fails with the version's current status (and the commit error, if it failed).
avsb datasets rollback
avsb datasets rollback pdp-relatedRe-activates the previous live version: the most recent committed version older than the current one. A confirmation prompt shows the flip first, for example v12 → v11. Pass -f to skip the prompt in scripts.
The dashboard, API, and CLI all enforce the same limits. A project holds up to 50 datasets, a version up to 500,000 rows, and a row up to 64 KB. Keys run up to 256 characters, and 5 rolled-back versions stay available. See What Are Datasets? for the full table.