Public API: Datasets

A dataset is a named, versioned keyed lookup table served from the edge. Each dataset has a fixed shape (a type: LIST for membership, TABLE for single-row lookups, FEED for ordered items, and a keyField) and a stream of immutable versions. You upload rows as NDJSON into a new version, commit it for ingestion, then activate it to make it the live version. Activating an older version is how you roll back.

All datasets endpoints live under a project. Every one authenticates with a Bearer token, either a service token or a personal access token, and the org is taken from the token, so the path carries only {projectId}, not {orgId}.

Plain text
https://app.avsb.cloud/api/v1/projects/{projectId}/datasets
Plain text1 line

Scopes

A scope is a named permission on your token. It controls exactly what the token can read or change.

OperationScope
List, get, list versions, key lookupdatasets:read
Create, update, delete, create version, discard version, commit, activatedatasets:write

Every /api/v1 token is rate-limited: a scoped token gets 600 reads and 120 writes per minute, and an admin:* token gets 600 requests per minute, reads and writes together. Every write call below also accepts an optional Idempotency-Key header for safe retries. See Conventions for both.

The upload → commit → activate flow

A new dataset version moves through a small state machine. The write calls below are the whole lifecycle:

1

Create a version

POST .../versions mints a version in UPLOADING and returns a short-lived presigned R2 uploadUrl.

2

Upload the NDJSON

PUT your newline-delimited JSON rows to the uploadUrl with Content-Type: application/x-ndjson. This is a direct upload to storage; it does not go through the AvsB API.

3

Commit the version

POST .../versions/{versionId}/commit flips the version to COMMITTING and hands it to the ingestion pipeline, which streams the rows into the edge store and marks the version COMMITTED (or FAILED).

4

Activate the version

POST .../versions/{versionId}/activate promotes a COMMITTED version to ACTIVE, repoints the dataset's live pointer, and republishes the datafile so the new data ships. Activating an older COMMITTED version rolls back.

Changed your mind before a version goes live? DELETE .../versions/{versionId} discards any non-ACTIVE version; see Discard a version.

Info

slug, type, and keyField are fixed when the dataset is created; only name, description, and piiPolicy can be changed later. The key namespace and row validation depend on slug, type, and keyField.

List datasets

GET /api/v1/projects/{projectId}/datasets: cursor-paginated, newest first, up to 100 per page (default 20). Requires datasets:read.

curl 'https://app.avsb.cloud/api/v1/projects/<projectId>/datasets?limit=20' \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{  "data": [    { "id": "cmd...", "shortId": 12, "projectId": "cmp...", "slug": "vip-customers", "name": "VIP customers", "description": null,      "type": "LIST", "source": "USER", "keyField": "key", "piiPolicy": "BLOCK", "liveVersionId": "cmv...",      "liveVersionShortId": 23, "liveVersionNumber": 4, "createdAt": "2026-06-18T00:00:00.000Z", "updatedAt": "2026-06-18T00:00:00.000Z" }  ],  "page": { "nextCursor": "MjAyNi0wNi0xOFQwMDowMDowMC4wMDBafGNtZC4uLg", "hasMore": true }}
JSON8 lines

liveVersionShortId is that version's global serving address; liveVersionNumber is its per-dataset "v4" label, the one the dashboard shows. Pass page.nextCursor back as ?cursor= for the next page; treat it as an opaque string.

Create a dataset

POST /api/v1/projects/{projectId}/datasets: slug, name, and type are required. keyField defaults to "key" and piiPolicy defaults to "BLOCK" (PII-shaped rows fail ingestion) when left out. Returns the new dataset, 201 Created. Requires datasets:write.

The smallest request that works:

curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets \  -X POST -H "Authorization: Bearer avsb_svc_..." -H "Content-Type: application/json" \  -d '{ "slug": "vip-customers", "name": "VIP customers", "type": "LIST" }'
Shell3 lines

Adding the optional fields sends description (free text for your team), a non-default keyField (the NDJSON row field to look rows up by, here an email address instead of "key"), and piiPolicy: "WARN" (let a commit succeed with a warning instead of failing on PII-shaped rows). Send Idempotency-Key on this call too, to retry safely without creating a duplicate:

Fuller request
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets \  -X POST -H "Authorization: Bearer avsb_svc_..." -H "Content-Type: application/json" -H "Idempotency-Key: $(uuidgen)" \  -d '{ "slug": "vip-customers", "name": "VIP customers", "type": "LIST", "description": "Shopify VIP segment, refreshed nightly", "keyField": "email", "piiPolicy": "WARN" }'
Shell3 lines
Response
{  "data": {    "id": "cmd...", "shortId": 13, "projectId": "cmp...", "slug": "vip-customers", "name": "VIP customers",    "description": "Shopify VIP segment, refreshed nightly", "type": "LIST", "source": "USER", "keyField": "email",    "piiPolicy": "WARN", "liveVersionId": null, "liveVersionShortId": null, "liveVersionNumber": null,    "createdAt": "2026-06-18T00:00:00.000Z", "updatedAt": "2026-06-18T00:00:00.000Z"  }}
JSON8 lines

slug must be kebab-case and unique in the project. Some slugs (e.g. product-costs) are reserved and pin type/keyField (product-costs must be TABLE with keyField: "sku"); the API returns 400 validation_failed if you create one with the wrong shape, before it reaches the database:

400: reserved slug used with the wrong type
{ "error": { "code": "validation_failed", "message": "Request body failed validation", "details": { "issues": [{ "param": "type", "path": ["type"], "code": "custom", "message": "The \"product-costs\" slug is reserved and must use type TABLE" }] }, "docUrl": "https://docs.avsb.cloud/docs/developer-reference/public-api/conventions#validation-errors", "requestId": "req_9f2c41ab7e0b4d1e8c35a6f0d2b91e77" } }
JSON1 line

Get a dataset with its versions

GET /api/v1/projects/{projectId}/datasets/{datasetId}: the dataset plus its versions, newest first. Requires datasets:read.

curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId> \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{  "data": {    "dataset": { "id": "cmd...", "slug": "vip-customers", "liveVersionShortId": 23, "liveVersionNumber": 4, "...": "..." },    "versions": [      { "id": "cmv...", "shortId": 23, "number": 4, "datasetId": "cmd...", "label": "June import", "status": "ACTIVE",        "rowCount": 1204, "byteSize": 88231, "error": null, "createdById": null, "createdAt": "2026-06-18T00:00:00.000Z", "committedAt": "2026-06-18T00:01:00.000Z" }    ]  }}
JSON9 lines

shortId is the version's global serving address (never renumbered); number is its per-dataset "vN" label, the one shown in the dashboard.

404: dataset not found
{ "error": { "code": "not_found", "message": "Dataset not found", "docUrl": "https://docs.avsb.cloud/docs/developer-reference/public-api/conventions#not-found-errors", "requestId": "req_9f2c41ab7e0b4d1e8c35a6f0d2b91e77" } }
JSON1 line

Update a dataset

PATCH /api/v1/projects/{projectId}/datasets/{datasetId}: partial update of name, description, and piiPolicy only. Send only the fields you are changing. Requires datasets:write.

curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId> \  -X PATCH -H "Authorization: Bearer avsb_svc_..." -H "Content-Type: application/json" \  -d '{ "name": "VIP customers (2026)" }'
Shell3 lines
Response
{ "data": { "id": "cmd...", "name": "VIP customers (2026)", "...": "..." } }
JSON1 line

slug, type, and keyField cannot be changed. Sending any of them (the schema is strict about unknown fields) is refused rather than silently ignored:

400: slug is not editable
{ "error": { "code": "validation_failed", "message": "Request body failed validation", "details": { "issues": [{ "param": "(body)", "path": [], "code": "unrecognized_keys", "keys": ["slug"], "message": "Unrecognized key(s) in object: 'slug'" }] }, "docUrl": "https://docs.avsb.cloud/docs/developer-reference/public-api/conventions#validation-errors", "requestId": "req_9f2c41ab7e0b4d1e8c35a6f0d2b91e77" } }
JSON1 line

Delete a dataset

DELETE /api/v1/projects/{projectId}/datasets/{datasetId}: returns the deleted dataset's id. Requires datasets:write.

curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId> \  -X DELETE -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{ "data": { "id": "cmd..." } }
JSON1 line
Warning

SYSTEM datasets (managed by the recommendation engine) cannot be deleted; the request is refused with 409 and dataset_state_conflict.

409: SYSTEM dataset
{ "error": { "code": "dataset_state_conflict", "message": "System datasets are managed by the recommendation engine and cannot be deleted", "docUrl": "https://docs.avsb.cloud/docs/developer-reference/public-api/conventions#conflict-errors", "requestId": "req_9f2c41ab7e0b4d1e8c35a6f0d2b91e77" } }
JSON1 line

List versions

GET /api/v1/projects/{projectId}/datasets/{datasetId}/versions: cursor-paginated, newest first, up to 100 per page. Same request shape as List datasets above, just a longer path; ts/python only change the URL. Requires datasets:read.

cURL
curl 'https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions?limit=20' \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{  "data": [    { "id": "cmv...", "shortId": 23, "number": 4, "status": "ACTIVE", "rowCount": 1204, "...": "..." }  ],  "page": { "nextCursor": null, "hasMore": false }}
JSON6 lines

Create a version (get an upload URL)

POST /api/v1/projects/{projectId}/datasets/{datasetId}/versions: returns the new UPLOADING version and a presigned uploadUrl, 201 Created. label is optional, free text for your own reference. Requires datasets:write.

curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions \  -X POST -H "Authorization: Bearer avsb_svc_..." -H "Content-Type: application/json" \  -d '{ "label": "June import" }'
Shell3 lines
Response
{ "data": { "version": { "id": "cmv...", "shortId": 24, "number": 5, "status": "UPLOADING", "rowCount": 0, "byteSize": 0, "error": null, "...": "..." },    "uploadUrl": "https://<r2-host>/datasets/.../24.ndjson?X-Amz-Signature=..." } }
JSON2 lines

Then PUT your NDJSON rows straight to uploadUrl (a direct storage upload, not an AvsB API call, so the same request works from any language). The URL expires in one hour:

Shell
curl -X PUT "<uploadUrl>" -H "Content-Type: application/x-ndjson" --data-binary @rows.ndjson
Shell1 line

Discard a version

DELETE /api/v1/projects/{projectId}/datasets/{datasetId}/versions/{versionId}: discards a single non-ACTIVE version and reclaims its storage. Use it to clean up an abandoned or superseded upload. Same request shape as Delete a dataset above. Requires datasets:write.

cURL
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions/<versionId> \  -X DELETE -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{ "data": { "id": "cmv..." } }
JSON1 line

Discarding the live version returns the same 409 dataset_state_conflict shown under Delete a dataset above. Discarding one still COMMITTING returns the same code but at 412 instead of 409, an inconsistency confirmed in source rather than a documented status.

Commit a version

POST /api/v1/projects/{projectId}/datasets/{datasetId}/versions/{versionId}/commit: only UPLOADING versions can be committed; any other status is refused with 409 dataset_state_conflict. Same no-body request shape as above; only the URL and method change per language. Requires datasets:write.

cURL
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions/<versionId>/commit \  -X POST -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{ "data": { "id": "cmv...", "status": "COMMITTING", "...": "..." } }
JSON1 line

Ingestion runs asynchronously: poll GET .../versions (or the single dataset) until the version reads COMMITTED. On a bad row it reads FAILED instead, with the reason in the error object, for example { "line": 482, "reason": "missing required key field \"email\"" }.

Activate a version

POST /api/v1/projects/{projectId}/datasets/{datasetId}/versions/{versionId}/activate: only COMMITTED versions can be activated. This is also how you roll back: activate an older committed version. Requires datasets:write.

cURL
curl https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/versions/<versionId>/activate \  -X POST -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{ "data": { "dataset": { "id": "cmd...", "liveVersionId": "cmv...", "liveVersionShortId": 24, "liveVersionNumber": 5, "...": "..." },    "version": { "id": "cmv...", "status": "ACTIVE", "...": "..." } } }
JSON2 lines

Only the 5 most recently committed versions stay activatable. Older COMMITTED versions move to ARCHIVED and can no longer be rolled back to.

Test a key against the live version

GET /api/v1/projects/{projectId}/datasets/{datasetId}/lookup?key=...: reads a single key from the dataset's live version, exactly as the edge serves it. 404 when the dataset has no live version. Same query-string request shape as List datasets above. Requires datasets:read.

cURL
curl 'https://app.avsb.cloud/api/v1/projects/<projectId>/datasets/<datasetId>/lookup?key=user-42' \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines
Response
{ "data": { "found": true, "value": { "tier": "gold" }, "slug": "vip-customers", "version": 23, "versionOrdinal": 4 } }
JSON1 line

version is the live version's global serving address; versionOrdinal is its per-dataset "v4" label. When the key is absent the live version returns { "data": { "found": false, "slug": "vip-customers", "version": 23, "versionOrdinal": 4 } }.

404: no live version
{ "error": { "code": "not_found", "message": "Dataset has no live version", "docUrl": "https://docs.avsb.cloud/docs/developer-reference/public-api/conventions#not-found-errors", "requestId": "req_9f2c41ab7e0b4d1e8c35a6f0d2b91e77" } }
JSON1 line

Next steps

Was this helpful?