Troubleshooting Datasets

When an upload fails, the cause is almost always one of three things: the browser can't reach storage, the ingestion service can't be reached, or the shared secret between them doesn't match. The upload self-test checks all three by name so you don't have to guess.

Running the upload self-test

Open Projects → Commerce → Datasets and click Run self-test in the Upload self-test panel. It runs three checks and, for anything that fails, shows the exact fix.

CheckWhat it provesIf it fails
Storage upload (browser → R2)Your browser can upload directly to storageStorage isn't accepting uploads from the dashboard origin: the CORS rule on the uploads bucket needs the dashboard origin allowed for PUT
Ingestion worker reachableThe dashboard can reach the ingestion serviceThe worker URL isn't set, the worker isn't deployed, or a bot-protection rule is blocking the dashboard's requests
Shared secret validThe dashboard and worker agree on the internal secretThe secret is set on one side but not the other, or the two values differ

The storage check runs a real upload from your browser, then verifies it landed and cleans it up, so it exercises the exact path a real upload takes, including any cross-origin rules a server-side check would miss.

All green?

When every check passes, uploads are ready. If uploads still fail after a green self-test, the problem is in the file itself: check the version's failure reason in the version table (for example an invalid row, or a PII block).

A version is stuck on "Processing"

A version normally moves from Processing (committing) to Committed within seconds, or a couple of minutes for a large file. If it sits on Processing for much longer, its background indexing job was lost.

  • Wait for the recovery: a version that is still Processing cannot be discarded, because the upload is mid-write and removing the version would strand the rows it is still saving. A recovery job marks long-stuck versions as failed, and a failed version discards normally. Every recovery is recorded, so a stuck version is never silently abandoned.
  • Then discard and upload again: once the version reads Failed, use its Discard action and upload the file again.

A stuck version never serves and never becomes live (only committed versions can be activated), so nothing is at risk while you wait.

Confirm it's fixed: after you discard and re-upload, watch the version's status badge. It's fixed once the new version reaches Committed, with a row count next to it, within seconds to a couple of minutes.

An upload failed with a PII block

If a version fails with pii_blocked, its rows contained values that look like personal data: emails, phone numbers, card-like numbers, or postal codes. Datasets block PII by default because their rows are served from edge caches and read by client-side code.

  • Best fix: remove the personal data and key rows by opaque identifiers (user IDs, SKUs) instead.
  • If it's a false positive (or you're cleared to store the data): open the dataset's Edit dialog and switch its PII policy to Warn, confirming the acknowledgment. Warn lets the upload through and only flags the suspected rows; activating a flagged version then requires a second acknowledgment.

Numeric data such as SKUs, prices, and costs never trips the check.

Confirm it's fixed: upload the file again. On Block, it's fixed once the version reaches Committed with no failure text next to it. On Warn, it commits either way; look for a Possible PII badge next to the status to see whether any rows are still flagged.

Did discarding or deleting actually remove the data?

Discarding any non-live version reclaims its stored rows and uploaded file. Deleting a whole dataset reclaims all of its versions' storage (the rows behind every version and every uploaded file), so no orphaned storage is left behind. The reclaim runs in the background; if it can't be scheduled immediately, a nightly sweep reclaims the storage as a backstop.

Confirm it's fixed: the discarded version or deleted dataset disappears from the dashboard immediately, and the action appears in your organization's Audit log as a Deleted entry. There's nothing further you need to check: the storage cleanup behind it always follows within moments, or by the next nightly sweep at the latest.

Deletes are permanent

Deleting a dataset removes it and all of its versions. There is no undo. Roll back to an older version (rather than deleting) whenever you might need the current data again.

Was this helpful?