Troubleshooting Datasets
When an upload fails, the cause is almost always one of three things: the browser can't reach storage, the ingestion service can't be reached, or the shared secret between them doesn't match. The upload self-test checks all three by name so you don't have to guess.
Running the upload self-test
Open Projects → Commerce → Datasets and click Run self-test in the Upload self-test panel. It runs three checks and, for anything that fails, shows the exact fix.
| Check | What it proves | If it fails |
|---|---|---|
| Storage upload (browser → R2) | Your browser can upload directly to storage | Storage isn't accepting uploads from the dashboard origin: the CORS rule on the uploads bucket needs the dashboard origin allowed for PUT |
| Ingestion worker reachable | The dashboard can reach the ingestion service | The worker URL isn't set, the worker isn't deployed, or a bot-protection rule is blocking the dashboard's requests |
| Shared secret valid | The dashboard and worker agree on the internal secret | The secret is set on one side but not the other, or the two values differ |
The storage check runs a real upload from your browser, then verifies it landed and cleans it up, so it exercises the exact path a real upload takes, including any cross-origin rules a server-side check would miss.
When every check passes, uploads are ready. If uploads still fail after a green self-test, the problem is in the file itself: check the version's failure reason in the version table (for example an invalid row, or a PII block).
A version is stuck on "Processing"
A version normally moves from Processing (committing) to Committed within seconds, or a couple of minutes for a large file. If it sits on Processing for much longer, its background indexing job was lost.
- Wait for the recovery: a version that is still Processing cannot be discarded, because the upload is mid-write and removing the version would strand the rows it is still saving. A recovery job marks long-stuck versions as failed, and a failed version discards normally. Every recovery is recorded, so a stuck version is never silently abandoned.
- Then discard and upload again: once the version reads Failed, use its Discard action and upload the file again.
A stuck version never serves and never becomes live (only committed versions can be activated), so nothing is at risk while you wait.
Confirm it's fixed: after you discard and re-upload, watch the version's status badge. It's fixed once the new version reaches Committed, with a row count next to it, within seconds to a couple of minutes.
An upload failed with a PII block
If a version fails with pii_blocked, its rows contained values that look like personal data: emails, phone numbers, card-like numbers, or postal codes. Datasets block PII by default because their rows are served from edge caches and read by client-side code.
- Best fix: remove the personal data and key rows by opaque identifiers (user IDs, SKUs) instead.
- If it's a false positive (or you're cleared to store the data): open the dataset's Edit dialog and switch its PII policy to Warn, confirming the acknowledgment. Warn lets the upload through and only flags the suspected rows; activating a flagged version then requires a second acknowledgment.
Numeric data such as SKUs, prices, and costs never trips the check.
Confirm it's fixed: upload the file again. On Block, it's fixed once the version reaches Committed with no failure text next to it. On Warn, it commits either way; look for a Possible PII badge next to the status to see whether any rows are still flagged.
Did discarding or deleting actually remove the data?
Discarding any non-live version reclaims its stored rows and uploaded file. Deleting a whole dataset reclaims all of its versions' storage (the rows behind every version and every uploaded file), so no orphaned storage is left behind. The reclaim runs in the background; if it can't be scheduled immediately, a nightly sweep reclaims the storage as a backstop.
Confirm it's fixed: the discarded version or deleted dataset disappears from the dashboard immediately, and the action appears in your organization's Audit log as a Deleted entry. There's nothing further you need to check: the storage cleanup behind it always follows within moments, or by the next nightly sweep at the latest.
Deleting a dataset removes it and all of its versions. There is no undo. Roll back to an older version (rather than deleting) whenever you might need the current data again.