Public API: Results

The results API surfaces the same statistical output the AvsB dashboard shows (per-variation lift, metric results, health guardrails, and segment analysis), so you can pull experiment results into your own reporting pipelines or alerts.

All results endpoints authenticate with a service token. The org is taken from the token, so the path carries only {experimentId}, not {orgId}.

Plain text
https://app.avsb.cloud/api/v1/experiments/{experimentId}/results
Plain text1 line

Scopes

OperationScope
Get experiment resultsresults:read
Compare all enginesresults:read

Every /api/v1 token is rate-limited: a scoped token gets 600 reads and 120 writes per minute, and an admin:* token gets 600 requests per minute, reads and writes together. See Conventions.

Info

Results are read-only on the public API: there is no results:write. Results are computed from raw event data; they cannot be written through this API.

Info

Both endpoints also need the integrations_api_export plan feature. results:read gets you past the scope check. But the call still refuses with 403 feature_disabled if your org's plan does not include API export. See Refusals at 403.

Get experiment results

GET /api/v1/experiments/{experimentId}/results: statistical results for a single experiment under its official engine (or a requested engine override).

curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results" \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines

Add query parameters for a specific window, a non-default engine, and a segment slice:

Fuller request
curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results?from=2026-07-01&to=2026-07-31&engine=FREQUENTIST&segment=device:mobile" \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines

Query parameters

ParameterTypeDescription
fromstringStart of the analysis window, as a UTC calendar day (YYYY-MM-DD). Omit to use the experiment's launch date.
tostringEnd of the analysis window, as a UTC calendar day (YYYY-MM-DD), inclusive. Omit to use now.
engineBAYESIAN | FREQUENTIST | SEQUENTIALOverride the engine for this fetch. Defaults to the experiment's official engine.
revenueModenet | grossRevenue attribution mode for PURCHASE bindings. net (default) subtracts refunds; gross does not.
costCoveragecovered_only | include_allFor profit bindings (Phase 8): covered_only (default) drops orders with incomplete cost data; include_all keeps them.
segmentstring (repeatable)Filter results to a segment slice. Format: key:value (e.g. device:mobile). Repeat the same key to match any of its values (OR); repeat with a different key to narrow further (AND).

from and to are UTC calendar days, both ends inclusive, never local ones: from=2026-04-01 means 2026-04-01T00:00:00Z, and to=2026-04-30 covers all of the 30th in UTC, whatever time zone you're reading from. That's what keeps two people looking at the same report seeing the same numbers. A from or to that isn't a real YYYY-MM-DD date, or a from later than to, is refused with 400 validation_failed before any query runs (see Errors).

Response

JSON
{  "data": {    "experimentId": "cldabc123",    "dateRange": {      "from": "2026-01-15",      "to": "2026-06-01"    },    "officialEngine": "BAYESIAN",    "renderedEngine": "BAYESIAN",    "isAATest": false,    "variations": [      {        "variationId": "ctrl",        "variationName": "Control",        "visitors": 12400,        "conversions": 1240,        "conversionRate": 0.1,        "improvementPct": null,        "probabilityToBeatControl": null,        "credibleIntervalLow": null,        "credibleIntervalHigh": null,        "goalAlignedLiftPct": null,        "probabilityBetter": null,        "relativeLiftCi": null,        "metricDirection": "INCREASE",        "totalRevenue": 0,        "aov": null      },      {        "variationId": "var-a",        "variationName": "Variant A",        "visitors": 12350,        "conversions": 1408,        "conversionRate": 0.114,        "improvementPct": 0.14,        "probabilityToBeatControl": 0.97,        "credibleIntervalLow": 0.05,        "credibleIntervalHigh": 0.22,        "goalAlignedLiftPct": 0.14,        "probabilityBetter": 0.97,        "relativeLiftCi": { "lower": 0.05, "upper": 0.23 },        "metricDirection": "INCREASE",        "totalRevenue": 0,        "aov": null      }    ],    "timeSeries": [],    "metrics": [],    "healthScore": {      "dataQuality": { "status": "green", "description": "No SRM detected" },      "statisticalConfidence": { "status": "green", "description": "97% probability to beat control" },      "trafficHealth": { "status": "green", "description": "Traffic split looks healthy" }    },    "availableSegments": [      { "key": "device", "name": "Device", "values": ["mobile", "desktop", "tablet"], "isDefault": true }    ],    "analysisConfig": {      "engine": "BAYESIAN",      "varianceReduction": "AUTO",      "customAlpha": 0.05    },    "segmentLift": null,    "computedAt": "2026-01-08T10:42:00.000Z",    "peeks": {      "looks": 7,      "distinctDays": 4,      "countedToday": false,      "firstPeekAt": "2026-01-02T09:00:00.000Z",      "lastPeekAt": "2026-01-08T10:42:00.000Z"    }  }}
JSON72 lines

Field notes

  • visitors and conversions are exact counts, not estimates. The server deduplicates by visitorId at query time, using ClickHouse's exact uniqExact rather than its approximate uniq. A visitor exposed twice, or converting twice, is still counted once.
  • improvementPct is a fraction (0.14 = +14%). null on the control row. This is the raw engine lift: it is always "did the number go up?", never "did the metric get better?". For a decision, read goalAlignedLiftPct.
  • probabilityToBeatControl is 0–1; null on the control row and on frequentist/sequential engines (use isSignificant instead). Like improvementPct, this is raw: it is P(variation > control) regardless of which direction is good.
  • isSignificant is a two-sided test: true means the variation differs from control, not that it beat it. A significant loser also reports true. To decide a winner, combine it with goalAlignedLiftPct > 0 (the direction-normalized lift below), not with the raw sign of improvementPct.
  • pValueRaw (optional) is the raw two-sided p-value before any multiple-comparison correction. When your experiment has more than two variations the headline pValue and isSignificant reflect the correction across the metric family; pValueRaw is the uncorrected value. It is present for binary metrics that carry a p-value, and null/absent on quantile and ratio metrics and on older cached payloads. The compare-engines view reads the gap between pValueRaw and the corrected pValue to explain when a correction is the reason the Frequentist column disagrees.
  • credibleIntervalLow / credibleIntervalHigh are the 95% credible interval bounds on improvementPct. Bayesian engine only; null otherwise.

Goal-aligned fields (direction-normalized)

Some metrics are better when they go down (bounce rate, load time, refunds). The four fields below apply the metric's direction for you, so that positive always means good, on every metric. Read these to decide; read the raw twins above only if you want the engine's unmodified output.

  • goalAlignedLiftPct is improvementPct normalized for direction. For a metricDirection: "INCREASE" metric it equals improvementPct; for "DECREASE" it is negated. A bounce rate falling 20% (improvementPct: -0.20) reports goalAlignedLiftPct: 0.20, a 20% win. null on the control row.
  • probabilityBetter is the goal-aligned win probability: the chance this variation is genuinely better at the metric's own goal. Equals probabilityToBeatControl for a higher-is-better metric, and its complement (1 - p) for a lower-is-better one: an 8% chance the bounce rate rose is a 92% chance it genuinely dropped. null on the control row and on the frequentist/sequential engines.
  • relativeLiftCi is the goal-aligned confidence/credible interval around the size of the lift, as { "lower": number, "upper": number } with lower <= upper. This is the uncertainty on the thing you actually decide on (how big the improvement is), not on the raw rates. It is null on the control row, and null when the engine could not form an interval (a zero-observation arm, or a degenerate/unidentifiable case). Treat null as "no interval available" and do not infer a value from the other fields.
  • metricDirection is "INCREASE" or "DECREASE": which direction counts as a win, and the direction the three fields above were normalized under.

Each metric in metrics also carries direction (the same value, per metric) and cupedApplied, a boolean that is true only when that metric's displayed numbers were actually computed from the CUPED-adjusted series. It is deliberately narrower than reductionDecision.applied: a variance-reduction decision is computed for every metric, but most metrics render raw numbers regardless, so cupedApplied is the field to trust for "were these numbers adjusted?".

  • computedAt is an ISO-8601 timestamp of when the server assembled this payload. It is the freshness of these numbers, computed live on every request.
  • totalRevenue and aov are in minor units of the project's currency (cents for USD, yen for JPY, fils for KWD). Divide by the currency's ISO-4217 exponent before display: 499900 USD is $4,999.00. aov may be fractional.
  • healthScore statuses are green, yellow, or red. A red dataQuality means a sample-ratio mismatch (SRM) was detected, so results should be treated with caution.
  • weightsChangedAt is the ISO-8601 instant of the last time this experiment's traffic split changed, or null when it has never changed. When set, the sample-ratio-mismatch check behind healthScore.dataQuality only counts visitors first exposed after that instant, so changing the split partway through doesn't itself trigger a false SRM alarm.
  • renderedEngine reflects the engine actually used for this response (the engine override if supplied, otherwise officialEngine).
  • timeSeries contains 4-hour buckets for the charted conversion-rate over time. Empty when there is insufficient data or no date-range overlap. A bucket can carry null for a variation with no traffic in that window: null means there's no rate to report, not a 0% conversion rate, so don't chart a gap as a zero. The same applies to each metric's own timeSeries array under metrics.
  • metrics contains per-metric result breakdowns (same structure per metric: variationId, conversions, conversionRate, improvementPct, etc.).
  • segmentLift is null unless the experiment has a primary metric and segment data. When present, it contains per-segment heterogeneous-treatment-effect analysis.

Fields the example above omits

The example response leaves out six fields that are real, but each only appears under its own condition. All six are omitted entirely (not null) when their condition is not met, except where the table says otherwise:

FieldAppears whenWhat it is
trustGradeOfficial-engine pass (no ?engine= override) with real data. null when official but ungradeable.The A to F trust grade: a 100-point rubric covering pre-registration, sample size, peeking, traffic health, and runtime.
honestLiftOfficial-engine pass whose leading arm crossed the engine's decision boundary. null on an official pass where nothing was selected. Absent on an ?engine= override.The winner's-curse-corrected ("honest") lift for the leading arm's primary metric: shrunkLiftPct alongside the raw observedLiftPct, plus wrongSignRisk and exaggerationRatio.
decisionRiskOfficial-engine pass whose primary metric has a money basis, with at least 3 distinct exposure days. null when computed but suppressed for thinner data.The "risk of shipping / upside of waiting" monthly money projection, in the currency's minor units.
guardrailSummaryOfficial-engine pass on an experiment with at least one guardrail-role metric.Breached-guardrail counts. Any breached > 0 means the experiment cannot verdict as "Ship".
excludedAny bot, preview-link, or internal traffic was filtered out of these numbers.A count per exclusion reason (bot, preview, internal).
trafficFilterUnavailableThe traffic-source column needed to filter bot/preview/internal traffic was missing, so this response could not exclude it.Always true when present; its absence means filtering ran normally.
recommendationThe experiment's variation code fired rec:impression or rec:click events (the headless recommendations API).Per-variation recommendation stats: impressions, clicks, click-through rate, and add-to-cart follow-through.

decisionRisk, guardrailSummary, trustGrade, and honestLift are official-decision surfaces. Each is computed once, against the experiment's official engine, not per exploratory view. So all four are withheld on an ?engine= override, and on every row of Compare all engines. recommendation is also absent from compare responses; it is only ever added on the single-results endpoint. excluded and trafficFilterUnavailable are the exception: both are data-quality disclosures, not decision surfaces, so they still appear on every compare row when they apply.

peeks: how often the results were looked at

peeks reports how many times a human checked this experiment's live results before it reached its decision point, in distinct user-days (one look per person per UTC day). It is what the dashboard's peek counter shows, and it feeds the experiment's trust grade.

  • looks: distinct user-days recorded.
  • distinctDays: how many separate UTC days anyone looked.
  • countedToday: always false on this endpoint (see below).
  • firstPeekAt / lastPeekAt: ISO-8601 instants, or null when nothing has been recorded.

Three properties of this field are worth reading carefully:

  • Reading results through this API is never counted. A token has no human behind it, and an integration polling on a schedule is not somebody making repeated decisions, so this request does not add to looks, and countedToday is always false here. Only dashboard views are counted. (Inside the dashboard the count does include the view being served, which is why countedToday exists at all.)
  • null is not zero. "peeks": null means the ledger could not be read, so the count is unknown. Genuinely-never-looked-at is {"looks": 0, ...}. Do not collapse the two: reporting "nobody has looked" on the strength of a failed query is exactly the kind of confident wrongness this field exists to avoid.
  • The field may be absent on payloads predating the feature, and counting starts at the first recorded look; history before that was never recorded and is not backfilled.

Compare all engines

GET /api/v1/experiments/{experimentId}/results/compare: results computed under all three engines (BAYESIAN, FREQUENTIST, SEQUENTIAL) in a single request, useful for comparing how engine choice affects conclusions.

Shell
curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results/compare" \  -H "Authorization: Bearer avsb_svc_..."
Shell2 lines

Supports the same from, to, and segment query parameters as the single-results endpoint. Does not accept engine (all three are always returned), revenueMode, or costCoverage.

Response

JSON
{  "data": {    "experimentId": "cldabc123",    "dateRange": {      "from": "2026-01-15",      "to": null    },    "officialEngine": "BAYESIAN",    "segmentLift": null,    "engines": {      "BAYESIAN": { "...": "full ExperimentResults object" },      "FREQUENTIST": { "...": "full ExperimentResults object" },      "SEQUENTIAL": { "...": "full ExperimentResults object" }    }  }}
JSON16 lines

Field notes

  • engines contains three ExperimentResults objects, one per engine, but each is a narrower "exploratory" cut of the single-results shape. peeks, recommendation, trustGrade, honestLift, decisionRisk, and guardrailSummary are never present on any of the three: those are official-decision surfaces computed once, not per engine. See Fields the example above omits for what each one is.
  • segmentLift is computed once against the officialEngine and shared at the top level; the per-engine entries all carry segmentLift: null (HTE is engine-independent for comparison).
  • officialEngine echoes which engine the experiment is configured to use as its decisive result.

Errors

Both endpoints (and their flag-rule twins) share these result-specific errors.

Info

No results endpoint caps how wide a from/to window you ask for. Your plan's retention window narrows it, and the window actually read comes back as dateRange in the response. The two endpoints on this page and both flag-rule twins behave the same way, so a range one of them answers is a range all of them answer.

400 validation_failed

Returned when from or to isn't a valid YYYY-MM-DD date, or when from is later than to. This is a request-shape problem: the query never reaches ClickHouse.

JSON
{  "error": {    "code": "validation_failed",    "message": "Invalid from date"  }}
JSON6 lines

The exact message tells you which bound was the problem: "Invalid from date", "Invalid to date", or "from must not be after to".

503 results_unavailable

The analytics store could not serve the queries behind the decision numbers.

JSON
{  "error": {    "code": "results_unavailable",    "message": "Results are temporarily unavailable. Please retry."  }}
JSON6 lines

This is a transient condition: retry with backoff. It is returned instead of a 200 with zeroed counts, so a 200 response always means the numbers are real. Treat results_unavailable as "unknown", never as "no conversions".

Descriptive extras degrade quietly rather than failing the whole request: if only the time-series or segment-discovery queries fail, you still get a 200 with the full decision numbers and an empty timeSeries array.

404 not_found

Returned when the experiment does not exist or is not in your token's organization. Cross-organization reads are indistinguishable from missing experiments by design, since a 403 would confirm that an ID exists.

Next steps

Was this helpful?