Public API: Results
The results API surfaces the same statistical output the AvsB dashboard shows (per-variation lift, metric results, health guardrails, and segment analysis), so you can pull experiment results into your own reporting pipelines or alerts.
All results endpoints authenticate with a service token. The org is taken from the token, so the path carries only {experimentId}, not {orgId}.
https://app.avsb.cloud/api/v1/experiments/{experimentId}/resultsScopes
| Operation | Scope |
|---|---|
| Get experiment results | results:read |
| Compare all engines | results:read |
Every /api/v1 token is rate-limited: a scoped token gets 600 reads and 120 writes per minute, and an admin:* token gets 600 requests per minute, reads and writes together. See Conventions.
Results are read-only on the public API: there is no results:write. Results are computed from raw event data; they cannot be written through this API.
Both endpoints also need the integrations_api_export plan feature. results:read gets you past the scope check. But the call still refuses with 403 feature_disabled if your org's plan does not include API export. See Refusals at 403.
Get experiment results
GET /api/v1/experiments/{experimentId}/results: statistical results for a single experiment under its official engine (or a requested engine override).
curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results" \ -H "Authorization: Bearer avsb_svc_..."const res = await fetch('https://app.avsb.cloud/api/v1/experiments/<experimentId>/results', { headers: { Authorization: `Bearer ${process.env.AVSB_SERVICE_TOKEN}` },})const { data } = await res.json()import os, requestsr = requests.get("https://app.avsb.cloud/api/v1/experiments/<experimentId>/results", headers={"Authorization": f"Bearer {os.environ['AVSB_SERVICE_TOKEN']}"})data = r.json()["data"]Add query parameters for a specific window, a non-default engine, and a segment slice:
curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results?from=2026-07-01&to=2026-07-31&engine=FREQUENTIST&segment=device:mobile" \ -H "Authorization: Bearer avsb_svc_..."Query parameters
| Parameter | Type | Description |
|---|---|---|
from | string | Start of the analysis window, as a UTC calendar day (YYYY-MM-DD). Omit to use the experiment's launch date. |
to | string | End of the analysis window, as a UTC calendar day (YYYY-MM-DD), inclusive. Omit to use now. |
engine | BAYESIAN | FREQUENTIST | SEQUENTIAL | Override the engine for this fetch. Defaults to the experiment's official engine. |
revenueMode | net | gross | Revenue attribution mode for PURCHASE bindings. net (default) subtracts refunds; gross does not. |
costCoverage | covered_only | include_all | For profit bindings (Phase 8): covered_only (default) drops orders with incomplete cost data; include_all keeps them. |
segment | string (repeatable) | Filter results to a segment slice. Format: key:value (e.g. device:mobile). Repeat the same key to match any of its values (OR); repeat with a different key to narrow further (AND). |
from and to are UTC calendar days, both ends inclusive, never local ones: from=2026-04-01 means 2026-04-01T00:00:00Z, and to=2026-04-30 covers all of the 30th in UTC, whatever time zone you're reading from. That's what keeps two people looking at the same report seeing the same numbers. A from or to that isn't a real YYYY-MM-DD date, or a from later than to, is refused with 400 validation_failed before any query runs (see Errors).
Response
{ "data": { "experimentId": "cldabc123", "dateRange": { "from": "2026-01-15", "to": "2026-06-01" }, "officialEngine": "BAYESIAN", "renderedEngine": "BAYESIAN", "isAATest": false, "variations": [ { "variationId": "ctrl", "variationName": "Control", "visitors": 12400, "conversions": 1240, "conversionRate": 0.1, "improvementPct": null, "probabilityToBeatControl": null, "credibleIntervalLow": null, "credibleIntervalHigh": null, "goalAlignedLiftPct": null, "probabilityBetter": null, "relativeLiftCi": null, "metricDirection": "INCREASE", "totalRevenue": 0, "aov": null }, { "variationId": "var-a", "variationName": "Variant A", "visitors": 12350, "conversions": 1408, "conversionRate": 0.114, "improvementPct": 0.14, "probabilityToBeatControl": 0.97, "credibleIntervalLow": 0.05, "credibleIntervalHigh": 0.22, "goalAlignedLiftPct": 0.14, "probabilityBetter": 0.97, "relativeLiftCi": { "lower": 0.05, "upper": 0.23 }, "metricDirection": "INCREASE", "totalRevenue": 0, "aov": null } ], "timeSeries": [], "metrics": [], "healthScore": { "dataQuality": { "status": "green", "description": "No SRM detected" }, "statisticalConfidence": { "status": "green", "description": "97% probability to beat control" }, "trafficHealth": { "status": "green", "description": "Traffic split looks healthy" } }, "availableSegments": [ { "key": "device", "name": "Device", "values": ["mobile", "desktop", "tablet"], "isDefault": true } ], "analysisConfig": { "engine": "BAYESIAN", "varianceReduction": "AUTO", "customAlpha": 0.05 }, "segmentLift": null, "computedAt": "2026-01-08T10:42:00.000Z", "peeks": { "looks": 7, "distinctDays": 4, "countedToday": false, "firstPeekAt": "2026-01-02T09:00:00.000Z", "lastPeekAt": "2026-01-08T10:42:00.000Z" } }}Field notes
visitorsandconversionsare exact counts, not estimates. The server deduplicates byvisitorIdat query time, using ClickHouse's exactuniqExactrather than its approximateuniq. A visitor exposed twice, or converting twice, is still counted once.improvementPctis a fraction (0.14= +14%).nullon the control row. This is the raw engine lift: it is always "did the number go up?", never "did the metric get better?". For a decision, readgoalAlignedLiftPct.probabilityToBeatControlis 0–1;nullon the control row and on frequentist/sequential engines (useisSignificantinstead). LikeimprovementPct, this is raw: it isP(variation > control)regardless of which direction is good.isSignificantis a two-sided test:truemeans the variation differs from control, not that it beat it. A significant loser also reportstrue. To decide a winner, combine it withgoalAlignedLiftPct > 0(the direction-normalized lift below), not with the raw sign ofimprovementPct.pValueRaw(optional) is the raw two-sided p-value before any multiple-comparison correction. When your experiment has more than two variations the headlinepValueandisSignificantreflect the correction across the metric family;pValueRawis the uncorrected value. It is present for binary metrics that carry a p-value, andnull/absent on quantile and ratio metrics and on older cached payloads. The compare-engines view reads the gap betweenpValueRawand the correctedpValueto explain when a correction is the reason the Frequentist column disagrees.credibleIntervalLow/credibleIntervalHighare the 95% credible interval bounds onimprovementPct. Bayesian engine only;nullotherwise.
Goal-aligned fields (direction-normalized)
Some metrics are better when they go down (bounce rate, load time, refunds). The four fields below apply the metric's direction for you, so that positive always means good, on every metric. Read these to decide; read the raw twins above only if you want the engine's unmodified output.
goalAlignedLiftPctisimprovementPctnormalized for direction. For ametricDirection: "INCREASE"metric it equalsimprovementPct; for"DECREASE"it is negated. A bounce rate falling 20% (improvementPct: -0.20) reportsgoalAlignedLiftPct: 0.20, a 20% win.nullon the control row.probabilityBetteris the goal-aligned win probability: the chance this variation is genuinely better at the metric's own goal. EqualsprobabilityToBeatControlfor a higher-is-better metric, and its complement (1 - p) for a lower-is-better one: an 8% chance the bounce rate rose is a 92% chance it genuinely dropped.nullon the control row and on the frequentist/sequential engines.relativeLiftCiis the goal-aligned confidence/credible interval around the size of the lift, as{ "lower": number, "upper": number }withlower <= upper. This is the uncertainty on the thing you actually decide on (how big the improvement is), not on the raw rates. It isnullon the control row, andnullwhen the engine could not form an interval (a zero-observation arm, or a degenerate/unidentifiable case). Treatnullas "no interval available" and do not infer a value from the other fields.metricDirectionis"INCREASE"or"DECREASE": which direction counts as a win, and the direction the three fields above were normalized under.
Each metric in metrics also carries direction (the same value, per metric) and cupedApplied, a boolean that is true only when that metric's displayed numbers were actually computed from the CUPED-adjusted series. It is deliberately narrower than reductionDecision.applied: a variance-reduction decision is computed for every metric, but most metrics render raw numbers regardless, so cupedApplied is the field to trust for "were these numbers adjusted?".
computedAtis an ISO-8601 timestamp of when the server assembled this payload. It is the freshness of these numbers, computed live on every request.totalRevenueandaovare in minor units of the project's currency (cents for USD, yen for JPY, fils for KWD). Divide by the currency's ISO-4217 exponent before display:499900USD is $4,999.00.aovmay be fractional.healthScorestatuses aregreen,yellow, orred. AreddataQualitymeans a sample-ratio mismatch (SRM) was detected, so results should be treated with caution.weightsChangedAtis the ISO-8601 instant of the last time this experiment's traffic split changed, ornullwhen it has never changed. When set, the sample-ratio-mismatch check behindhealthScore.dataQualityonly counts visitors first exposed after that instant, so changing the split partway through doesn't itself trigger a false SRM alarm.renderedEnginereflects the engine actually used for this response (theengineoverride if supplied, otherwiseofficialEngine).timeSeriescontains 4-hour buckets for the charted conversion-rate over time. Empty when there is insufficient data or no date-range overlap. A bucket can carrynullfor a variation with no traffic in that window:nullmeans there's no rate to report, not a 0% conversion rate, so don't chart a gap as a zero. The same applies to each metric's owntimeSeriesarray undermetrics.metricscontains per-metric result breakdowns (same structure per metric:variationId,conversions,conversionRate,improvementPct, etc.).segmentLiftisnullunless the experiment has a primary metric and segment data. When present, it contains per-segment heterogeneous-treatment-effect analysis.
Fields the example above omits
The example response leaves out six fields that are real, but each only appears under its own condition. All six are omitted entirely (not null) when their condition is not met, except where the table says otherwise:
| Field | Appears when | What it is |
|---|---|---|
trustGrade | Official-engine pass (no ?engine= override) with real data. null when official but ungradeable. | The A to F trust grade: a 100-point rubric covering pre-registration, sample size, peeking, traffic health, and runtime. |
honestLift | Official-engine pass whose leading arm crossed the engine's decision boundary. null on an official pass where nothing was selected. Absent on an ?engine= override. | The winner's-curse-corrected ("honest") lift for the leading arm's primary metric: shrunkLiftPct alongside the raw observedLiftPct, plus wrongSignRisk and exaggerationRatio. |
decisionRisk | Official-engine pass whose primary metric has a money basis, with at least 3 distinct exposure days. null when computed but suppressed for thinner data. | The "risk of shipping / upside of waiting" monthly money projection, in the currency's minor units. |
guardrailSummary | Official-engine pass on an experiment with at least one guardrail-role metric. | Breached-guardrail counts. Any breached > 0 means the experiment cannot verdict as "Ship". |
excluded | Any bot, preview-link, or internal traffic was filtered out of these numbers. | A count per exclusion reason (bot, preview, internal). |
trafficFilterUnavailable | The traffic-source column needed to filter bot/preview/internal traffic was missing, so this response could not exclude it. | Always true when present; its absence means filtering ran normally. |
recommendation | The experiment's variation code fired rec:impression or rec:click events (the headless recommendations API). | Per-variation recommendation stats: impressions, clicks, click-through rate, and add-to-cart follow-through. |
decisionRisk, guardrailSummary, trustGrade, and honestLift are official-decision surfaces. Each is computed once, against the experiment's official engine, not per exploratory view. So all four are withheld on an ?engine= override, and on every row of Compare all engines. recommendation is also absent from compare responses; it is only ever added on the single-results endpoint. excluded and trafficFilterUnavailable are the exception: both are data-quality disclosures, not decision surfaces, so they still appear on every compare row when they apply.
peeks: how often the results were looked at
peeks reports how many times a human checked this experiment's live results before it reached its decision point, in distinct user-days (one look per person per UTC day). It is what the dashboard's peek counter shows, and it feeds the experiment's trust grade.
looks: distinct user-days recorded.distinctDays: how many separate UTC days anyone looked.countedToday: alwaysfalseon this endpoint (see below).firstPeekAt/lastPeekAt: ISO-8601 instants, ornullwhen nothing has been recorded.
Three properties of this field are worth reading carefully:
- Reading results through this API is never counted. A token has no human behind it, and an integration polling on a schedule is not somebody making repeated decisions, so this request does not add to
looks, andcountedTodayis alwaysfalsehere. Only dashboard views are counted. (Inside the dashboard the count does include the view being served, which is whycountedTodayexists at all.) nullis not zero."peeks": nullmeans the ledger could not be read, so the count is unknown. Genuinely-never-looked-at is{"looks": 0, ...}. Do not collapse the two: reporting "nobody has looked" on the strength of a failed query is exactly the kind of confident wrongness this field exists to avoid.- The field may be absent on payloads predating the feature, and counting starts at the first recorded look; history before that was never recorded and is not backfilled.
Compare all engines
GET /api/v1/experiments/{experimentId}/results/compare: results computed under all three engines (BAYESIAN, FREQUENTIST, SEQUENTIAL) in a single request, useful for comparing how engine choice affects conclusions.
curl "https://app.avsb.cloud/api/v1/experiments/<experimentId>/results/compare" \ -H "Authorization: Bearer avsb_svc_..."Supports the same from, to, and segment query parameters as the single-results endpoint. Does not accept engine (all three are always returned), revenueMode, or costCoverage.
Response
{ "data": { "experimentId": "cldabc123", "dateRange": { "from": "2026-01-15", "to": null }, "officialEngine": "BAYESIAN", "segmentLift": null, "engines": { "BAYESIAN": { "...": "full ExperimentResults object" }, "FREQUENTIST": { "...": "full ExperimentResults object" }, "SEQUENTIAL": { "...": "full ExperimentResults object" } } }}Field notes
enginescontains threeExperimentResultsobjects, one per engine, but each is a narrower "exploratory" cut of the single-results shape.peeks,recommendation,trustGrade,honestLift,decisionRisk, andguardrailSummaryare never present on any of the three: those are official-decision surfaces computed once, not per engine. See Fields the example above omits for what each one is.segmentLiftis computed once against theofficialEngineand shared at the top level; the per-engine entries all carrysegmentLift: null(HTE is engine-independent for comparison).officialEngineechoes which engine the experiment is configured to use as its decisive result.
Errors
Both endpoints (and their flag-rule twins) share these result-specific errors.
No results endpoint caps how wide a from/to window you ask for. Your plan's retention window narrows it, and the window actually read comes back as dateRange in the response. The two endpoints on this page and both flag-rule twins behave the same way, so a range one of them answers is a range all of them answer.
400 validation_failed
Returned when from or to isn't a valid YYYY-MM-DD date, or when from is later than to. This is a request-shape problem: the query never reaches ClickHouse.
{ "error": { "code": "validation_failed", "message": "Invalid from date" }}The exact message tells you which bound was the problem: "Invalid from date", "Invalid to date", or "from must not be after to".
503 results_unavailable
The analytics store could not serve the queries behind the decision numbers.
{ "error": { "code": "results_unavailable", "message": "Results are temporarily unavailable. Please retry." }}This is a transient condition: retry with backoff. It is returned instead of a 200 with zeroed counts, so a 200 response always means the numbers are real. Treat results_unavailable as "unknown", never as "no conversions".
Descriptive extras degrade quietly rather than failing the whole request: if only the time-series or segment-discovery queries fail, you still get a 200 with the full decision numbers and an empty timeSeries array.
404 not_found
Returned when the experiment does not exist or is not in your token's organization. Cross-organization reads are indistinguishable from missing experiments by design, since a 403 would confirm that an ID exists.