Viewing Rule Results

Every A/B test rule has a dedicated results page. Open it by clicking Results on a running rule. For each variation (one specific version being tested, control or a challenger) the page shows how many visitors it got and how many converted. Below that sit a time-series chart of the conversion rate, a summary of winning probability and lift, and a health-guardrails section. That last section flags problems. It watches for SRM (sample ratio mismatch, a red flag that fires when the traffic split between variations does not match what you configured), low traffic, and statistical confidence concerns.

Opening the results page

Results are available for any rule whose type is A/B Test, at any status. You can only reach the page from inside the app while the rule is Running, though:

  • From the ruleset list, click the Results button on the rule's card.
  • From the rule detail panel, click View results in the actions menu.

Both buttons disappear once you pause or conclude the rule. The page itself keeps working after that. A Concluded badge appears next to the rule name, and the numbers stay exactly as they were. Bookmark the results URL while the rule is running, so you can still find it later.

Ready rules start with their environment

A rule sitting at Ready goes live automatically when its environment starts: it is promoted to Running in the same step and begins collecting the visitors this page reports on. A rule-level Run click is only needed when the environment is already running.

Both entry points open a dedicated page at /projects/<projectId>/flags/<flagId>/rules/<ruleId>/results?env=<envKey>. The breadcrumb back-link returns you to the flag detail, with the same env tab pre-selected.

Opening that address for a targeted delivery rule with no metrics attached shows "This rule does not produce results". Nothing is wrong: that rule measures nothing, so waiting for visitors would never fill the page.

An A/B test rule with no metric yet still shows who was exposed: a chart of visitors per variation over time, under the note "Add a metric to this rule to see conversions". "No data collected yet" means no visitor has been exposed to the rule in the selected dates.

  1. Export CSV downloads every variation across every metric, named after the rule. It stays greyed out until the rule has metric results to export.
  2. Baseline changes which variation the table below compares against. The cards above always stay relative to control.
  3. Observed Lift vs control is a gross total across the non-control variations. It is not a projected or an incremental figure.

Once a rule concludes or stops early

Concluding a rule writes its own entry in the audit log, just like an experiment. A normal conclusion logs a Completed entry, and stopping early under the Frequentist engine logs an Early Stopped entry naming the engine, the reason you typed, and the winning variation. Either appears alongside the usual "Rule updated" entry for the publish, so the audit log tells you exactly when and why a rule concluded.

Locked settings are enforced everywhere

A rule's stats engine and variance-reduction settings lock once it launches, and a Concluded rule is read-only. Both locks are enforced on every write path, including the API and CLI: a request that changes a launched rule's analysis settings, or edits a concluded rule at all, is rejected. A mid-run change would make the numbers on this page describe two different tests spliced together, so the platform refuses it.

Reading the summary cards

Two or three summary cards sit above the chart. Whether you see the third depends on whether the rule's primary metric (the one metric an experiment is actually judged on) tracks revenue:

  • Winning Probability / Statistical Significance / Sequential Decision: engine-aware. Bayesian rules report the leading variation's probability-to-beat-control. Frequentist rules report the p-value (the probability of seeing a result this extreme if the variation actually changed nothing) and a Significant badge once the threshold is crossed. Sequential rules report "Safe to stop" or "Inconclusive" based on always-valid bounds.
  • Observed Lift vs control: the leading variation's improvement over control, with the credible-interval range below.
  • Revenue observed: the total revenue recorded on the non-control variations, shown when the rule has a revenue-tracking primary metric. This is a plain sum of what happened. It is not a projection, and it does not subtract what control would have earned. Hidden when there is no positive revenue to show.

Reading the time-series chart

The conversion-rate chart bins exposures and conversions into 4-hour buckets, fixed across all date ranges. The experiment results page charts its own conversion rate the same way, in the same 4-hour buckets. A flag rule and an experiment read equally smooth trend lines because of this. Buckets with zero exposures render as gaps in the line: the chart never interpolates over absent data.

Date range cap

The custom-range picker caps at 90 days end-to-end. At the 90-day cap with five variations the chart draws ~2,700 points, comfortable for the chart engine but noticeable on slower devices.

Reading the variation table

The Primary Metric card expands to show a per-variation table. Each row reports the variation name, total visitors, total conversions, conversion rate, observed lift over the baseline, and the engine-appropriate significance signal. Use the Baseline dropdown in the filter bar to change which variation other rows are compared to. This is useful when you want to see how variation B performs relative to variation A, instead of control.

Multiple metrics get corrected together

A rule tracking more than one metric gets a multiple-comparisons correction: an adjustment for the extra chances a false positive has to slip through. The default method is Tiered. It corrects the primary metric with Bonferroni, and secondary metrics (metrics tracked for context, which do not decide the result on their own) with Benjamini-Hochberg. This runs automatically behind the p-values and Significant badges you see here. An older method called "Tiered Bonferroni" worked differently: it left the primary metric uncorrected. It is no longer offered on new rules.

Two different confidence bars

The "Significant" badge above, and the winner tick in this table, are judged against your rule's Bayesian win-probability threshold (95% by default). The Statistical Confidence signal further down this page uses your project's confidence level instead: a separately stored setting that also defaults to 95%. The two usually agree. If your project has customized one without the other, the summary card and the guardrail can disagree about the same rule.

Filtering by attributes

Click the Attributes dropdown in the filter bar to slice variation performance by registered attributes: plan, country, deviceType, or any custom attribute your team has registered. The picker is grouped by Built-in (auto-detected from the user context) and Custom (registered explicitly in your project).

Each attribute group lists the distinct values observed in this rule's exposures within the active date range, merged with any suggested values you registered for that attribute. Pick one or more values to slice on. Attributes with no observed values, or whose exposures all fall outside your date range, are hidden from the picker.

  • Built-in attributes are auto-detected from the visitor context, like device or country.
  • Custom attributes are the ones your team registered explicitly for this project.
  • Type a value into Add custom value and press Enter to filter on it, even if it is not in the rendered list yet.

Searching long lists

Groups with more than 20 values render the first 20 by default, with a per-attribute search field above the list. Type to filter through up to the 200 most-frequent values for that attribute. For values past that cap, or values you know but haven't been recorded yet, use the Add custom value input below the list. Type the value and press Enter: it applies as a filter even when it isn't in the rendered list.

Combining filters

Multiple values for the same attribute combine with OR. Selecting plan: pro and plan: enterprise shows visitors on either plan. Different attributes combine with AND. Adding country: US on top of the plan filter narrows the slice to US visitors on those plans.

Bookmarking and sharing

Active filters appear in the URL as repeated ?attribute=key:value params. You can bookmark a sliced view, share it with a teammate, or paste it into a doc. Opening the link restores the same filters on load. The breadcrumb back-link does not carry the filter forward: filters are page-scoped to the current rule.

Empty slice

If a slice has no visitors, the page shows "No visitors match these filters" in place of the variation tables, with a Clear filters button to return to the unfiltered view.

Health guardrails

Below the metric tables, the Flag Rule Health Guardrails section reports three signals. These checks flag a problem with the test's own mechanics, like a broken traffic split or too few visitors, rather than with anything the rule is measuring.

  • Sample Ratio Mismatch (SRM): chi-square test of observed vs configured traffic split. Green at p ≥ 0.01, yellow between 0.001 and 0.01, red below 0.001.
  • Statistical Confidence: engine-aware. Reports whether the leading variation has crossed the engine's decision threshold yet, or whether more data is needed.
  • Traffic Health: visitor-count floor, read from the smallest variation. Green at 1,000 visitors or more; yellow from 100 up to 1,000; red below 100.

See the Health Guardrails reference for the full description of each signal, and what the red, yellow, and green statuses mean in context.

  1. Sample Ratio Mismatch checks whether traffic actually split the way you configured it.
  2. Statistical Confidence reports whether the leading variation has crossed this project's confidence bar yet.
  3. Traffic Health turns red when the smallest variation has under 100 visitors, too few to trust.

Getting alerted when a signal goes red

You don't have to keep this page open. A webhook is an automatic HTTP request A vs B sends to your server when something happens. While a rule is Running, A vs B can send a red signal to Slack, Teams, Jira, or a webhook:

  • flag.srm_failed: the SRM signal is red and the rule has traffic. A red SRM with no visitors yet means "can't be checked yet", not "the split is broken". It never raises an alert on its own.
  • flag.guardrail_breached: the traffic-health signal is red. This one does fire at zero visitors. A live rule that nobody is landing in is exactly the problem worth hearing about.

Both are checked whenever this page recomputes results. Each fires at most once per rule per day, so refreshing the page will not re-notify anyone. Alerts are raised only for the rule's official engine: previewing results under a different engine never notifies anyone. They are also raised only while the rule is Running. Opening the results of a paused or concluded rule recomputes the guardrails you see on screen, but never pages anyone.

Turn them on per destination under Project Settings → Notifications.

Segment Lift on flag rules

Below the secondary metrics, the Segment Lift section answers "for whom did this rule help, and for whom did it hurt?" It breaks results down by every registered flag attribute (plan, country, deviceType, custom attributes) instead of experiment segments. The same panel runs on the experiment results page, using identical statistics. Both use Benjamini-Hochberg correction across the family, a chi-square interaction p-value per attribute, and minimum-sample-size masking. One explanation covers both surfaces. See the Segment Lift reference for how the numbers are computed and how to interpret them.

Segment Lift always reads conversions: a visitor counts as converted when they fired the rule's primary metric event at least once inside the attribution window. Rule metrics are event-based, so every rule metric works here. On the experiment results page the same panel is limited to conversion primary metrics for the same reason, and shows a short note when the primary metric measures something else.

On the flag-rule results page specifically, the panel respects any active attribute filter from the picker. HTE runs within the filtered slice. Exporting the rule's CSV (below) appends a Segment Lift section with per-cell rows whenever the panel is populated.

Exporting results

Click Export CSV in the page header to download a CSV of every variation across every metric. Every metric's rows carry a Kind column marking it Binary or Continuous. The value column is formatted to match: a percentage for a binary conversion rate, or a plain number to two decimals for a continuous metric like time-on-page or average order value. Filenames come from the rule name, with non-alphanumeric characters replaced by dashes (e.g. New-Checkout-V2-results.csv). When the Segment Lift panel has data, the export also appends per-cell rows and Top Movers sections under a Segment Lift heading.

Was this helpful?