How the statistics work

Every number on the experiment screen is derived, on each page load, from four integers: impressions and conversions for control, and the same for variant. Nothing else is stored. This page explains what those derived numbers mean and how to read them.

The model: one Beta distribution per arm

For each arm we keep a Beta distribution that represents our current belief about its true conversion rate. It starts as Beta(1, 1) — a flat line that says "any rate between 0% and 100% is equally plausible." Each visitor updates it:

ts
alpha = 1 + conversions
beta  = 1 + (impressions − conversions)

With little data the distribution is wide — we genuinely don't know the rate yet. As impressions accumulate it narrows around the observed rate. There is no point at which the math "turns on"; the belief just gets sharper. That is why you can watch the dashboard continuously without invalidating anything — there is no penalty for looking early and often.

Thompson sampling

Assignment and inference are the same operation. To assign a visitor, the SDK draws one random sample from each arm's Beta distribution and sends the visitor to whichever arm's sample came out higher.

  • When the two distributions overlap heavily (early on), each arm wins roughly half its draws — traffic stays near 50/50 and both arms keep collecting data.
  • As one arm pulls ahead, it wins more draws, so more traffic flows to it — you spend less of your visitors on the worse option.
  • It self-corrects. If the trailing arm was just unlucky, its distribution stays wide enough to keep winning some draws and clawing traffic back.

There is no fixed split to configure and no exploration rate to tune.

P(variant > control)

The headline number. We draw 10,000 paired samples — one from each arm's distribution — and count the fraction where the variant sample is higher. 0.87 means: in 87% of plausible worlds consistent with the data so far, variant's true rate is higher than control's.

It is a direct probability about your experiment, not a p-value. 0.5 means the data can't yet tell the arms apart. 0.95+ or 0.05− is where the dashboard flags a strong result.

Expected loss

Probability alone doesn't tell you how much is at stake. Expected loss does. From the same 10,000 paired samples:

  • ifChooseVariant — average amount of conversion rate you give up if you ship variant and control was actually better. 0.004 = about 0.4 percentage points.
  • ifChooseControl — the mirror: what you forgo by keeping control if variant was actually better.

Read it next to the probability. 82% with an expected loss of 0.1% is a safer call than 88% with an expected loss of 3% — in the second case, being wrong is expensive. The dashboard uses an expected-loss ceiling of 0.5% as one of the conditions before it recommends concluding.

Credible intervals

Each arm shows a 95% credible interval: "there's a 95% probability this arm's true conversion rate is between X and Y."

At low traffic these are wide, and often they overlap. That is the model being honest — the data simply doesn't pin the rate down yet. A narrow interval you didn't earn with sample size would be the bug. Intervals tighten on their own as impressions grow.

Relative lift

(variant rate − control rate) / control rate, using the distribution means. It's the practical-size question sitting next to the probability question. A 40% lift from 2% to 2.8% is worth shipping; a 1% lift from 5.00% to 5.05% probably isn't, however confident the probability gets. Both are shown so you weigh them together.

The decision label

The engine collapses the above into one of five labels. The thresholds:

LabelCondition
conclude_variantP(variant > control) ≥ 0.95 and expected loss if you choose variant < 0.5%
conclude_controlP(variant > control) ≤ 0.05 and expected loss if you choose control < 0.5%
strong_signalP(variant > control) ≥ 0.80
directionalP(variant > control) ≥ 0.60
no_signaleverything else — the arms look the same so far

When the label reaches conclude_variant or conclude_control, a banner appears on the experiment. The experiment stays live — nothing transitions automatically. You decide when to conclude. The banner can be dismissed and does not reappear for that browser session.

The forward projection

While a result is still inconclusive, the dashboard estimates how much longer you need. It is based on:

  • your traffic rate since launch — total impressions divided by days elapsed;
  • the pooled conversion rate observed so far;
  • a sample-size estimate for reaching 95% confidence, scaled to the effect size the data suggests.

It has guard rails:

  • Nothing is shown in the first 2 days or under 10 total impressions — too little to extrapolate from.
  • If the projection lands beyond 90 days, the dashboard says the experiment may not reach 95% confidence within 90 days rather than printing a number you can't act on.
  • Once the probability crosses 0.95 or 0.05, the projection disappears — you're there.

It is a planning aid built from current velocity, not a promise. Real traffic and conversion rates move.

The full object

This is what the stats endpoint returns for a live experiment (and what gets frozen into the record when you conclude):

json
{
  "probVariantWins": 0.87,
  "expectedLoss": {
    "ifChooseVariant": 0.004,
    "ifChooseControl": 0.021
  },
  "relativeLift": 0.34,
  "credibleIntervals": {
    "control": [0.031, 0.078],
    "variant": [0.048, 0.101]
  },
  "decision": "strong_signal"
}

For the experiment's state machine — when stats freeze, what happens to the data on conclude and archive — see Experiment lifecycle.