Statistics you can defend

Frequentist results

The p-value, the z-score and the interval, computed the way somebody who wants to argue with you would compute them.

Two-proportion z-testCorrectedSigned to the variant

The frequentist half answers one question: how often would a difference at least this large appear by chance, if the variant changed nothing at all. Lower is harder to explain away, and below 0.05 is the usual bar.

The comparison is a two-proportion z-test between each variant and the control, pooled the standard way. The z-score is signed in the direction of the variant, so a variant that is ahead reads positive and agrees with the lift interval computed beside it. That sounds trivial and is not: a sign convention that disagrees between two columns of one table is how a losing variant gets shipped.

The significant verdict is the p-value after multiple-comparison correction, not before it. With one variant the correction has no effect and the bar is simply 0.05; with several it tightens, and the console states the threshold rather than leaving you to work it out.

A deployed variant is excluded from the comparison. It receives no randomised traffic, so treating it as a test arm would compare it against a control it no longer competes with and inflate the number of comparisons everything else is corrected for.

Better experiments, better conversions

Test on all visitors with the world’s lightest script and make confident decisions powered by real-time reporting.