Monitoring

Program report

A look back at the testing programme: how often tests win, how fast they launch, whether winners ship, and which ideas tend to work.

Finished tests onlyRead onlyUp to 24 months

The Program report is the retrospective half of monitoring. The health monitor shows what needs a person now; this shows how the programme has been doing. It reads only tests that have finished, so the early numbers of a running test never sit next to a final result, and opening it changes nothing.

Every finished test gets one outcome. Won means at least one variant beat the control with confidence and none did confidently worse; lost is the reverse; mixed is one of each, which somebody needs to look at; inconclusive is neither. Outcomes are decided on conversions with Holm-Bonferroni correction across variants, and confident means significant with the whole interval on one side of zero.

A deployed test counts as a win even when the statistics were inconclusive, because deploying is a person deciding it won. That decision is kept out of the numbers it could inflate: winners shipped is counted on statistical winners only, and the tables by expected mechanism and by PXL importance use the statistical result, so the ideas people like to roll out do not look better for it.

The velocity chart puts launches per month over the outcomes of what concluded. If the line rises while the share of wins stays flat, the programme is running more tests without learning more from them, which a count of launches on its own cannot show.

Better experiments, better conversions

Test on all visitors with the world’s lightest script and make confident decisions powered by real-time reporting.