Monitoring

Stop on significance

Stop early when the evidence genuinely justifies it, and never when it only looks like it does.

Two tests must agreeEnd date requiredTrust gated

Stopping the moment a test crosses p below 0.05 is textbook optional stopping. Peeking repeatedly and stopping on the first crossing inflates the false-positive rate far past the nominal five per cent, and an automated check peeks every single minute, which is the worst case of it.

So the rule is conservative. An automated stop requires the sequential test and the corrected p-value to agree, on the same variant, in the same direction. The sequential test is valid under continuous monitoring; the corrected p-value is a reading aid for a human who looks occasionally and is not a valid stopping rule on its own.

This has a known consequence, which the platform enforces rather than hides. The sequential test is sized for a fixed target lift, so an experiment with a real but smaller effect leaves it reading no effect while the p-value eventually reads significant. They never agree, nothing fires, and the experiment runs to its scheduled end. That is why a scheduled end date is mandatory alongside this setting: significance may only ever end an experiment early, never be the only thing that ends it.

Three trust gates sit in front of the whole path. A split that stopped matching, a full day without traffic, or interference from another running experiment each block the stop, because the reading it would be based on cannot be trusted.

Better experiments, better conversions

Test on all visitors with the world’s lightest script and make confident decisions powered by real-time reporting.