Program Monitor Report

See how often your tests win, how fast you launch them, whether winners get rolled out, and which kinds of ideas tend to work.

5 min read


The Program report looks back at your testing programme. It shows how often your tests win, how fast you launch them, whether winning variants get rolled out, and which kinds of ideas tend to work.

You'll find it under Monitors → Program.

The other two Monitors screens look at the present. Health shows what needs attention, and Performance shows how running experiments are trending. Program only looks at tests that have finished, so a running test's early numbers never sit next to a final result.

The report only reads your data. Opening it changes nothing.

Choosing what you see

The strip at the top of the page has two controls.

Scope

  • All websites (agencies see All accounts): every test you have access to.
  • Account specific: only the account you have selected.
  • Only mine: only tests you created or started.

Time window

Choose 3, 6, 12 or 24 months. The default is 12. Windows are whole calendar months ending with the current month, so "12 months" always gives twelve monthly bars.

How a test's result is decided

Each finished test gets one of these outcomes:

  • Won: at least one variant beat the control with confidence, and none did confidently worse.
  • Lost: at least one variant did confidently worse than the control, and none confidently beat it.
  • Mixed: one variant confidently won and another confidently lost, so someone needs to look closer.
  • Inconclusive: the test finished without a confident result either way.

Results are based on conversions, with a correction for comparing several variants at once (Holm-Bonferroni). "Confident" means the result is statistically significant and its whole range of likely lifts is on one side of zero.

Deployed tests count as wins. Deploying a variant means someone on your team decided it won, so the report counts it as a win even when the statistics were inconclusive. These tests appear as Won · deployed in the list. They still count if the deployment was later ended.

When a test counts as finished: on the date it was first deployed or the date it ended, whichever came first. Traffic from a rollout after deployment doesn't count toward the test's length or its sessions.

Tests with no statistics at all (for example, a test that never collected data) are left out of the whole report.

The headline numbers

  • Win rate: won tests as a share of tests with a clear result (won + lost + mixed). Inconclusive tests don't count against you. The note underneath shows the raw count and how many wins came from deployment. Below 5 decided tests it says too few to read.
  • Median winning lift: the typical improvement over the control among winning tests. If some wins have no measured lift (for example, deployed inconclusive tests), the note says how many wins it is based on.
  • Launched: tests started in the window, the average per month, and how many are running now.
  • Median test length: typical number of days from start to finish.
  • Idea to launch: typical number of days from creating a test to starting it.
  • Winners shipped: of the tests that won on the statistics, the share that were deployed or marked as implemented. It uses statistical winners only, so deploying tests can't push this number up by itself.

Velocity chart

For each month, the line shows how many tests you launched and the stacked bars show how finished tests turned out (Won, Mixed, Lost, Inconclusive). Hover over a month to see its numbers.

How to read it: if the line goes up but the green share of the bars stays the same, you're running more tests without learning more from them.

Win rate by expected mechanism

This table groups tests by the expected mechanisms chosen on each test's hypothesis:

  • Increased attention / visibility
  • Improved clarity of next step
  • Reduced cognitive load
  • Reduced friction
  • Increased perceived value or trust

A test with several mechanisms counts under each one. Tests without a mechanism appear as Not stated.

Win rate by PXL importance

This table groups tests by their PXL importance score. It shows whether the way you prioritise ideas predicts which ones win.

  • High importance (7–10)
  • Medium importance (4–6)
  • Low importance (0–3)
  • Not PXL-scored

Both tables use the statistical result only. Deployed wins are left out, because counting them would make high-priority ideas look better just because people choose to roll those out.

The columns in both tables:

  • Significant: tests with a clear result.
  • Won and Lost: how many of those went each way.
  • Win rate: won as a share of significant.
  • Win lift: the median lift of the winners.

Rows with fewer than 5 significant results are faded, because there are too few to compare fairly.

Tip: these tables only become useful once your team fills in expected mechanisms and PXL scores on each test's hypothesis.

Concluded tests

This is every test that finished in the window, newest first, 25 per page. Each row shows the outcome, the lift of the winning variant, the test length, the number of sessions and the date it finished. Click a row to open the experiment.

  • Filter by outcome with the All / Won / Mixed / Lost / Inconclusive tabs.
  • Search by test name or website.

Good to know

  • Results can change slightly over time. Statistics for finished tests are recalculated regularly, so an improvement to how Pertento calculates results can move an older test's outcome.
  • A dash (—) means there is no value, for example a win rate with no decided tests. It is not zero.
  • Nothing showing? No test started, ran or finished in the selected window. Try a longer one.