Why We Stopped Trusting the Sample Size
The test was called a winner on day four - the sample size looked big enough. It just wasn't the sample size that mattered.
An A/B test that hits a testing tool's default significance threshold feels like a finished result. Often it isn't - it's a snapshot of a result still in motion, read at the moment it happened to cross a line. This is what it looked like when calling a test early cost more than waiting would have.
A pricing-page variant was live against the control for four days. By day four, the testing dashboard showed the variant ahead by a wide margin, with the tool's built-in significance indicator lit up green. The traffic volume looked healthy - well past the minimum the platform's own guidance suggested. The call was made to end the test early and roll the variant out to everyone.
Two weeks after the rollout, overall conversion was flat against the pre-test baseline - not the lift the four-day result had promised. Nothing about the four-day number had been calculated wrong. It just hadn't finished happening yet.
The test's four days had included one unusually high-traffic day - a promotional email had gone out to the full list on day two, sending a burst of already-warm visitors through both variants at once. That single day's behavior didn't reflect the site's normal traffic mix, but the significance calculation had no way to know that. It just saw enough total visitors and enough of a gap between the two groups, and reported confidence.
Significance answers a narrower question than it sounds like it answers: given the data collected so far, is the gap between variants big enough that it's unlikely to be random noise. It says nothing about whether the traffic collected so far is actually representative of a normal week.
The only way to know if the early result was real was to look past it - comparing the two full weeks after the rollout against the same two weeks' baseline from before the test began, rather than trusting the in-test dashboard's own read of itself.
With the flat post-rollout number in hand, the four test days were reviewed individually instead of as a single pooled total. Day two stood out immediately once traffic was broken out by day rather than averaged across the whole test.
Rather than banning early looks at the dashboard entirely, a minimum run time - covering at least one full week, so every day-of-week pattern appears at least once - was set as a condition before a result gets acted on, promotional email or not.
Email sends, sales, and other planned traffic spikes now get logged against the test calendar as they're scheduled, so a future review doesn't have to reconstruct after the fact which day was unusual and why.
No new tooling - the testing platform's significance calculation was never wrong about what it was measuring. The fix was a process change: a minimum run time, and a habit of checking a day-by-day traffic breakdown before trusting a pooled total. Reading a test result for what it actually says, not what the dashboard's green checkmark implies, is part of the Analytics & Data practice on every test that gets called, not just the ones that later look wrong.
Yes - the math was sound given the data it had. The problem wasn't the calculation, it was calling the test before the data it was calculating from was representative of a normal period.
Long enough to include at least one full cycle of whatever pattern the traffic follows - usually a full week at minimum, so weekday and weekend behavior both appear, longer if the business has a monthly or seasonal rhythm.
Not by itself - a bigger sample collected entirely during one atypical event is still an atypical sample. Representativeness and sample size are different questions, and a significance reading only checks one of them.
Review traffic broken out by day while a test is still running, not just the pooled total, and cross-check against anything else scheduled to reach the audience during that window - an email send, a sale, a press mention.
Checking a rollout against its pre-test baseline usually settles it in an afternoon.