Tests should run for a minimum of two full weeks and ideally four to six weeks to account for weekly traffic patterns, seasonal variations, and business cycle fluctuations. Running tests in complete seven-day increments captures weekday-versus-weekend behavior differences. Stopping early based on preliminary results is the most common cause of false conclusions in A/B testing.
The two-week minimum is an industry baseline, but several factors extend the required duration. If the sales cycle is longer than two weeks, the test needs to run at least one full cycle to capture the complete buying behavior. Multivariate tests with more variants require longer durations (often one to two months) because traffic is split across more combinations, slowing the time to statistical significance. "Peeking" at results before the test reaches its calculated duration is a documented problem: approximately 70% of incomplete tests show apparent significance that disappears when the test runs to completion. This happens because random fluctuations in early data can mimic real effects. Running tests too long also creates problems, as cookie deletion and data pollution accumulate over extended periods. The window of four to six weeks balances the need for complete data against the risk of data degradation, and the test should end when the predetermined sample size is reached, not when results look favorable.