Conversion Optimization

    Most of Your A/B Test Winners Never Happened

    An underpowered A/B test does not give you a weak answer, it gives you a confident wrong one. The sample-size gate every test passes before launch, the four-week rule that kills most tests on most pages, why bigger changes are the way out, and the discount to apply to every winner.

    Asha Frazier
    7 min read
    Most of Your A/B Test Winners Never Happened

    A test goes live on Monday. By Wednesday the new variant is up 30% and the dashboard says 96% confidence. Somebody screenshots it, the team calls it, the variant ships to everyone, and a month later the conversion rate is where it was before anyone touched the page.

    Nothing broke. The win was never there.

    I have killed tests that hit "statistical significance" in three days because the next eleven days reverted the result. Early readings on small samples swing hard in both directions, and a team that checks daily and stops at the first good number will reliably ship the swings. Do that across a year of testing and you end up with a roadmap full of winners and a conversion rate that has not moved.

    The fix is unglamorous and it happens before the test, not after. It is also the reason I tell a lot of companies to stop A/B testing most of their pages.

    The gate every test passes before it goes live

    Every test starts with a sample-size calculation. It needs four inputs:

    • Baseline conversion rate, from at least 30 days of data on that page
    • Minimum detectable effect, the smallest lift you would actually act on, usually 5 to 15%
    • Statistical power of 80%
    • Significance of 95%

    Put those into any sample-size calculator and it returns the number of visitors each variant needs. Then apply the rule:

    If the page cannot deliver that sample to every variant within four weeks, do not run the test.

    Here is what that looks like on a real-sized page. A landing page converting at 3% gets 15,000 visitors a month. You want to detect a 15% lift, meaning 3% to 3.45%. The calculator says roughly 24,000 visitors per variant, so 48,000 for a two-way test. That is more than three months of traffic. The test fails the gate.

    You have three ways out, and only three. Move the test to a page with more traffic. Raise the minimum detectable effect, which means testing a bigger change. Or drop the test and make the call on judgment, knowing that is what you are doing.

    What you do not do is run it anyway and read the result. An underpowered test does not give you a weak answer. It gives you a confident wrong one, and that is worse than no answer, because nobody goes back and questions a winner.

    I ran into the small-sample version of this with Travello, a travel app, in early 2017. We set up A and B versions of a landing page for each city on their own domain, with an automated rotator and a short link per test, and paired them with push notification tests. The London list for the notification test was 83 users. My note at the time was that 83 was enough to start with and that we would need to re-test in a week or so once there were more users. That is the honest way to treat 83 people: a read on which message deserves the bigger test, never a result you ship to everyone.

    If you want a second opinion on which of your own pages can actually carry a test, that is a good use of a 20-minute conversation.

    Why bigger changes are the way out

    The gate has a useful side effect. It forces you up the test hierarchy, which is where most teams should have been all along.

    Sample size falls fast as the effect you are hunting gets bigger. On that same 3% page, a 30% minimum detectable effect needs about 6,500 visitors per variant instead of 24,000. Thirteen thousand visitors total, under four weeks of that page's traffic. Same page, same traffic, and now the test passes.

    A button color is never going to move a page 30%. A different value proposition can. So the order I test in runs from the things that decide whether anyone cares down to the things that decide how it looks:

    • Value proposition, offer, headline, social proof, the call to action. Start here.
    • Layout, imagery, form length, how the price is displayed. Middle.
    • Button color, fonts, image placement, spacing. Last, and only on pages with traffic to burn.

    Most teams run this list upside down, because the bottom of it is easy to build and nobody has to agree on anything.

    At Therma, now GlacierGrid, the homepage converted at 2.1% behind an abstract hero. Replacing it with a line about preventing a $50K freezer failure, the outcome the buyer actually lost sleep over, lifted demo requests 47%. Run the arithmetic on a 2% page: detecting a 47% lift takes about 4,000 visitors per variant, while detecting 15% takes about 37,000. Big swings are not only where the upside lives. On a B2B site they are often the only tests the traffic can afford.

    How long it runs, and what counts as a win

    Even with the sample in hand, a test runs a minimum of two full business cycles, which is usually 14 days in B2B and 7 in consumer. It runs through at least one full email cycle if email sends traffic to the page. It stays clear of major holidays, conference weeks and product launches. And nobody changes the tracking while it is live, because a mid-test attribution change makes the whole data set uninterpretable.

    Before launch, every test gets a one-page doc: the hypothesis, written as "We believe that [change] will [outcome] for [audience] because [rationale]," the baseline and target metric, the required sample and estimated duration, the variants, who is included and excluded, what counts as a win and a loss, and what you will do in each case.

    That last line is the one that matters. It is the protection against the conversation where the conversion rate did not lift but time on page did, so somebody wants to call it a win anyway. If the test did not move the metric you named in advance, it did not win.

    Discount the winner before you celebrate it

    A 95% significance level means that if there were no real difference, you would see a result this strong about 5% of the time. It does not mean a 95% chance the winner replicates, and it does not mean the lift estimate is accurate.

    Two practical consequences. Treat a test lift as an upper bound, because the lift you keep in production usually lands at 50 to 80% of what the test showed once the novelty wears off. Then re-test confirmed winners against a holdout about 90 days later. A win at 95% that replicates at 95% is robust. A single 95% win is a guess with good formatting.

    The same honesty applies to multi-variant tests. Four challengers against a control, each judged at 95% with no correction, gives you close to a one-in-five chance of crowning a variant that did nothing. Either correct for the comparisons or cut the variants.

    This is also where testing earns its place in the growth system. Conversion gains lower what every paid click costs, and the email and owned traffic you already have is often the cheapest place to run the test in the first place. A team running eight disciplined tests a month, winning 30% of them at a 12% average lift, compounds to something like 40 to 80% a year on the stage it is optimizing. The same eight tests without the discipline produce a noisy single-digit number nobody can attribute to anything.

    What to do tomorrow

    List every test currently live or planned. For each one, pull 30 days of baseline and traffic for the page and run the sample-size calculation at 95% significance and 80% power.

    Any test that cannot reach its sample within four weeks either moves to a bigger page, gets rebuilt as a bigger change from the top of the hierarchy, or comes off the roadmap. Write the one-page doc for whatever survives, including the decision rule, before it goes live.

    Then go back through the last three winners you shipped and check whether the conversion rate in production actually moved. If it did not, you already know what the next test should be.

    A/B testing
    conversion optimization
    CRO
    sample size
    experimentation
    growth systems

    Growth notes, occasionally.

    What I'm seeing in the numbers across engagements — sent when there's something worth saying, not on a schedule. Reply to unsubscribe; no sequence, no spam.

    Want this applied to your business?

    The Pain Ladder Diagnostic reads your site and your market, finds where the funded pain is, and tells you what I'd do next. Free, a few minutes.

    Prefer to start with your own numbers? Run the growth calculator.