Back to Blog

The Winner Was a Coin Flip

One creative came back at 3.1% and one at 2.4%, so you scaled the winner. At that volume the gap you acted on was inside the noise.

5 min read

You ran two versions of an ad. After a week, one was converting at 3.1% and the other at 2.4%. That is a 29% difference, which is enormous, so you did the obvious thing. You turned off the loser and put the budget behind the winner.

Here is the uncomfortable part. At the volume most of these tests run at, a gap that size between two creatives is not a finding. It is the amount that two identical ads would differ by, just from the order the traffic happened to arrive in.

The number you needed was bigger than you think

To say with reasonable confidence that a 2.4% creative and a 3.1% creative are genuinely different, and not just noise, you need roughly 8,600 visitors on each version. About 17,000 in total. That works out to somewhere between 200 and 270 conversions per side before the comparison means anything.

Most creative tests get called at a small fraction of that. Forty conversions a side feels like plenty when you are staring at it. It is not enough to distinguish a good ad from an identical ad.

The reason is not complicated. Conversion is a rare event, and rare events are lumpy. When a version has 40 conversions, moving four of them from one side to the other changes the result by a fifth, and four conversions is a slow Thursday. The percentage on your screen is precise to two decimal places and that precision is completely fake.

What you can actually detect at your volume

Flip the question around. Instead of asking how long to run, ask what your traffic is capable of proving. Starting from a 2.4% baseline, here is roughly what it takes to catch a difference of a given size:

A 10% improvement needs around 67,000 visitors per version. A 25% improvement needs around 11,500. A 50% improvement needs about 3,200. A doubling needs under 1,000.

Read that from the bottom up, because that is the useful direction. If you can put 1,000 visitors on each version, you can detect a creative that is twice as good. You cannot detect one that is 15% better, and 15% better is what most creative differences actually are. So a test at that volume has exactly two honest outcomes: this new idea is dramatically better, or we learned nothing. There is no third result, no matter what the percentages say.

That is not an argument against testing. It is an argument for testing bigger swings. At real-world volumes, small variations are unmeasurable, so spending your test budget on button colors and headline tweaks buys you noise. Testing a genuinely different angle, offer, or format is the only kind of test most accounts have the traffic to resolve.

Checking every day manufactures winners

This is the part that does the most damage, and almost nobody knows they are doing it.

You start a test. You look at it every morning. The moment one version pulls ahead by a margin that looks convincing, you call it and move on. Reasonable behavior. It also breaks the math completely.

We simulated it. Two creatives with exactly identical true performance, 2.4% each, no real difference at all, 20,000 simulated tests. Look once at the end, and you wrongly declare a winner about 5% of the time, which is what you would expect and accept. Look every day for a week and stop when it looks significant, and you declare a false winner 17% of the time. Two weeks of daily checking gets you to 22%. A month gets you to 28%.

To be clear about what that means: better than one in five of those tests hands you a confident winner between two ads that are the same. Not similar. The same. And every one of those false winners becomes a lesson about what your audience responds to, which then shapes the next round of creative, which is how an account ends up with a whole theory of its customers built on coin flips.

The fix is dull and it works. Decide the sample size before you start, then do not act until you reach it. Watch it if you like, but looking is not deciding.

What to do when you cannot get the volume

Most accounts genuinely cannot feed 17,000 visitors into a creative test, and pretending otherwise is not useful. Some things that work anyway:

Test upstream, where the numbers are bigger. Click-through rate resolves far faster than conversion rate, because clicks are common and conversions are rare. You will not learn whether a creative converts, but you will learn whether it earns attention, and you will learn it in days instead of never.

Stop testing variants and start testing swings. If your traffic can only prove a 50% difference, only run tests capable of producing one. Different format, different angle, different offer. Two headlines that mean the same thing are not a test, they are a way to spend a month.

Let the losers accumulate. One test that says nothing is worthless. Six months of tests that all pointed weakly the same direction is a real signal, and it costs nothing extra to keep the records. Most accounts throw that away by never writing down what they ran.

Judge on money, not rate. Conversion rate is the noisiest number in the account. Cost per acquisition and total conversions are steadier, and they are what you actually care about.

Where we land on this

The reason to care is that a false winner is not a neutral outcome. It is worse than having run no test at all, because it leaves you more confident and pointed in a random direction. You scale it, you brief the next round against it, and you carry the belief forward into decisions that get progressively more expensive. Nothing about it feels like an error, which is precisely why it survives.

This applies to anything anyone shows you with a percentage attached, including us. If a number gets put in front of you as evidence that something worked, the only question worth asking is how many conversions it rests on and whether the decision was made before or after somebody went looking. A lift with 30 conversions behind it is a story. Ask for the denominator. If nobody can produce it, you already know.

Book a free strategy call