Back to Blog

You Can Make Forty Ads. You Can Test Three.

AI made creative production nearly free. Measuring it did not get cheaper. How many ads your traffic can actually settle, and what to do with the rest of them.

4 min read

Your agency sends over forty ad variants on a Tuesday. Different hooks, different openers, three aspect ratios each. Last year that was a month of work and a real invoice. This year it took an afternoon and cost almost nothing. So you load all of them, split the budget evenly, and wait for a winner.

Three weeks later the dashboard has forty rows, most of them with single digit conversions, and the top one is up 60%. You scale it. It goes back to average.

The cost that moved, and the one that did not

Making creative got cheap. Finding out whether it worked did not. Production was never the thing limiting how fast you learned, and now that production is nearly free it is easy to miss that the real limit never moved. You have only so much traffic, and every ad you add has to share it.

Here is the part that surprises people. An ad needs a certain number of visitors before the result means anything, and that number is much larger than it feels. On a site converting at 2%, you need somewhere around nineteen thousand visitors on a single ad before you can tell a genuine 20% improvement from a good week.

Nineteen thousand. On one ad.

What that does to your calendar

Say you send twenty thousand paid visits a month. Testing two ads against each other takes about two months before the result is trustworthy. Testing ten takes closer to ten months, by which point the creative is stale, the offer has changed twice and nobody remembers what you were testing.

That is the whole trap. Forty ads feels like forty times the learning. It is actually the same learning spread so thin that none of it arrives.

There is a way out, and it is not more traffic. Small improvements are expensive to prove and big ones are cheap. Chasing a 20% lift on that same site takes about nineteen thousand visitors per ad. Chasing a 50% lift takes about three thousand. Same confidence, six times less traffic, just because you stopped trying to measure something small.

The winner is often just the luckiest ad

There is a second problem with a big test, and it is worse because it shows up disguised as good news.

Any single ad can beat another by luck. That is unavoidable, and with one challenger against one control it is rare enough to live with. The trouble is that the odds stack every time you add another ad to the test. One of them is going to look great, and with enough of them running, one of them looking great is almost guaranteed whether or not any of them are actually better.

Put ten ads in a test and the chance that at least one of them posts a fake win is better than one in three. That is true even in the case where all ten ads are identical. Nothing is really winning. Something just has to come first.

And the one that came first is the one you scale. That is why the lift so often evaporates the moment you put real budget behind it.

What to do with the other thirty seven

  • Test arguments, not wording. Two ads making genuinely different cases can beat each other by half, which you can afford to find out. Two rewrites of the same headline will differ by a few percent, which you cannot.
  • Run three at a time, not thirty. If you are not sure how many your traffic supports, three is a safe answer for most accounts, and it is always better than ten.
  • Save the volume for after the test. Once a concept wins, that is when forty variations earn their keep. You are improving something you know works instead of searching with a budget too small to search.
  • Decide when you will stop before you start. Pick the end date up front and hold to it. A test you call the moment it starts looking good is not a test, it is a mood.

Where we land on this

Cheap creative is a real gain. It just moved the bottleneck somewhere less visible. The honest version of “we can make forty ads now” is “we can make forty ads and afford to learn from three of them,” and the teams getting the most out of it are the ones choosing the three on purpose.

The other way to get more out of the same traffic is to stop losing most of it on arrival. Most of the people those ads bring in never identify themselves, so they never enter a test, a list or a report at all. That is the part we work on.