Too little traffic for an A/B test? What to do instead

An A/B test tells you which page performs better, with a known chance of being wrong. That guarantee is paid for in visitors. Here is the arithmetic, worked through, and what to do when you do not have them.

Published 2 October 2026 · 5 min read

An A/B test answers one question well: which of two versions of a page performs better for the average visitor, with a known chance that the answer is wrong. That guarantee is not free. It is paid for in visitors, and most new sites do not have enough of them.

This post works through the arithmetic so you can check it yourself, then covers what to do when the numbers say no.

The arithmetic

The standard calculation for comparing two conversion rates needs four inputs: your current conversion rate, the improved rate you want to be able to detect, how confident you want to be that a detected difference is real (significance), and how likely you want to be to detect the improvement if it really exists (power). The usual choices are 95% confidence and 80% power.

The formula for the number of visitors needed in each variant is:

n = (z_a + z_b)^2 × [p1 × (1 - p1) + p2 × (1 - p2)] / (p2 - p1)^2

p1  = current conversion rate
p2  = the rate you want to be able to detect
z_a = 1.95996  (95% confidence, two-sided)
z_b = 0.84162  (80% power)

Take a 2% conversion rate, and an improvement to 2.4%, which is a 20% relative lift. Step by step:

(z_a + z_b)^2           = (1.95996 + 0.84162)^2 = 7.8489
p1 × (1 - p1)           = 0.02 × 0.98   = 0.0196
p2 × (1 - p2)           = 0.024 × 0.976 = 0.023424
sum                     = 0.043024
(p2 - p1)^2             = 0.004^2       = 0.000016

n = 7.8489 × 0.043024 / 0.000016 ≈ 21,106 per variant
Computed for this post. Other standard versions of the calculation, such as a pooled variance or an arcsine transform, land within a few dozen visitors of this figure.

So about 21,100 visitors per variant, or about 42,200 for a simple two-way test. These have to be visitors who actually reach the page being tested, not total site traffic.

What that means in time, at a few traffic levels:

Visitors to the page per weekWeeks to reach about 42,200
500about 84
2,000about 21
10,000about 4

And how the number moves when the inputs change, using the same formula:

Current rateRate to detectVisitors per variant
2%3% (50% relative lift)about 3,800
2%2.4% (20% relative lift)about 21,100
2%2.2% (10% relative lift)about 80,700
10%12% (20% relative lift)about 3,800

Two patterns are worth remembering. Halving the lift you want to detect roughly quadruples the visitors you need. And a higher baseline rate makes everything cheaper, which is why tests on a click-through step finish long before tests on paid sign-ups. You can put your own numbers into the formula above, or into any online sample size calculator for two proportions.

Why running it anyway does not help

It is tempting to run the test regardless and see what happens. Two things go wrong.

An underpowered test mostly comes back with no significant difference, even when one version really is better, because there is not enough data to separate the effect from noise. You spend weeks and learn nothing.

Worse, when an underpowered test does declare a winner, the measured improvement tends to be exaggerated. With little data, only a large chance fluctuation can cross the significance line, so the results that do cross it overstate the true effect. Stopping a test early the moment it looks significant makes this worse still: checking repeatedly and stopping at the first good result raises the chance of a false winner well above the 5% you thought you had chosen.

What to do instead

None of the alternatives below gives you the clean answer of a well-powered test. They give you something else: a good chance of finding large problems, which on a new site are usually the ones that matter.

Talk to users and watch them. Call five people who match your audience. Ask about the last time they had the problem, then share your page and ask them to think aloud as they read it. Problems that block people tend to show up within the first handful of sessions. You will not get a conversion rate, but you will hear exactly where the page loses them.

Run five-second and first-click tests. Show the first screen for five seconds and ask what the product does. Or ask where they would click to do a specific task. You can do this by hand with a few people from your audience, or with one of the services built for it.

Watch session recordings, by source. Pick one traffic source and watch twenty recordings from it. Look for where people stop scrolling, what they click that is not a link, and where they leave. Mask form inputs, and check your consent obligations, since recordings may need consent depending on where your visitors are.

Compare before and after, carefully. Change one thing, then compare a fixed period after the change with an equally long period before it, covering the same days of the week. Keep a dated log of every change you make. Be honest about what this cannot do: the mix of visitors changes week to week, a single post that does well somewhere can shift every number, and a before-and-after cannot separate your change from everything else that happened. Use it to notice large shifts, read it by source, and treat small differences as noise.

If you do test, test something big. A test that compares two different offers or two genuinely different pages can detect a large effect with far fewer visitors than a test of button colours. And if a step earlier in the funnel has a higher rate, such as clicks on the main button, it reaches a result sooner, as the table shows. Just remember that a proxy is not the goal: more clicks do not always mean more customers.

Treat personalization as a different tool. Personalization changes what each visitor sees based on context, such as where they came from. That can be a sensible bet when your visitors arrive with different questions, but it is a bet, made on reasoning rather than evidence. It does not tell you whether it helped, and with low traffic you could not measure that reliably anyway. If you use it, keep doing the qualitative checks above.

When you do have the traffic

Test. A well-powered A/B test is the best evidence a website can get, and nothing in this post replaces it. The comparison of IntentGrid and Optimizely covers when a dedicated testing platform is the right choice, which is whenever your traffic lets tests finish in weeks rather than months.

For transparency: IntentGrid, which we build, belongs to the last category in the list above. It rewrites copy per visitor, it is in early access, and it measures nothing.