Too little traffic for an A/B test? What to do instead
An A/B test tells you which page performs better, with a known chance of being wrong. That guarantee is paid for in visitors. Here is the arithmetic, worked through, and what to do when you do not have them.
Published 2 October 2026 · 5 min read
An A/B test answers one question well: which of two versions of a page performs better for the average visitor, with a known chance that the answer is wrong. That guarantee is not free. It is paid for in visitors, and most new sites do not have enough of them.
This post works through the arithmetic so you can check it yourself, then covers what to do when the numbers say no.
The arithmetic
The standard calculation for comparing two conversion rates needs four inputs: your current conversion rate, the improved rate you want to be able to detect, how confident you want to be that a detected difference is real (significance), and how likely you want to be to detect the improvement if it really exists (power). The usual choices are 95% confidence and 80% power.
The formula for the number of visitors needed in each variant is:
n = (z_a + z_b)^2 × [p1 × (1 - p1) + p2 × (1 - p2)] / (p2 - p1)^2
p1 = current conversion rate
p2 = the rate you want to be able to detect
z_a = 1.95996 (95% confidence, two-sided)
z_b = 0.84162 (80% power)Take a 2% conversion rate, and an improvement to 2.4%, which is a 20% relative lift. Step by step:
(z_a + z_b)^2 = (1.95996 + 0.84162)^2 = 7.8489
p1 × (1 - p1) = 0.02 × 0.98 = 0.0196
p2 × (1 - p2) = 0.024 × 0.976 = 0.023424
sum = 0.043024
(p2 - p1)^2 = 0.004^2 = 0.000016
n = 7.8489 × 0.043024 / 0.000016 ≈ 21,106 per variantSo about 21,100 visitors per variant, or about 42,200 for a simple two-way test. These have to be visitors who actually reach the page being tested, not total site traffic.
What that means in time, at a few traffic levels:
| Visitors to the page per week | Weeks to reach about 42,200 |
|---|---|
| 500 | about 84 |
| 2,000 | about 21 |
| 10,000 | about 4 |
And how the number moves when the inputs change, using the same formula:
| Current rate | Rate to detect | Visitors per variant |
|---|---|---|
| 2% | 3% (50% relative lift) | about 3,800 |
| 2% | 2.4% (20% relative lift) | about 21,100 |
| 2% | 2.2% (10% relative lift) | about 80,700 |
| 10% | 12% (20% relative lift) | about 3,800 |
Two patterns are worth remembering. Halving the lift you want to detect roughly quadruples the visitors you need. And a higher baseline rate makes everything cheaper, which is why tests on a click-through step finish long before tests on paid sign-ups. You can put your own numbers into the formula above, or into any online sample size calculator for two proportions.
Why running it anyway does not help
It is tempting to run the test regardless and see what happens. Two things go wrong.
An underpowered test mostly comes back with no significant difference, even when one version really is better, because there is not enough data to separate the effect from noise. You spend weeks and learn nothing.
Worse, when an underpowered test does declare a winner, the measured improvement tends to be exaggerated. With little data, only a large chance fluctuation can cross the significance line, so the results that do cross it overstate the true effect. Stopping a test early the moment it looks significant makes this worse still: checking repeatedly and stopping at the first good result raises the chance of a false winner well above the 5% you thought you had chosen.
What to do instead
None of the alternatives below gives you the clean answer of a well-powered test. They give you something else: a good chance of finding large problems, which on a new site are usually the ones that matter.
Talk to users and watch them. Call five people who match your audience. Ask about the last time they had the problem, then share your page and ask them to think aloud as they read it. Problems that block people tend to show up within the first handful of sessions. You will not get a conversion rate, but you will hear exactly where the page loses them.
Run five-second and first-click tests. Show the first screen for five seconds and ask what the product does. Or ask where they would click to do a specific task. You can do this by hand with a few people from your audience, or with one of the services built for it.
Watch session recordings, by source. Pick one traffic source and watch twenty recordings from it. Look for where people stop scrolling, what they click that is not a link, and where they leave. Mask form inputs, and check your consent obligations, since recordings may need consent depending on where your visitors are.
Compare before and after, carefully. Change one thing, then compare a fixed period after the change with an equally long period before it, covering the same days of the week. Keep a dated log of every change you make. Be honest about what this cannot do: the mix of visitors changes week to week, a single post that does well somewhere can shift every number, and a before-and-after cannot separate your change from everything else that happened. Use it to notice large shifts, read it by source, and treat small differences as noise.
If you do test, test something big. A test that compares two different offers or two genuinely different pages can detect a large effect with far fewer visitors than a test of button colours. And if a step earlier in the funnel has a higher rate, such as clicks on the main button, it reaches a result sooner, as the table shows. Just remember that a proxy is not the goal: more clicks do not always mean more customers.
Treat personalization as a different tool. Personalization changes what each visitor sees based on context, such as where they came from. That can be a sensible bet when your visitors arrive with different questions, but it is a bet, made on reasoning rather than evidence. It does not tell you whether it helped, and with low traffic you could not measure that reliably anyway. If you use it, keep doing the qualitative checks above.
When you do have the traffic
Test. A well-powered A/B test is the best evidence a website can get, and nothing in this post replaces it. The comparison of IntentGrid and Optimizely covers when a dedicated testing platform is the right choice, which is whenever your traffic lets tests finish in weeks rather than months.
For transparency: IntentGrid, which we build, belongs to the last category in the list above. It rewrites copy per visitor, it is in early access, and it measures nothing.
More from the blog
Traffic problem or first-screen problem? How to tell
Before you go looking for more visitors, find out what the ones you have are doing. Split your numbers by where people came from, and the problem often turns out to be the first screen rather than the traffic.
Your landing page has no error message
A broken deploy tells you it is broken. Copy that nobody understands looks exactly like copy that works. Here is how to find out which one you have, and the order of work that avoids the problem.