Launching soon on the Shopify App Store

testfeed

not enough traffic to ab test

Not Enough Traffic to A/B Test? Do This Instead

Not Enough Traffic to A/B Test? Do This Instead

Low traffic is not a testing problem. It is a maths problem, and most of the advice you will find quietly tells you to cheat the maths.

If you have searched “not enough traffic to A/B test”, you have probably already sensed the honest answer and are hoping someone talks you out of it. I am not going to. I ran a store before I ran a market research company, and I have watched founders burn a quarter waiting on a split test that was never going to conclude. What I can do is show you exactly why the test will not finish, what the popular workarounds actually cost, and the two jobs you should be doing instead. One of them fixes the page you already have. The other decides the offer before you have spent a penny putting traffic behind it.

Why low traffic breaks A/B testing

An A/B test is a statistical instrument. It compares two conversion rates and tries to tell you whether the difference is real or just noise. To do that reliably it needs a minimum number of visitors in each variation, and that minimum is set by three things: your baseline conversion rate, the size of the improvement you want to catch, and how sure you want to be.

Ron Kohavi, who ran thousands of experiments at Microsoft, Amazon and LinkedIn, is blunt about this in his Seven Rules of Thumb for Web Site Experimenters: you have to have enough users, the minimum is fixed by your metric’s variance and the size of the change you want to catch, and in practice that means thousands, often far more. The widely used shorthand puts the minimum at roughly 16 times the variance of your metric divided by the square of the effect you want to detect, per variation, for 80% power at 95% confidence. For a conversion rate, the variance is close to your rate times one minus your rate. You do not need to love the algebra. You need to see what it does to real numbers.

The average Shopify store converts at about 1.4% of sessions, per Littledata’s benchmark of 2,800 stores. Call it a healthy 2% for a store that is doing fine. Plug that in, aim to catch a 10% relative improvement, which is a good result in the real world, and the maths asks for around 78,000 visitors per variation. That is roughly 156,000 visitors in total for one test. Even a strong store converting at 5% needs about 31,000 per variation, or 62,000 in total.

Now put a clock on it. A store sending 2,000 visitors a week to the page you are testing would take about 30 weeks to finish the friendly 5% version of that test, and well over a year for the realistic 2% one. Most small stores do not send 2,000 visitors a week to a single tested page. So the test does not fail. It simply never reaches a point where the number it shows you means anything.

That is the whole problem in one sentence: the lower your conversion rate and the smaller the win you are hunting, the more traffic you need, and small stores have neither the high base rate nor the volume. If you want to see it for your own numbers before you take my word for it, drop them into VWO’s test duration calculator and watch the required weeks climb.

The advice that quietly lowers the bar

Search this topic and you will find a stack of tactics for low-traffic A/B testing. Some are legitimate. Several are just permission to fool yourself, dressed up as pragmatism. Here is the honest read on each.

Lower the confidence threshold. The testing tools will let you do it. Optimizely, for instance, tells low-traffic users they can drop the significance setting so a winner is declared sooner. Dropping from 95% to 80% confidence does shorten the wait. It also means that instead of a 1-in-20 chance of shipping a change that does nothing, you accept something closer to 1-in-5. You have not found more signal. You have agreed to be wrong more often. For a tiny reversible tweak, fine. For a pricing or positioning change, no.

Stop the moment it looks good. This one is the most common and the most damaging, because it feels rigorous. You watch the dashboard, you see p drop below 0.05, you call it. The statistician Evan Miller wrote the definitive takedown of this in How Not To Run An A/B Test. The significance figure your tool shows assumes you fixed the sample size in advance. If you instead peek and stop as soon as you see a winner, the real false-positive rate is nowhere near the 5% you think. Miller shows it can climb past 25%. Peek ten times and what you believe is 1% significance is really about 5%. Repeated peeking does not speed up the truth. It manufactures winners that are not there.

Count sessions instead of users. Switching the unit from unique visitors to sessions inflates your sample and shortens the test. It is legitimate only when the change you are testing plausibly affects behaviour within a single visit rather than across the whole relationship. Use it for a checkout tweak, not for a brand-level change, and know you are trading a little cleanliness for speed.

Test fewer variations. Sound advice, as far as it goes. Every extra variant splits your thin traffic further and pushes significance further away. On a low-traffic store you run one challenger against the control, never a spread of five. This does not solve the traffic problem. It just stops you making it worse.

Switch to a Bayesian tool. You will be told a Bayesian testing tool lets you decide sooner on less data. Be careful with that promise. Bayesian methods change how the result is expressed, in probabilities rather than a pass-or-fail p-value, but they do not conjure information that is not in the data. A Bayesian read on a few dozen conversions is still a few dozen conversions, and it is just as capable of handing you a confident answer that is noise. The engine is not the constraint. The traffic is. If you want the fuller version of this argument, it runs through the guide to A/B testing tools.

None of these turn a store without enough traffic into a store with enough traffic. At best they buy a slightly faster read on a small, reversible change. At worst they hand you a confident number that is noise, and you build a quarter of decisions on top of it.

The reframe that actually helps

Here is the shift that gets founders unstuck. A/B testing was only ever good at one narrow job: telling you which of two versions of an existing page performs better, once real visitors are flowing through it. It was never the tool for deciding what to build, what to sell, what to charge or which ad to run. People reach for it there anyway, because it is the testing word everyone knows, and then they get stuck when the traffic is not there.

So separate the work into two jobs, because they need different tools.

The first job is polishing a page you already have live. The second is deciding an offer before you have spent to send traffic at it. Low traffic makes the first job slow and the second job impossible to answer with on-site testing at all, because there is often no page and no spend yet to test. Once you name which job you are actually doing, the right method is obvious.

What to do instead, by situation

Match your situation to the row. The right move is rarely a faster A/B test. It is usually a different kind of evidence.

Your situationThe honest moveWhy it beats a low-traffic A/B test
I want to improve a live pageQualitative research: 5 to 8 user sessions, on-site polls, session replayTells you why people leave, which a split test never does, and needs a handful of people not thousands
I am choosing a big page redesignTest one bold challenger, pre-register the sample size, accept it may take monthsA big change needs a big effect to detect, which cuts the required sample dramatically
I am deciding an offer, price or claimPre-spend purchase-intent testing against an audience like your buyersGets a directional read before you build or spend, when no live test is even possible
I am picking between ad creativesPut them in front of target buyers before you pay to run themKills weak creative cheaply instead of paying media budget to learn the same thing slowly
I genuinely have moderate trafficBigger swings, one variant, micro-conversions, full business cyclesMakes the traffic you do have go as far as the maths allows

If you want to improve a page you already have

You do not need a split test to learn that your product page confuses people. You need to watch people use it. Five to eight moderated sessions, where you ask someone to buy and narrate their thinking, will surface more real problems than a test that runs for six months. This is not a consolation prize. Agencies that generate test ideas from user research report noticeably higher win rates than those working from data or opinion alone, precisely because the ideas are grounded in a real reason someone hesitated.

Layer on the cheap continuous stuff. A one-question on-site poll that asks “what nearly stopped you buying today?” puts the shopper’s own words next to the behaviour you just watched. Session replay shows you where thumbs hover and where they leave. None of it needs statistical mass. The research methods that suit a small team are collected in the voice-of-customer toolkit, and the order to fix things in is laid out in the ecommerce CRO framework for low-traffic stores. Fix what the evidence already proves, rather than trying to re-prove it on traffic you do not have.

If the change is genuinely large, a full redesign or a new offer on the page, you can still run it as a single bold challenger. Bigger changes produce bigger effects, and bigger effects need far less traffic to detect, so a radical redesign is more testable on thin traffic than a button colour ever was. Pre-register the sample size, then leave it alone until you hit it.

If you are deciding an offer, price, claim or ad before you have the traffic

This is the job A/B testing cannot touch, and the one where most expensive mistakes are made. The wrong hero product, the wrong price, the wrong core claim, the wrong ad: these are decided before a single visitor lands, and no on-site test can help you when there is no page and no spend yet.

The move here is to put the decision in front of your target buyer up front, rather than pouring ad budget behind it and reading the wreckage later. For pricing specifically there are proper methods for this, like the Gabor-Granger approach to find a revenue-sensible price before you commit. For the broader question of whether an audience wants the thing at all, you test the audience before you spend instead of using your live conversion rate as a slow, confounded proxy.

This is the one place my own company fits the problem on the page, so I will say it plainly rather than slip it in. TestFeed lets you put a product, pack, price, claim, name, concept or ad in front of your target shoppers and get back a purchase-intent read, the shoppers’ reasons in their own words, and a clear next move, in days rather than weeks. We built it working with challenger brands like Bae Juice and Sol Bevi. It is a pre-spend, directional signal, not a sales forecast or a guarantee, and it does not judge taste, texture or smell, so it will not tell you whether the product is nice to use. What it does well is triage: killing weak offers cheaply, so that later, when you finally have the traffic to A/B test, you are only ever optimising something already worth optimising.

If you genuinely have moderate traffic

Maybe you are not at zero. You have a few thousand visitors a week and want to test properly. Then make every visitor count. Test one bold change rather than a timid tweak, because the maths rewards larger effects. Consider a higher-funnel metric such as add-to-cart, which converts at a higher rate than purchase and so reaches significance faster, as long as you know the relationship between that step and the sale. Run for full business cycles of at least two to four weeks so a good Tuesday and a bad bank holiday both land inside the window. And fix the sample size before you start, then hold your nerve until you reach it. That last discipline, refusing to peek, is worth more than any tool.

The sequence that works

Put those together and the order almost writes itself. Decide the offer first, using pre-spend research, so you are building around something people have already said yes to. Then get traffic to it. Then, once that traffic is real and steady, use A/B testing for what it is good at: polishing a proven offer at the margins. Founders get this backwards. They try to learn what to sell from a live conversion rate they do not have the volume to read, and skip the cheap up-front decision that would have told them in days.

If you take one thing from this: not enough traffic to A/B test is not a signal to lower your standards. It is a signal that you are reaching for the wrong instrument. Match the tool to the job, and low traffic stops being the thing that blocks you.

Frequently asked questions

How much traffic do you need to A/B test?

Enough to reach the sample size the maths demands, which is usually more than people expect. For a store converting at 2%, spotting a 10% improvement needs about 78,000 visitors per variation, roughly 156,000 in total. Even a strong 5% converting store needs about 31,000 per variation. As a rough gate, if a single page is not seeing tens of thousands of visitors during the test window, an A/B test will not settle the question.

Can you A/B test with low traffic?

You can run one. Reaching a trustworthy result is the hard part. With low traffic you either wait many months, or you loosen the statistics to declare a winner sooner, which just means you are wrong more often. For most small stores the honest answer is to test something other than on-page variants until traffic grows.

What can I do instead of A/B testing?

Split the problem. To improve a page you already have, use qualitative research: five to eight moderated user sessions, on-site polls and session replay will tell you why people leave. To decide an offer, price, claim, name or ad before you have spent on it, use pre-spend purchase-intent testing against an audience built to look like your buyers. Both give you a decision without needing thousands of live visitors.

Is it okay to lower the significance level for low-traffic tests?

Only if you are honest about the trade. Dropping from 95% to 80% confidence, or calling a test the moment it looks good, does not manufacture certainty. It raises the chance you ship a change that does nothing or hurts. For a small, reversible tweak that is a defensible bet. For anything that touches price, positioning or a full-funnel change, it is not.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first test with you.

Direct install from the Shopify App Store arrives in a few weeks.