How to A/B Test When You Don't Have the Traffic

A couple of months ago I was on a panel at an online conference, sitting there as the CRO guy while a couple thousand people watched and fired questions through in real time.
Most of them were small business owners.
One question came through that I've been chewing on ever since. Something like: can you even do A/B testing when you're really small and don't have the traffic to support big tests?
And I kind of fumbled it.
I gave the answer you're supposed to give. That you won't always reach statistical significance when you're a smaller brand, so you should test bigger things like pricing and offers instead of little tweaks.
Which is true. But it wasn't a good answer.
It was the answer you give when you've got 90 seconds and a few thousand people watching. Once the call ended it kept nagging at me, because I don't actually think small brands should sit testing out. I reckon you should be testing from day zero.
The real question is how you do it without chasing a number down a rabbit hole for three weeks, then finding out the test was underpowered and your "winner" was just noise.
So I went and worked out the answer I wish I'd given on the call. This is it.
Why small stores can't test the way big ones do
The thing nobody tells you when you start testing is that traffic isn't really what powers a test. Conversions are.
Every visitor who lands and leaves without buying is noise. You don't know if your variant nudged them or not. The only people who tell you anything are the ones who converted.
So when you run an A/B test, you're really comparing two piles of conversions. The bigger the piles, the more confident you can be that one is genuinely bigger than the other.
This is where small stores get burned.
Say your variant pulls 30 conversions and your control pulls 10. Feels like a monster win. You're already mentally shipping it and telling your mates you tripled conversions.
But 30 versus 10 on small numbers is almost certainly noise. Move a handful of random buyers from one bucket to the other and the whole result flips.
Deborah O'Malley from GuessTheTest has looked at thousands of these tests, and she says the lift on low traffic "appears enormous when, in reality, it's just the difference between a few random conversions." Kohavi, who ran experimentation at Microsoft and Airbnb, is blunter about it. After the first round of obvious fixes, he says 30% improvements to key metrics are "in the realm of fairy tales and untrustworthy results."
That's the trap. The smaller you are, the bigger and more exciting your fake wins look.
So what should you actually test?
If you're small, the changes you test have to be big enough to punch through that noise.
Statistical significance comes down to effect size versus noise. You can't shrink the noise, that's fixed by your traffic. The only lever you control is how big the change is.
So test things that shift behaviour, not things that polish it.
That means the offer itself. Your pricing, your bundles, your free shipping threshold, gift-with-purchase, subscribe and save, the actual deal the customer is weighing up. Changes that hit every single buyer, not just the handful a new button colour might sway.
These are what I think of as tier one experiments. They tend to move things 15% to 40%, which means you can actually read them in three or four weeks even on modest traffic.
Pricing is almost always where I'd start. Test a 10-15% price bump on your top five or ten SKUs. If conversion drops less than the price went up, your net revenue goes up. And if it tanks, you've learned exactly where your ceiling is. Either way you got a real answer. I wrote a whole piece on running profitable pricing experiments if you want the detail on that one.

Below the offer sits how people actually buy. Mini cart drawer versus a full cart page. Skipping the cart and going straight to checkout. Where you drop your cross-sell. These changes move things 8% to 20%, so they take a bit longer to read but they're still very much worth running.
And right at the bottom you've got the presentation layer. Copy, trust badges, social proof placement, image tweaks, layout. This stuff only moves things 2% to 8%. On a small store a change that size needs 60 days or more to show up, if it ever does.
So don't lead with it. Everyone in the CRO world jokes about testing button colours, and this is why. These changes do matter. They just move things so little that at your traffic level you'll never see the result before you die of old age.
How do you squeeze more signal out of the traffic you have?
Two things help a lot here.
The first is to stop measuring purchases and start measuring the step before them.
Purchases are rare. Add to carts and checkout starts happen far more often, sometimes five to fifty times more often. And remember, conversions are what power the test. So if you measure higher up the funnel, your piles of "conversions" are suddenly way bigger and the noise drops off a cliff.
Same test, measured two ways.

Same visitors, same time period, less than half the noise. Add to cart reaches a readable result roughly three times faster than purchase rate does.
The catch is you have to trust the link between the two. As long as you know your add to cart to purchase ratio is fairly stable, a lift in add to carts is a decent proxy for a lift in sales. Keep an eye on the actual purchase rate as a guardrail. If add to carts jump but sales quietly fall, something further down is broken and you don't ship it. There's more on picking the right funnel metric in my post on micro conversions.
The second thing is to stop running one test at a time.
Loads of people believe you can only have one experiment live at once, or you'll "contaminate" the results. This belief absolutely kills your testing velocity, which is the last thing a low traffic store can afford.
It's also mostly wrong. There's a Microsoft paper from 2023 called A/B Interactions: A Call to Relax, where they ran hundreds of tests across millions of users and found the interference between overlapping experiments was less than 0.002%. Basically nothing.

So you can run several tests at once, as long as they're not fighting over the same pixels. Split them across different sections of the store. One up on the hero, one in the mid-page content, one down on the reviews, maybe one in the nav. They just can't touch, because two tests changing the same block are genuinely changing the same variable.
The catch is your testing tool has to actually support it, and on Shopify most of them don't. This one tripped me up for months. We were on Shoplift, which I still rate, but it doesn't do randomised participation. Every extra test you stack on a single page splits that page's traffic further, from a half to a third to a quarter, so your power drops through the floor and you're basically back to one test per page anyway. Intelligems is the only native Shopify tool I've found with a proper randomised participation model, where one visitor can sit in several tests at once and each test still gets clean data. That's the whole reason we moved. I went down the full rabbit hole on this in running parallel A/B tests on Shopify with Intelligems if you want the technical detail.
On a store doing 30-40k sessions a month I'll happily run five to seven tests at once this way.
One hard rule though. Don't test on Shopify checkout. It's already heavily optimised and you'll likely make it worse. Stick to the mini cart, cart page, product pages and navigation.
What if you genuinely can't test at all?
Some of you reading this are below the line where testing makes any sense. That's fine. It's worth knowing where the line is.
Under roughly 1,000 sessions or 20-30 conversions a month, skip A/B testing entirely. You don't have the data, and pretending you do is worse than not testing, because you'll act on noise.
But you're not stuck. You just switch from testing to finding.

Start by fixing what's obviously broken. Walk your store against a heuristic framework like LIFT and be honest about where the value prop is fuzzy, where there's needless anxiety, where you're distracting people right before they buy. Half the stuff you find has an obvious fix that needs no test at all. Hidden shipping costs, no guest checkout, a product page missing the one bit of info everyone wants. Just fix it.
Then get some eyes on real behaviour.

Install Clarity, it's free and it'll auto-flag rage clicks and dead clicks so you're not watching hours of footage blind. Watch 20 or 30 recordings a week focused on your product and checkout pages. Run a survey asking what nearly stopped people buying. None of this needs statistical significance because you're hunting for repeated signals, not a p-value.
The magic is when they line up. If your heatmap shows people ignoring the CTA, and your survey says nobody gets the offer, and your recordings show hesitation on the same page, you don't need a test. Three independent things are pointing at one problem. Just fix it. Nielsen Norman have been banging this drum for years, warning that if you spend all your time fiddling with 1-2% A/B improvements you'll walk straight past the 100% improvements sitting in your qualitative research.
You can also run before and after tests, which is what I usually suggest at the very low end. Measure your baseline for a couple of weeks, change one big thing, measure again. It's directional, not significant, and you're doing a bit of guessing. The big risk is that outside stuff like a sale or a seasonal swing corrupts the comparison. So annotate every change in your analytics, keep the two windows the same length, and whatever you do, don't run one across a promo period.
How do you know your data is actually telling the truth?
This was the real question buried inside that conference question, and it's the bit I fumbled. So here are the guard rails I use.
The number I keep in my head is around 350 conversions per group. That's roughly what you need for a well powered test detecting a 10-15% lift. Below that, treat any result as a hint, not a verdict.
Work out your minimum detectable effect before you start, not after. If your traffic can only reliably spot a 20% change, and you're testing something that'll realistically move things 3%, the test is dead on arrival. Don't run it. There are free calculators for this, CXL's is the one I reach for, and it takes about two minutes.

Watch for a lopsided traffic split too. If you set up a 50/50 test and the traffic came in 58/42, something's broken in the setup, and the whole test is invalid no matter how good the numbers look. Bin it and start again.
And you're allowed to lower the bar. The sacred 95% confidence threshold is a big-traffic luxury. For a cheap change that's easy to reverse, calling it at 85% or 90% is a perfectly reasonable trade. Testing is really just about how much risk you're comfortable with, and a reversible change carries almost none.
That's the honest answer to the conference question. You can absolutely A/B test as a small brand. You just test bigger things, measure higher up the funnel, run more tests at once and stay ruthless about which numbers you're actually allowed to believe.
The stores that win are the ones that keep testing without fooling themselves.
Common questions about low traffic A/B testing
How much traffic do you need to run an A/B test?
Aim for around 350 conversions per variation, or roughly 1,000 conversions a month, for reliable A/B tests. Below about 20-30 conversions a month, skip testing and use session recordings, surveys and heuristic fixes instead. Traffic matters less than conversion count, since conversions are what actually power the result.
What should a small store test first?
Test your offer before anything else. Pricing, bundles, free shipping thresholds and other changes to the actual deal move conversion 15-40%, so they reach a readable result in three to four weeks even on low traffic. Copy and UI tweaks only move things 2-8% and need months of traffic you don't have.
Can you A/B test with fewer than 1,000 visitors a month?
Not reliably. At that level any "winner" is likely random variation between a few conversions. Spend the effort on qualitative research instead: install Microsoft Clarity, watch session recordings, run on-site surveys and fix the obvious usability problems. When three separate signals point at the same issue, fix it without testing.
How long should you run a low traffic test?
Run it at least two to three weeks, ideally a full business cycle of four to five weeks. Pricing and offer tests especially need time for buying patterns to settle. Stopping early because the first few days look good is one of the most expensive mistakes in testing, since early numbers swing wildly on small samples.
Are micro conversions like add to cart reliable to test on?
Yes, as long as the ratio between add to cart and purchase stays stable. Add to cart happens far more often than purchase, so it reaches significance about three times faster and cuts your noise. Always keep purchase rate as a guardrail. If add to carts rise but sales fall, don't ship the change.
Sources
- The Winner's Curse in A/B Testing — Deborah O'Malley, GuessTheTest
- The Ultimate Guide to A/B Testing — Ronny Kohavi
- How to Do Conversion Optimization With Very Little Traffic — Peep Laja, CXL
- A/B Split Testing for Low Traffic Sites — VWO
- Putting A/B Testing in Its Place — Nielsen Norman Group
- A/B Testing With a Small Sample Size — Georgi Georgiev, Analytics Toolkit
- A/B Interactions: A Call to Relax — Microsoft Research, 2023


