Launching soon on the Shopify App Store

testfeed

ab testing tools

The Best A/B Testing Tools (And When You Can't Use Them)

The Best A/B Testing Tools (And When You Can't Use Them)

Every “best A/B testing tools” list has the same blind spot. It ranks the tools, scores the features, tallies the prices, and never asks the one question that decides whether any of them will work for you: do you have the traffic to run a valid test? Skip that question and the fanciest platform on the market becomes an expensive way to ship random noise with a confidence score attached.

So this guide does it the other way round. First the maths that tells you whether you are even allowed to play, then the actual A/B testing tools grouped by the job you are trying to do, with current pricing and the trade-off each one asks you to accept. And a straight answer for the majority of stores, which is that split testing is not their tool yet, and here is what to use instead.

The maths that decides whether any of these tools will work

An A/B test, also called split testing, shows two versions of a page to different visitors and measures which one wins. The catch is the word “wins”. A difference between two variations only means something once it clears statistical significance, and significance needs a volume of traffic that most stores do not have.

Here is the arithmetic the tactic lists leave out. Run the standard two-proportion calculation at 95% confidence and 80% power, and detecting a realistic 10% lift on a 2% conversion rate needs roughly 80,000 visitors per variation. An optimistic 20% lift still needs around 20,000 per variation. You can run your own numbers in any sample-size calculator and watch the requirement balloon as the effect you are hunting gets smaller. A store doing 8,000 sessions a month would wait years to power a single test on a modest change.

Two more truths make it harder, not easier. The first is peeking. If you check an ongoing test and stop the moment you see a winner, you are not running one test, you are running several, and each look gives noise another chance to pass for signal. Evan Miller’s How Not To Run an A/B Test shows that stopping at the first sight of p below 0.05 pushes your real false-positive rate towards 26%, not the 5% you think you are getting. The discipline is to fix the sample size in advance and look once.

You will hear that a Bayesian testing tool sidesteps all of this, that it lets you peek freely and get a read at lower traffic. Be careful with that pitch. Bayesian methods change how the result is expressed, in probabilities rather than a pass-or-fail p-value, but they do not conjure information that is not in the data. A Bayesian test on thirty conversions is still thirty conversions of noise. The engine is not the constraint. The traffic is.

The second is that most tests do not win. When Ron Kohavi and Stefan Thomke wrote up years of experimentation at Microsoft for Harvard Business Review, the headline was sobering: only about a third of well-designed experiments moved the target metric, a third did nothing, and a third made things worse. The effects that do land are often fractions of a percent. Read those two facts together and the traffic requirement stops looking like a technicality. You need volume precisely because the real wins are small and easy to miss.

Which is why the honest first step is not choosing a tool. It is finding your row in this table.

Monthly sessionsCan you A/B test?What to buy
Under ~10,000No, not reliablyNo testing tool yet. Qualitative research and pre-launch testing instead.
~10,000 to ~50,000Only large, obvious changes, run patiently over weeksA Shopify app or a mid-market platform, used sparingly on big changes.
Over ~50,000Yes, run a proper experimentation programmeA full-stack platform, or an open-source tool if you have engineers.

If you are in the top row, and most stores are, skip to the section on what to do instead. If you are in the lower two, read on, because now the question of which tool is worth answering.

The best A/B testing tools, by the job you are doing

There is no single best A/B testing tool, only the best fit for what you are testing and the stack you are testing it on. Three jobs cover almost everyone.

If you run a Shopify store and want to test on-store

Shopify ships with no built-in A/B testing, which is the whole reason this app category exists. These three run tests inside your store without you touching a separate experimentation platform.

Shoplift. The pick for theme, page and section testing. It plugs into the Shopify theme editor so you can build variations without leaving the admin, then splits traffic and reports the result. Pricing is visitor-based: a Core plan from 99 a month for up to 50,000 monthly visitors, an Advanced tier from 299, and Pro from 699, with roughly a quarter off annual billing, per Shoplift. Best for: stores redesigning pages and templates. Trade-off: priced on visitors, so the bill rises with the very traffic that makes your tests valid.

ABConvert. Broader test types from a lower entry price. It covers URL redirect, theme and content tests on its Basic plan from 79 a month, and adds price, shipping and checkout testing on the Advanced plan at 149, with a Plus tier at 249, by its own pricing. Best for: merchants who want price and offer tests without jumping straight to a premium plan. Trade-off: the interface is denser than Shoplift’s, and the cheapest tier leaves price testing locked.

Intelligems. The incumbent for price and offer testing on Shopify, and the deepest of the three on that specific job. Plans open around 49 a month for redirect tests and climb, with price and shipping testing gated to a Plus plan at 499 a month, on the Shopify App Store, and the cost scales with your order volume. Best for: brands whose main question is what to charge. Trade-off: the feature most people come for sits on the top tier, so it is the priciest way in if pricing is all you need.

One caveat that applies to all three: they are still A/B testing tools, so the traffic maths above still governs them. A theme test on 5,000 sessions a month is theatre no matter how neatly the app draws the chart. Use these when your volume can actually feed them.

If you want a full-stack experimentation platform

These run experiments across your whole site, and in some cases your product and server-side logic too. They are built for teams doing CRO as a standing programme, not an occasional tweak.

VWO. The mid-market default for web testing, with a visual editor, heatmaps and session recordings in one suite. Pricing is based on monthly tracked users and the modules you switch on, opening around 300 a month on the entry testing tier and moving to custom quotes above that, per VWO. Its long-standing free plan is being wound down to a 30-day trial, so the free on-ramp is closing. Best for: teams that want testing, heatmaps and recordings from one vendor. Trade-off: costs scale with tracked users, and the useful tiers are quote-based.

Convert. The serious testing tool that stays affordable and leans hard on privacy. Its Growth plan runs about 299 a month billed annually for 100,000 tested users, with a Pro tier around 420 adding multivariate and full-stack testing, on Convert’s pricing. It became one of the common landing spots for brands orphaned when Google sunset Google Optimize on 30 September 2023, leaving a large group of stores without a free tool overnight. Best for: brands that want proper experimentation without an enterprise contract. Trade-off: fewer bundled extras than VWO, and still priced on traffic.

Optimizely. The enterprise default, and priced like it. There is no self-serve sign-up and no published price. It is sold on custom annual contracts that procurement data puts from the mid five figures and rising well into six figures with traffic and modules. Best for: large organisations running experimentation across web and product with a dedicated team. Trade-off: the cost and the sales cycle put it out of reach for almost every store reading this.

AB Tasty. Enterprise web experimentation with personalisation bolted alongside. Like Optimizely it is quote-only, with contracts commonly running from the mid five figures a year, and implementation fees on top. Best for: bigger brands that want testing and personalisation from one vendor. Trade-off: enterprise pricing and no free trial, so you cannot try before you commit.

If you have engineers and want free or open-source

If you have developers who can wire in an SDK, you can run a real experimentation programme for little or nothing. The cost moves from a licence to your team’s time.

GrowthBook. Open source under the MIT licence and free to self-host with unlimited experiments, feature flags and traffic, keeping your data in your own warehouse. The hosted cloud is free for up to three users, with a Pro tier at 40 per user a month adding features like sequential testing, per GrowthBook. Best for: data-literate teams that want warehouse-native testing and full control. Trade-off: you own the setup and maintenance.

PostHog. Experimentation as one feature inside a wider product-analytics suite that also gives you session recordings and feature flags. The free tier is generous, covering roughly a million events a month before usage-based pricing kicks in, on PostHog’s pricing. Best for: product teams that want testing living next to their analytics. Trade-off: it is a product-analytics tool first, so it suits apps and SaaS better than a classic storefront.

Statsig. Feature flags free at any scale, with experimentation and analytics on top, and a free Developer tier covering two million events a month. Paid plans open around 150 a month, per Statsig. Best for: teams that want to ship behind flags and experiment as they go. Trade-off: it is built for engineering-led workflows, not marketers clicking through a visual editor.

When you cannot use any of them, which is most of the time

Here is the part the other lists will not say, because they are usually paid to sell you a subscription. If you are under roughly 10,000 sessions a month, none of the tools above will give you a trustworthy result. Buying one does not get you A/B testing. It gets you a dashboard that turns randomness into a green “winner” banner, which is worse than no data because it feels like data.

One nuance worth keeping, though. The full-stack suites like VWO, Convert and PostHog bundle diagnostics, heatmaps, session recordings and on-site polls, that need no statistical significance at all to be useful. If you want one of those for its recordings, buy it for that and simply leave the split-testing switched off until your traffic can feed it. It is the experiment that needs volume, not the heatmap.

So do not buy a testing tool for its testing. Do the things that actually work at your traffic.

Watch real people. Ten session recordings of your checkout, or a five-person usability test where you sit shoppers who resemble your customer in front of your store and watch them stumble, will teach you more than a split test you cannot power. Nielsen Norman Group’s long-standing finding is that testing with five users surfaces around 85% of a site’s usability problems, because the big problems recur fast. You will not get a clean conversion delta from five people, and you are not trying to. At your volume, finding the problem is the job.

Ask, then act on proof. A one-question on-site poll (“what nearly stopped you buying today?”) puts the shopper’s own words next to the behaviour you just watched. Then fix with what the evidence already proves rather than re-proving it on traffic you do not have. The research methods that fit a small team are collected in the voice-of-customer toolkit.

Test the decision before you build it. The most expensive mistakes happen before a single visitor lands: the wrong product, the wrong price, the wrong hero claim, the wrong ad. Do not try to learn those from a live split test, even if you have the traffic, because the signal is slow and tangled with a dozen other variables. Put the thing in front of your target buyer first. The sequencing is covered in how to test your audience before you spend.

This is the one place my own company fits the problem on the page, so I will say it plainly. TestFeed lets you put a product, pack, price, claim, name, concept or ad in front of your target shoppers and get back a purchase-intent read, the shoppers’ reasons in their own words, and a clear next move, in days rather than weeks. We built it working with challenger brands like Bae Juice and Sol Bevi. It is a pre-spend, directional signal, not a sales forecast or a guarantee, and it does not judge taste, texture or smell, so it will not tell you whether the product is nice to use. What it does is help you kill weak options and back strong ones before you have spent the ad budget, so that later, when you finally have the traffic to A/B test, you are optimising something already worth optimising.

Save the A/B testing tools for the day your traffic earns them. Until then, reasoning beats a rigged coin toss.

A/B testing tools compared (2026)

All figures are entry-level and in US dollars, and pricing on these tools moves, so confirm on the vendor’s own page before you commit.

ToolJob it doesPricing modelEntry priceBest forWatch-out
ShopliftShopify theme and page testsVisitor tiers~99/moStore and template redesignsBill rises with traffic
ABConvertShopify price, offer and content testsFlat app tiers~79/moPrice tests on a budgetPrice testing on Advanced tier
IntelligemsShopify price and offer testingOrder-volume tiers~49/moDeciding what to chargePrice testing on Plus at 499
VWOFull-stack web testingTracked-user tiers~300/moTesting plus heatmaps in oneUseful tiers are quote-based
ConvertWeb and full-stack testingTested-user tiers~299/moSerious testing, mid budgetFewer bundled extras
OptimizelyEnterprise web and productCustom annualMid five figures/yrLarge teams at scaleCost and sales cycle
AB TastyEnterprise testing and personalisationCustom annualMid five figures/yrTesting plus personalisationNo free trial, setup fees
GrowthBookWarehouse-native, open sourceFree / per seatFree self-hostedTeams with engineersYou own the setup
PostHogTesting inside product analyticsFree / usageFree to ~1M eventsProduct and SaaS teamsStorefront fit is weaker
StatsigFeature flags and experimentsFree / usageFree dev tierShip-behind-flags teamsBuilt for engineers

How to choose, in three questions

Do you have the traffic? Answer this before anything else. Under about 10,000 monthly sessions, the honest choice is no testing tool at all. Between 10,000 and 50,000, a single Shopify app used sparingly on large changes. Over 50,000, a full programme is worth building.

What are you testing? A page, offer or price on a Shopify store points to Shoplift, ABConvert or Intelligems. A web experiment across your whole funnel points to VWO or Convert, or Optimizely and AB Tasty at enterprise scale. A product or feature, with engineers on hand, points to GrowthBook, PostHog or Statsig.

Build or buy? If you have developers who can maintain an SDK and a warehouse, open source gives you the most control for the least money. If you do not, a managed app or platform is worth the licence, because a testing tool nobody can operate is the most expensive option of all.

Frequently asked questions

What are the best A/B testing tools?

There is no single best tool, only the best tool for the job. For Shopify stores testing pages, offers and prices, Shoplift, ABConvert and Intelligems are the strongest apps. For web experimentation at scale, VWO and Convert lead the mid-market and Optimizely and AB Tasty own the enterprise. For teams with engineers, GrowthBook, PostHog and Statsig give you experimentation on a free or open-source footing. Match the tool to what you are testing and, first, to how much traffic you have.

How much traffic do you need for A/B testing?

More than most stores think. Detecting a realistic 10% lift on a 2% conversion rate at 95% confidence and 80% power needs roughly 80,000 visitors per variation, and an optimistic 20% lift still needs around 20,000. As a working rule, below about 10,000 monthly sessions you cannot A/B test reliably, between 10,000 and 50,000 you can only test large, obvious changes patiently, and above 50,000 you can run a proper programme.

What is the best A/B testing tool for Shopify?

Shopify has no built-in A/B testing, so you need an app. Shoplift is the pick for theme, page and section tests, ABConvert covers price, shipping, checkout and content tests from a lower entry price, and Intelligems is the incumbent for price and offer testing. All three still depend on having enough traffic to reach significance, so check the maths before you subscribe.

Are there free A/B testing tools?

Yes, if you have engineers. GrowthBook is open source under the MIT licence and free to self-host with unlimited experiments. PostHog and Statsig both offer generous free tiers that include experimentation alongside analytics and feature flags. The catch is setup: these are developer tools that need an SDK wired in, not something a solo merchant installs in an afternoon.

What is the difference between A/B testing and split testing?

In everyday use they mean the same thing: showing two versions of a page to different visitors and measuring which performs better. Strictly, split testing sometimes refers to split-URL testing, where each variation lives on a separate URL, while A/B testing more often means changing an element within one page. Both need the same thing to be trustworthy, which is enough traffic to tell a real difference from noise.

The one-line version

Pick the tool for the job, but check the traffic first. If you are on Shopify with the volume to support it, Shoplift, ABConvert or Intelligems will do the on-store work; if you are running a real programme, VWO, Convert or the enterprise pair handle the web; if you have engineers, open source does it for next to nothing. And if you are under 10,000 sessions, the best A/B testing tool is not a tool at all. It is research and pre-launch testing, run until your traffic has earned the right to split-test.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first test with you.

Direct install from the Shopify App Store arrives in a few weeks.