Launching soon on the Shopify App Store

testfeed

copy testing

Copy Testing: The Measures That Actually Predict Performance

Copy Testing: The Measures That Actually Predict Performance

You can run a beautiful copy test and still back the wrong ad. Not because the method was flawed, but because you measured the thing that was easy to score instead of the thing that predicts a sale. Recall is easy to score. It is also, on the industry’s own evidence, close to useless on its own.

Copy testing is checking whether an ad, headline or claim will work before you spend money running it. Done well, it is the cheapest insurance you can buy on your media budget. Done the way most guides describe it, with recall as the headline metric, it quietly rewards the loud, forgettable ad over the one that would have sold. This piece is about the difference, and about the handful of measures that have been shown to predict performance for decades.

What copy testing actually is

Copy testing is a research method that puts your advertising in front of target buyers before, or instead of, betting the full budget on it. You show them the ad, or two or more versions of it, and measure how they respond. The goal is simple: launch the version most likely to sell, and kill the weak ones before they cost you anything.

The word “copy” is a hangover from print, where the text of an ad was the copy. Today it covers the whole execution: the video, the static, the hook, the headline, the claim, the landing page. If it carries the message, it can be copy tested.

Two terms get tangled up here, so it is worth separating them. Copy testing usually means testing an execution, the actual ad or headline. Message testing usually means testing the idea underneath it, the claim or angle you want to land, before you decide how to word it. In practice you test the message to pick the angle, then copy test the execution to pick the wording. The methods are the same. The order matters, because there is no point polishing the phrasing of an idea that was never going to move anyone.

Copy testing is one part of the wider job covered in our ad testing guide, which takes in audiences, placements and in-flight measurement as well as the creative itself. It also sits right next to concept testing, which does the same job for products and propositions. An ad is just a concept you can put in front of buyers before you commit. The logic carries straight across.

Why creative is the thing worth testing

Before we get to method, it helps to know why this is where your testing effort belongs at all. Marketers spend enormous energy tuning audiences, bids and placements, and comparatively little pressure-testing the creative itself. The evidence says that is backwards.

Nielsen Catalina Solutions analysed nearly 500 CPG campaigns and split the sales lift from advertising three ways: 49 per cent came from the creative, 36 per cent from media buying and planning, and 15 per cent from brand factors such as price and penetration (NCSolutions, Five Keys to Advertising Effectiveness). Creative was the single largest lever, ahead of everything the media plan could do.

Two caveats belong with that, because you will see the number quoted without them. The campaigns ran in 2016 and early 2017 and were all packaged goods, so treat it as a strong prior rather than a current reading of your category. And the gap has been closing: an earlier cut of the same work put creative at 65 per cent and media at just 15, so media planning has genuinely improved while creative stayed roughly where it was. Creative is still the biggest lever. It is no longer the only one that moves.

The ceiling is higher when the creative is actually good. Nielsen’s own framing is that strong creative becomes the overwhelming driver of in-market success, up to 80 per cent for TV and 89 per cent for digital. That is the prize a copy test is trying to find. If the creative is the biggest lever, the creative is what you should be testing hardest.

The cost of getting it wrong is not small. System1, working with Peter Field, found that dull ads that generate no emotional response need up to 2.6 times the media spend to reach the same business effect as interesting ones, and that emotional ads produced roughly six times the share growth of flat, rational ones (System1, The Extraordinary Cost of Dull). That gap is exactly what a good copy test is meant to catch before the money goes out the door.

The measures that predict sales, and the one that doesn’t

Not all copy testing measures are equal, and the industry worked out which ones matter a long time ago.

In the early 1990s the Advertising Research Foundation ran the Copy Research Validity Project, testing 35 different copy testing measures against the real sales results of matched pairs of TV commercials, one known winner and one known loser in each pair. The single clearest finding was that ads people liked generally outsold ads they did not, which made likability the standout predictor (Haley and Baldinger, Journal of Advertising Research). Later reanalysis argued that persuasion, the shift in how likely someone is to buy after seeing the ad, is at least as strong a signal. The debate between the two continues, but the practical takeaway does not change: likability and persuasion are the measures worth trusting.

Notice what is missing from that list. Raw recall, the “do you remember seeing this ad” question that dominates so many copy tests, is a poor standalone predictor of whether the ad sells. An ad can be highly memorable and do nothing for the brand, or worse, be remembered as a rival’s ad. Recall matters only when it is branded recall, tied clearly to you.

So the four measures worth building a copy test around are:

Persuasion and purchase intent. After seeing the ad, is the buyer more likely to choose you? This is the closest proxy to the behaviour you actually want, and it is the one the validity work keeps pointing back to.

Likability. Did the ad land, or fall flat? Not “is it clever”, but did it create a positive response. Given the cost-of-dull evidence, a flat emotional read is a genuine warning sign, not a soft one.

Branding. Do people know it was your ad, and not a competitor’s? An ad that builds the category and forgets to attach your name is a gift to whoever has the bigger shelf. Test whether the brand is remembered, not just the ad.

Comprehension. Did they actually understand the message? Confusion masquerades as indifference. If people cannot say back what you were offering, no amount of persuasion questioning will save it.

This is also why emotional response earns its place. Binet and Field’s analysis of nearly a thousand campaigns in the IPA effectiveness databank found emotional campaigns were almost twice as likely to deliver large profit growth as rational ones, and built more strongly over time (IPA, The Long and the Short of It). A copy test that only checks whether the message was understood, and never whether it was felt, is testing half the ad.

The methods, and when each one earns its keep

There are only a few ways to actually run a copy test. The academic taxonomy sorts them by whether they measure what people think, feel or do (Pechmann and Andrews, copy test methods). In plain merchant terms, these are your options.

Monadic survey pre-test. Each respondent sees one version of the copy and answers on the measures above. Because nobody sees the alternatives, you get a clean read on how each version performs on its own, the way a real viewer would meet it. This is the workhorse of pre-testing.

Comparative or A/B survey. Respondents see two or more versions side by side and pick. Fast and cheap, and fine for a quick preference read, but it flatters differences that a solo viewer would never notice, because real buyers never see your ads lined up next to each other.

In-flight A/B testing. You launch the variants against live traffic and measure real behaviour: clicks, cost per acquisition, return on ad spend. This is the only method that measures actual buying rather than stated intent, which is its great strength. Its weakness is volume. A conversion test needs real traffic and real time before the result is anything more than noise, and many smaller stores never reach that in a sensible window.

Qualitative reactions. A handful of in-depth responses or interviews, to catch the confusing line, the off-putting claim, the joke that does not land. Not for picking winners, but excellent for spotting a problem before you spend on a full quantitative test.

Here is the short version of which to use when.

SituationBest methodWhy
Choosing between finished ad variations before launchMonadic pre-testClean, solo read on the measures that predict
Sense-checking a claim or line earlyQualitative reactionsCheap way to catch confusion before you invest
High-traffic store, want proof on real behaviourIn-flight A/B testMeasures actual buying, has the volume to be valid
Low-traffic store, need a read this weekMonadic pre-testA/B testing will not reach significance in time
Deciding the angle before the executionMessage test, then copy testPick the idea first, then the wording

Plenty of ad testing tools will run one or more of these for you. The mistake I see most often is a small store running in-flight A/B tests it can never power, waiting a fortnight, and reading a coin-flip as a verdict. If you do not have the traffic, pre-test instead. You will get a directional answer in days rather than a false one in weeks.

The standard the industry already wrote

Before designing your own test, it helps to know the industry settled most of this in 1982. Twenty-one of the largest US agencies signed up to PACT, the Positioning Advertising Copy Testing principles, and the nine of them have aged well. Four are worth lifting straight into whatever you run.

Use more than one measurement. PACT is explicit that a single number is not adequate, which is the same conclusion the validity work reached from the other direction. If your test produces one score, it is a preference poll, not a copy test.

Test executions at the same level of finish. The more finished a piece of copy is, the more soundly it can be judged, and putting a polished hero video against a rough mock-up measures your production budget rather than your idea. This is the trap most small brands walk into when they test a new concept against an ad they have already paid to make.

Control the exposure context. Where and how someone meets the ad changes their reaction, so the context has to be identical across versions and as close to real life as you can manage.

Define the sample properly. Who you ask matters more than how many, which is the principle underneath recruiting category buyers rather than whoever a panel will sell you.

A copy testing scorecard you can run this week

You do not need a research department to do this properly. You need the right questions, asked in the right order, of the right people. Here is a template you can lift.

Recruit at least 100 people per version who genuinely buy your category, and 150 to 200 if you can get them. That range is not arbitrary: the traditional standard for a quantitative copy test is 125 to 200 respondents per cell. The audience matters more than the sample size either way: 120 real buyers beat 500 random panellists every time.

Show each person one version only, and put it in a realistic context. That phrase is doing real work. Traditional copy tests show the ad inside a reel of eight to ten other commercials, precisely because an ad viewed on its own flatters itself. Attention is free in a test and scarce in life. The digital equivalent is showing it inside a simulated feed with competing content above and below. Test your ad in a white void and a win only tells you it beat nothing.

Then ask, roughly in this order:

  1. What was that ad for? (Comprehension. Ask before you prompt anything, and see if they can say it back.)
  2. Whose ad was it? (Branding. Unprompted first, then prompted from a short list.)
  3. How did it make you feel? (Likability and emotion. Offer a simple positive-to-negative scale, plus one open box.)
  4. After seeing this, how likely are you to buy from this brand next time you need one? (Persuasion and purchase intent, ideally asked before and after exposure so you can see the shift.)
  5. What, if anything, put you off or confused you? (Open text. This is where the real gold is.)
  6. In your own words, what is the one thing you took away? (Open text. If it is not the thing you were selling, the message failed.)

Questions 5 and 6 are the ones people write badly, and a weak open question returns a weak answer, so it is worth borrowing from these open-ended question examples before you field it.

Score each version on the first four, weight persuasion and likability most heavily, and read every open-text answer, because the numbers tell you which version won and the words tell you why. A version that scores well on branding and comprehension but flat on likability is a message people understood and did not care about. That is a rewrite, not a launch.

One honest limit: stated intent is not the same as a sale. Copy testing gives you a directional signal on which creative is most likely to work, not a guaranteed sales figure. Treat it as a strong filter, not a forecast, and confirm the winner in-flight once it is live and you have the traffic.

Where TestFeed fits

This is the job TestFeed is built for. You can put an ad, a claim, a headline or a concept in front of your target buyers before you spend, and get back a purchase-intent read, the shoppers’ reasons in their own words, and a clear verdict on what to launch, in days rather than weeks. It is pre-spend signal, directional by design, so you use it to choose which creative deserves the budget and then confirm the winner once it is running. It will not tell you how something tastes or smells, and it is not a sales forecast. What it does is stop you pouring money into the ad that was never going to land.

The one thing to take away

Copy testing is not really about the method. Any of them will produce a number. It is about measuring the things that predict a sale, persuasion, likability, branding and comprehension, rather than the thing that is easy to score, recall. Pick the version buyers understand, remember as yours, and actually feel something about, and you have done the job. Everything else is tuning.

Frequently asked questions

What is copy testing?

Copy testing is a research method that checks whether an ad, headline, claim or piece of marketing copy will work before you put real budget behind it. You show the copy, or two or more versions of it, to a sample of your target buyers and measure how they respond, so you can launch the strongest version and cut the weak ones before they eat your media spend. It is sometimes called ad pre-testing or message testing.

What is the difference between copy testing and A/B testing?

Copy testing usually happens before you spend. You show the copy to a sample of target buyers and measure their reaction, which gives a fast directional read on what to launch. A/B testing happens in-flight, with live traffic and real money, and measures actual behaviour such as click-through and cost per acquisition. Copy testing is cheap, fast and predictive; A/B testing is slower and needs volume, but it measures real buying. Strong programmes use both, pre-testing to choose what to launch and A/B testing to confirm what scales.

What should you measure in a copy test?

Measure the things that predict buying: persuasion and purchase intent, likability, branding, and comprehension. Raw recall on its own is a weak predictor. The advertising industry’s Copy Research Validity Project found that ads people like generally outsell ads they do not, so likability paired with clear branding is a better signal than recall alone.

How many people do you need for a copy test?

For a quantitative read you generally want at least 100 target buyers per version, and 150 to 200 if you can get it, so the differences you see are unlikely to be noise. The people matter more than the number: a clean read from 120 genuine buyers of your category beats 500 random respondents. For early qualitative sense-checking, a handful of in-depth reactions can be enough to catch a confusing line before you spend on a full test.

Is copy testing the same as message testing?

They overlap heavily and the terms are often used interchangeably. Copy testing tends to refer to a finished or near-finished execution, the actual ad, headline or landing page copy. Message testing tends to refer to the underlying idea or claim, the thing you want to say before you decide how to say it. In practice you often test the message first to choose the angle, then copy test the execution to choose the wording.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first test with you.

Direct install from the Shopify App Store arrives in a few weeks.