Launching soon on the Shopify App Store

testfeed

monadic testing

Monadic Testing vs Sequential Monadic: Which to Use and When

Monadic Testing vs Sequential Monadic: Which to Use and When

Show one person three product concepts in a row and ask them to rate each, and you have already contaminated the result. The second concept gets judged against the first, the third against both, and the order you happened to show them in quietly shifts the scores. Monadic testing is the fix for that problem, and understanding it properly is the difference between a concept test you can act on and one that just feels rigorous.

This guide covers what monadic testing is, how it differs from its cheaper cousin sequential monadic, where straight side-by-side comparison fits, the research on why order matters, a decision table for choosing between them, and a worked example with the actual sample-size maths. By the end you will be able to design one yourself.

What monadic testing actually is

Monadic testing is a research design where each respondent evaluates a single concept in isolation. They see one idea, one product, one ad, one price, one pack design, and then answer rating questions about it. They never see a second option, so there is nothing to compare against and no order to be swayed by. Their score reflects how they feel about that concept on its own terms.

To compare several concepts, you do not show them all to the same people. You split your audience into separate groups, one group per concept, and each group runs its own isolated monadic test. Then you compare the groups. This is sometimes called a monadic cell design, because each concept sits in its own cell of respondents. Three concepts means three cells of people who never overlap.

That isolation is the whole point. Because no respondent ever holds two concepts in their head at once, the scores you collect are clean reads of individual appeal, not relative preferences shaped by whatever they saw first. When the stakes are high and you need to trust the number, this is the design researchers reach for.

The cost is obvious once you see the structure. If you want a solid read on each concept, every cell needs a proper sample. Three concepts at 150 respondents each is 450 respondents, not 150. The design that gives you the cleanest data is also the one that asks for the most people.

Sequential monadic testing: the cheaper cousin

Sequential monadic testing keeps the monadic idea, one concept evaluated at a time, but reuses the same respondents. Each person sees the first concept, rates it in full, moves on to the second, rates that, and so on through the set. It is still monadic in spirit, because at any given moment they are focused on one concept, but they work through several in sequence.

The appeal is money. Testing four concepts with a pure monadic design might need 400 people, 100 per cell. Sequential monadic could get you there with 100 people, each rating all four. Run that arithmetic across a typical three or four concept test and you are cutting the sample by 60 to 75 per cent, which is the sort of saving that decides whether a test happens at all when you are a growing brand rather than a research department with a standing budget. It is the trade SIS International names plainly: the design that needs the bigger sample brings the bigger expense and the longer delay with it.

It also buys a second thing that pure monadic cannot: within-person comparison. Because each respondent has seen the whole set, you can look at how an individual ranked the options, not just how two separate groups scored on average. For some decisions that relative read is genuinely useful, and it mirrors how a shopper actually behaves in an aisle, weighing options against each other rather than in a vacuum.

The trade-off is longer surveys and tired respondents. A pure monadic survey is short, maybe eight to ten minutes. Ask the same person to work through four concepts properly and you are closer to twenty or thirty minutes, and completion rates fall as attention drains. That fatigue is not free. It shows up in your data.

The catch: order and carryover effects

This is the part that decides whether you can trust a sequential monadic result at all.

When people rate several concepts in sequence, the order changes the scores. This is not a hunch. Research on sequential monadic concept tests has found a clear position effect: concepts are rated more highly when shown first than when the same concept is shown later in the sequence (Understanding the Order Effects in Sequential Monadic Product Tests). Simply being first is worth points.

There is a second, subtler effect on top of that. The quality of the concept a person just saw colours how they judge the next one. A strong concept makes the following concept look worse by contrast, and a weak concept makes the next one look better. This carryover lingers and fades as the respondent works through more options, but it is real and it is measurable. Friedman and Schillewaert demonstrated both effects in the Journal of Marketing Theory and Practice, and concluded that these methodological details matter enough to change which concept a study picks as the winner.

Pure monadic testing does not have this problem, because nobody ever sees a second concept. That is the core reason it is treated as the more trustworthy design.

You cannot delete order effects from a sequential monadic test, but you can manage them so they do not corrupt the answer. Randomise the order each respondent sees, so that across the whole sample every concept appears first, last, and in the middle in roughly equal measure. Randomisation does not remove the effect on any single response, but it spreads it evenly, so it stops favouring one concept over another. Keep the number of concepts per person low, three or four rather than eight, so fatigue and carryover have less room to build. And if the decision is close, record each concept’s position and check whether the ranking holds once you account for it. The classic industry treatment of how to balance these designs is still worth reading in full (Quirks: Balancing monadic-cell test designs).

And what about showing them side by side?

There is a third design that belongs in this comparison, and it is often the one people are actually asking about when they land on monadic testing. Comparative testing, sometimes called side-by-side or forced choice, shows one respondent every concept at once and asks them to pick a winner or rank the set. It is the cheapest and fastest of the three, and it produces a beautifully clear result.

The mistake is treating it as a budget monadic test. It is not a worse version of the same thing, it answers a different question. Monadic asks how much people want this concept. Comparative asks which of these wins a straight fight. Those come apart more often than you would think, because forced choice manufactures a winner whether or not one deserves it. Show three mediocre packs side by side and someone still has to pick. You will get a decisive-looking result that tells you which of three bad ideas is least bad, and nothing at all about whether any of them should be made.

So the rule is simple, and it is worth writing on the brief. If “none of these” is a permitted answer to your decision, you need a monadic read, because only an isolated score can tell you a concept is weak in absolute terms. If you are shipping one of these regardless and the only question is which, comparative is fine and you should take the cheap answer. Two hero images where a page needs one, two button colours, two names on a shortlist you have already committed to: pick side by side and move on. A launch you can still cancel: pay for monadic.

Monadic vs sequential monadic: a decision table

The choice is not about which design is better in the abstract. Pure monadic is cleaner. The question is whether the cleaner read is worth the extra sample for the decision in front of you.

QuestionPure monadicSequential monadic
How many concepts does one person see?OneSeveral, in sequence
Sample neededLarger, one full cell per conceptSmaller, one group rates all
CostHigher, multiplies per conceptLower, one sample covers the set
Survey lengthShort, 8 to 10 minLong, 20 to 30 min
Order and carryover biasNone by designPresent, must randomise
Within-person comparisonNoYes
Best forHigh-stakes launch calls, final validationEarly screening, tight budgets, many ideas

Read it as a spectrum, not a rule. The higher the stakes and the closer the concepts, the more you want pure monadic. The earlier the stage and the tighter the budget, the more sequential monadic earns its place, provided you randomise properly.

A worked example

Say you are launching a new snack line and you have three pack designs to choose between. Call them A, B and C.

The pure monadic route: you field three separate cells of 150 respondents, 450 in total. Nobody sees more than one pack. Design A scores 72 on purchase intent, B scores 68, C scores 61. Because the cells never overlapped and never compared, you can take that at face value. A wins, clearly, and B is close enough that you might test a refinement of both. The read is clean but you paid for 450 people.

The sequential monadic route: you field 150 respondents, each rating all three packs in a randomised order, so you spend on 150 rather than 450. Now you get a second layer. On raw average, A still leads, but you also see that 58 per cent of individuals ranked A first when it was theirs to compare, and you can watch how the ordering behaves. Before you trust it, you check the position data. If A only wins because it landed first more often, that is a warning, not a result. If A wins from every starting position, you can believe it, and you got there for a third of the cost.

The maths is the thing to hold on to. In a pure monadic design the sample is per concept, so it multiplies with the number of concepts. In sequential monadic it does not, which is exactly why it is cheaper and exactly why order has to be watched. A common benchmark for a directional read is 100 to 200 completes per concept, rising to 300 or more for high-stakes launch validation or when you need to read results within subgroups.

That benchmark hides the thing that actually sets your sample, though, which is the size of the gap you need to detect. Sample size is not a fixed price per concept. It is a function of how small a difference would still change your decision. Three packs that land 15 points apart will separate cleanly on 100 per cell. Three that land 4 points apart will not separate on 150, or on 250. So ask it in reverse before you field anything: how close do two concepts have to be before you would happily ship either one? If the honest answer is five points, you need a far bigger cell than the rule of thumb implies.

This is where the received wisdom that monadic is always the better design breaks down. An underpowered monadic cell is the worst of both worlds. You pay the full price of the clean design and still cannot tell your top two apart, so you end up guessing with a bigger invoice. If you cannot fund the sample the gap demands, a properly randomised sequential monadic test is the better buy, because the within-person ranking gives you a read on close concepts that an underpowered monadic cell never will. Clean data you cannot afford enough of is not cleaner than good-enough data you can.

How to run a monadic test this afternoon

You do not need an agency to run a competent monadic test. The steps are the same whichever design you pick.

Start by writing the single decision the test has to inform. Not “learn about the packs”, but “which of these three packs should we take to production”. A test built around a decision stays sharp. A test built around curiosity sprawls.

Fix your success metric before you field anything. For most concept work that is purchase intent, usually a top-two-box read on a five-point scale, so you are counting the people who say they would probably or definitely buy. Decide the threshold that means go before you see the data, so you are not moving the goalposts to fit the result you hoped for.

One trap to sidestep while you set that threshold. If the benchmark came from monadic data, whether that is your own history or a published industry norm, it does not carry over to a sequential monadic test. The position effect moves the absolute level of the scores, not just the running order, so a concept that clears a monadic benchmark by a point in a sequential monadic study has not really cleared it. Compare sequential scores with sequential scores and monadic with monadic. If monadic norms are all you have, use the sequential test to rank the field and a monadic one to make the call.

Choose the design using the table above. Pure monadic if the stakes justify the sample. Sequential monadic with randomised order if you are screening or watching the budget. Keep the concept count per respondent low either way.

Write a short, neutral question set: the purchase-intent question, one or two diagnostic questions on what stood out and what put them off, and just enough profiling to confirm you reached the right audience. Resist the urge to bolt on twenty extra questions because the survey is already open. Every extra question costs you completions.

Then field it, and read the winner against the threshold you set, not against the other concepts alone. If two concepts sit inside a few points of each other, treat that as a tie and let a diagnostic question or a follow-up break it, rather than reporting a false precision the sample cannot support.

Where AI panels fit

The reason monadic testing carries a cost problem is that clean reads need real sample, and real sample means either money or time. That is where an AI panel changes the arithmetic for early-stage work. Instead of recruiting hundreds of people per concept, you can put each concept in front of an audience matched to your own shoppers and get a directional read back in days rather than weeks.

This is the pass we built TestFeed for: test a concept, a pack, an ad, a claim or a price before you spend on a full panel, and get a purchase-intent score, the shoppers’ reasons in their own words, and a clear next move. It works best as a first screen that tells you which concepts deserve the expensive validation and which to cut early. Treat the output as pre-spend signal, not a launch forecast, and pair it with a recruited panel when the final call is high stakes. If you want to see how it sits alongside the traditional options, the roundup of concept testing platforms covers the full field, and the guide to audience testing covers the fundamentals first.

Whatever tool you use, the design logic in this guide still applies. Monadic keeps concepts isolated. Sequential monadic trades a little of that cleanliness for a lot of cost, and asks you to manage order in return.

Frequently asked questions

What is monadic testing? Monadic testing is a research design where each respondent evaluates a single concept in isolation, then answers rating questions about it. Nobody sees a second concept, so their scores are not influenced by comparison. To compare several concepts you split the audience into separate groups, one per concept, and compare the groups at the analysis stage.

What is the difference between monadic and sequential monadic testing? In a pure monadic test each person sees one concept only. In a sequential monadic test each person sees several concepts one after another and rates each. Sequential monadic is cheaper because one respondent does the work of several, but the order they see the concepts in affects the scores, so you have to randomise order and account for it.

What is the difference between monadic and comparative testing? Monadic testing shows each respondent one concept in isolation, so you learn how much people want that concept on its own terms. Comparative testing shows one respondent every concept at once and asks them to pick a winner. Comparative is cheaper and gives a clear ranking, but forced choice produces a winner whether or not one deserves it, so it cannot tell you that a concept is weak in absolute terms. If ‘none of these’ is a valid answer to your decision, use monadic.

When should you use sequential monadic testing? Use sequential monadic when budget or audience size is tight and you need to compare several concepts. It is a sensible default for early screening. Randomise the order each respondent sees, keep the number of concepts per person low, and read the results knowing the design carries an order effect.

How many respondents do you need for a monadic test? A common benchmark is 100 to 200 completed responses per concept for a directional read, rising to 300 or more per concept for high-stakes launch decisions or subgroup analysis. The number is per concept, so three concepts at 150 each means fielding 450 respondents in a pure monadic design, not 150.

Does monadic testing remove all bias? No. Monadic design removes the comparison and order bias that comes from seeing several concepts in a row, which is its main advantage. It does not remove sampling bias, question wording bias, or the gap between what people say in a survey and what they do at the shelf. Treat any concept test as directional signal, not a guaranteed sales number.

The decision, in one line

If the call is expensive to get wrong and the concepts are close, pay for pure monadic and trust the clean read, but fund a cell big enough to see the gap you care about. If you are screening early or watching the budget, run sequential monadic, randomise the order, and check that your winner wins from every position before you believe it. And if you are shipping one of them whatever the data says, stop overthinking it and test them side by side.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first test with you.

Direct install from the Shopify App Store arrives in a few weeks.