NevTan Engage lets you create automated email, push, SMS, and WhatsApp customer journeys, segment audiences, and deliver personalized campaigns powered by unified customer data. Even a well-built journey needs testing to improve.
You have a subject line, a push notification, or an email body. You want more conversions. Do you test one change at a time, or many at once?
The honest answer for most teams is one at a time — but understanding why, and knowing the specific conditions where multivariate wins, will save you from running a test that mathematically cannot finish.
A/B testing compares two versions of one element. Multivariate testing (MVT) tests several elements and their combinations at once. MVT needs dramatically more traffic — often an order of magnitude more — because sample requirements scale with the number of combinations, not variables. Use A/B unless you specifically need to measure how elements interact and have the volume to detect it.Running sequential A/B tests is almost always the better use of the same traffic.
The Core Difference
A/B testing splits traffic between two versions differing in one element. Version B wins or it doesn't.
Multivariate testing tests combinations. Three elements with two options each produces 2 × 2 × 2 = eight combinations, and traffic splits across all eight.
The thing MVT gives you that A/B cannot is the interaction effect: cases where one element's performance depends on another. A short headline might work better with a prominent button, and worse with a subtle one. A/B testing, by changing one thing at a time, is structurally blind to this.
The thing MVT costs you is sample size, and the cost is steeper than most people expect.
The Traffic Maths — Where Most MVT Plans Fail
This is the section worth reading carefully, because published guidance on it is frequently wrong, including in ways that encourage tests that cannot finish.
Required sample size depends on three things: your baseline conversion rate, the smallest effect you want to detect, and your tolerance for false positives and missed effects. Any blanket rule like "1,000 conversions per variant" is a rough heuristic for detecting moderate effects, not a law — detecting a 2% relative improvement needs vastly more data than detecting a 20% one.
With that caveat, the scaling problem is real. Work through an example:
A/B test | 8-combination MVT | |
|---|---|---|
Variants | 2 | 8 |
Conversions needed (at ~1,000 each) | 2,000 | 8,000 |
Visitors needed at 2% conversion | 100,000 | 400,000 |
Time at 100,000 visitors/month | ~1 month | ~4 months |
Four months, not four weeks. This is where plans fall apart: teams calculate the total sample correctly, then estimate a timeline as though traffic splits don't multiply duration. A four-month test is usually not a test at all — your product, pricing, audience, and season will all change inside the window, contaminating the result.
The useful threshold is conversions per month, not visitors per month. Those differ by your conversion rate, which is typically a factor of 20 to 100. A site with 10,000 monthly visitors at 2% conversion produces 200 conversions — nowhere near enough for MVT, and marginal even for A/B on small effects.
Pro Tip: Calculate required sample and expected duration before you build the test. If the answer exceeds about four weeks, redesign it — reduce combinations, test a higher-traffic page, or target a larger effect. Tests that run longer than a month rarely produce trustworthy results regardless of what the significance calculator says at the end.
The Alternative Most Teams Should Use
Before reaching for MVT, consider sequential A/B testing: test the headline this week, the winner against a new button next week, that winner against a new image the week after.
Sequential A/B | Multivariate | |
|---|---|---|
Traffic per test | Low | Very high |
Time to first insight | Days | Weeks to months |
Finds interaction effects | No | Yes |
Risk of inconclusive result | Low | High |
Compounding gains | Yes — each winner becomes the new baseline | Once, at the end |
Sequential testing gets you learning immediately, compounds each win into the next baseline, and works at traffic levels where MVT is impossible. What it can't do is reveal interactions — and for most teams, interactions are a refinement to chase after the obvious wins are gone, not before.
The honest guidance: use MVT when you have genuinely high volume, have already captured the large single-element wins, and have a specific reason to believe two elements interact. That's a narrow set of conditions, and it's narrower than most articles on this topic imply.
Open Rates Are No Longer a Reliable Test Metric
This affects every email subject line test you run, and it's missing from most testing guides.
Since Apple introduced Mail Privacy Protection, mail clients pre-fetch images for a large share of recipients — registering an "open" whether or not the person looked at the message. Other providers apply similar protections. The effect is that measured open rate is inflated by an amount that varies by your audience's device mix and doesn't stay constant over time.
What this means practically:
A subject line test won on open rate may have won on nothing
Open rate differences between variants can be driven by audience composition rather than copy
The inflation is not uniform, so you cannot simply adjust for it
Test on clicks and conversions instead. Subject lines affect clicks too — a message that doesn't earn the open can't earn the click — so click-through and downstream conversion remain valid measures of subject line performance. They need more traffic to reach significance than opens did, which is a real cost, but a slower honest signal beats a fast meaningless one.
Track click and conversion performance in campaign reports and campaign click reporting. The metrics glossary defines each so your variants are compared on the same basis.
How to Choose
1. Calculate conversions per month, not visitors. This is your real constraint. Multiply traffic by conversion rate before anything else.
2. Define what you're asking. "Which subject line performs better?" is an A/B question. "Does the discount offer work better with urgency framing or social proof framing?" is still A/B. "Does the effect of urgency depend on which image we use?" is the MVT question — and if you can't phrase your question that way, you don't need MVT.
3. Assess statistical literacy honestly. MVT requires understanding factorial design and interaction effects, and the results require more interpretation than "B won." A misread MVT is worse than no test, because it produces confident wrong conclusions that enter your playbook permanently.
4. Check your timeline. If the test can't complete inside roughly four weeks, redesign it.
5. Match the tool to the channel. Web testing tools and messaging platforms are different categories. Testing email subject lines, send times, and push copy happens in your engagement platform; testing landing page layouts happens in a web experimentation tool.
For most small and mid-sized teams, A/B testing is the right answer indefinitely, not just as a starting point. Graduating to MVT is a function of traffic, not maturity — a disciplined team at 50,000 monthly visitors should still be running sequential A/B tests.
What to Test, in Priority Order
Choosing what to test matters more than choosing how, and this is where most testing programmes leak value.
Priority | Element | Why |
|---|---|---|
1 | The offer | The largest single lever, and the one teams skip |
2 | Audience and targeting | The right message to the wrong segment loses regardless |
3 | Subject line / headline | What most recipients actually process |
4 | Send time | Often a larger effect than copy — see send-time optimization |
5 | Body copy and structure | Matters, but after the above |
6 | Channel sequence | Email-then-SMS versus SMS-then-email |
7 | Visual elements | Images, layout |
8 | Button colour | Last, and usually negligible |
Teams routinely spend months on items 7 and 8 while item 1 has never been tested. What to test and when covers the sequencing, and the subject line testing framework covers item 3 specifically.
Note that segmentation often beats optimisation. A message that underperforms overall may win decisively for one segment. Behavioural segmentation frequently produces larger gains than any variant test, because you're changing who receives the message rather than what it says.
How Each Method Works
A/B testing
Split the audience, show each half one version, compare. Two subject lines sent to 5,000 subscribers each; the one with the higher click rate wins — if the difference clears significance.
The critical qualifier is that last clause. A difference that looks decisive can easily be noise at small samples, which is why you decide your sample size before you start rather than watching until the gap looks convincing.
Multivariate testing
Three elements with two options each gives eight combinations. Traffic splits across all eight, and the analysis examines both main effects (does the short headline win on average?) and interactions (does the short headline only win with the prominent button?).
That second question is the entire justification for MVT. If you don't need the answer, you're paying an eightfold sample cost for information you could have obtained in three sequential A/B tests.
On significance
You need a pre-declared threshold — conventionally p < 0.05 in frequentist testing, or a posterior probability threshold in Bayesian approaches. These are different frameworks answering different questions, and they're not interchangeable despite often being presented as equivalent. Pick one and apply it consistently.
Three rules that matter more than which framework you choose:
Declare sample size and duration before starting. Checking repeatedly and stopping when results look good inflates false positives substantially — this is the single most common way invalid results enter a company's playbook.
Run full weeks. Behaviour varies by day. A test running Tuesday to Friday measures the midweek audience, not your audience.
Expect most tests to lose. Industry experience consistently shows only a minority of A/B tests produce a significant improvement, and MVT win rates are lower still. That's not failure — a test that rules out a change has saved you from shipping it. Teams that expect most tests to win stop testing when reality arrives.
Common Mistakes
1. Stopping early. Peeking and stopping on a favourable result is how false positives become company knowledge.
2. Running MVT on insufficient traffic. An eight-combination test on low volume will never conclude. The arithmetic decides this before you write a variant.
3. Confusing visitors with conversions. A 20× to 100× difference, and the source of most impossible test plans.
4. Testing on open rate. Mail privacy protections have made it unreliable. Use clicks and conversions.
5. Ignoring interaction effects in MVT. If you only read main effects, you've paid MVT's cost for A/B's answer.
6. Not segmenting results. A variant that loses overall may win on mobile or for repeat customers. Check before discarding.
7. Testing trivia first. Button colour before offer is the most common misallocation in testing programmes.
8. Not validating winners. Novelty effects fade. Re-test significant winners after a few months, especially anything that depends on being new.
9. Testing while deliverability is broken. Variants in the spam folder measure nothing — deliverability comes before optimisation.
Testing Across Messaging Channels
Testing in journeys differs from testing on web pages in ways worth planning for.
Your sample is your list, not your traffic. You can't wait for more visitors — you have the subscribers you have, which makes sample planning more constrained and makes segmentation more valuable as an alternative to testing.
Each recipient is testable once per send. Unlike a web page where visitors return, an email variant reaches a person once. This makes sequential testing across campaigns the natural structure.
Journeys can test continuously. An automated flow enrols new people constantly, so a variant test inside a welcome or cart recovery journey accumulates sample over time rather than in one burst — generally a better place to test than one-off campaigns. Review results in automation reports.
Channel sequence is a testable variable most teams never test. Whether SMS before email outperforms the reverse is often a larger effect than any copy change, and it's only testable when channels share a journey and a customer profile.
Reusable content blocks make variants cheap. Managing variations in the template editor rather than duplicating whole campaigns is what makes a sustained testing cadence practical.
Frequently Asked Questions
When should I use multivariate instead of A/B?
When you have high conversion volume, have already captured the obvious single-element wins, and have a specific hypothesis about two elements interacting. Absent all three, sequential A/B testing uses the same traffic better.
How much traffic does multivariate testing need?
Roughly the per-variant requirement multiplied by the number of combinations. Eight combinations needing 1,000 conversions each means 8,000 conversions — at a 2% conversion rate, 400,000 visitors. Run the calculation for your own baseline before planning the test.
How long does a multivariate test take?
Divide required sample by monthly traffic. The eight-combination example above, at 100,000 visitors monthly, needs about four months — which is long enough that seasonality and product changes will contaminate it. If your estimate exceeds four weeks, redesign the test.
Can I run multivariate tests on email and SMS?
Some messaging platforms support testing multiple variables in a journey. The constraint is usually list size rather than platform capability: multivariate on a 20,000-contact list won't reach significance on any realistic effect. Sequential A/B testing across campaigns is the practical approach at most list sizes.
What is an interaction effect?
When one variable's impact depends on another — a prominent button helping only alongside a short headline, say. A/B testing cannot detect this because it varies one thing at a time. Detecting interactions is MVT's sole unique capability.
Do I need a data scientist?
For A/B testing, no — platforms automate significance calculation. For MVT, interpreting interactions and segment-level results needs genuine statistical literacy, and a misread result is worse than no result because it gets acted on.
Why can't I test subject lines on open rate anymore?
Mail privacy protections pre-fetch images for many recipients, registering opens that didn't happen. The inflation varies by audience and isn't correctable. Test on clicks and conversions instead — slower to significance, but measuring something real.
What should I test first?
The offer. Then audience, then subject line, then timing. Most programmes invert this and optimise visual details of a message whose fundamental proposition was never questioned.
Test Fewer Things, More Carefully
A/B testing suits almost every team almost all the time. Multivariate earns its place only at high volume with a specific interaction question.
What matters more than the method is testing the right things in the right order, on metrics that still mean something, with sample sizes decided in advance.
NevTan Engage runs journeys across email, push, SMS, and WhatsApp on one unified customer profile — so you can test copy, timing, and channel sequence within the same flow, and compare results on a consistent basis rather than stitching reports across tools.
Start free — no credit card required. See customer case studies for production examples.
