What Is an Email A/B Test?
TL;DR
An email A/B test splits a campaign audience into two randomly assigned groups that receive versions differing in one element, then compares a chosen outcome between them. In cold outreach the outcome worth measuring is replies, because a version can win on opens and lose the conversation.
How do you run a valid email A/B test?
Change one element, assign contacts randomly rather than by list order, send both versions in the same window, and decide the metric before the campaign starts. List order correlates with how the list was built, so splitting on it compares two segments and calls the difference a result.
Volume is the constraint people skip. Reply counts in cold outreach are small, so the difference between two versions has to be large to be distinguishable from ordinary variation, and a test on a short list mostly measures which group happened to contain the more responsive contacts.
Why do most cold email A/B tests prove nothing?
Because they change several things at once, stop the moment one version pulls ahead, and judge on opens. Open tracking is unreliable on its own: Apple Mail Privacy Protection, introduced with iOS 15 in 2021, prefetches tracking images through a proxy, so opens are recorded for messages nobody read.
Stopping early is the subtler failure. An early lead in a small sample reverses often, and a test that is checked repeatedly and ended when it looks decided will produce a winner almost every time, whether or not one exists.
What is worth testing in cold outreach?
Audience and offer first, because those move results by margins a test can actually detect. Which segment you send to, and what you offer them, change reply behaviour far more than any rewrite of the same message to the same people.
Then the ask, then the subject line. Wording tests are the most common and the least informative, since the effect size is usually smaller than the noise in the volumes a single campaign provides.
Frequently asked questions
Until both versions have completed their follow-up cycle, not until one looks ahead. Replies to outbound arrive over days, and a test read on the first afternoon is reading the fastest responders rather than the audience.