Answer

What is A/B testing in sales and marketing?

Done well, A/B testing turns opinions about copy and design into evidence. Done poorly, it produces false winners that quietly erode results for months because the sample was too small, the test ran against too many variables, or the result was called before the data stabilized.

Short answer

A/B testing is a randomized experiment that compares two versions of a message, page, or offer by changing one variable at a time, then measuring which version wins on a chosen metric. Sales and marketing teams run A/B tests on subject lines, calls to action, send times, and landing page copy. The winner is the version whose result is unlikely to come from random chance given the sample size.

Key points

What matters most.

Six ideas that separate a real A/B test from a split of the audience that looks like a test and misleads the team for months.

Definition

Randomized experiment, one variable.

A/B testing splits a population at random into two groups, shows each group a different version of exactly one thing (the variable), and measures a predefined outcome. Random assignment is what makes it an experiment instead of a comparison. Changing one variable at a time is what lets the result be attributed to anything at all.

What gets tested

Subject lines, CTAs, send times, pages.

In email: subject line, preview text, sender name, send time, call to action. On landing pages: headline, hero image, button copy, form length, social proof. In sales outreach: opener, personalization depth, cadence spacing, channel order. The useful tests target the smallest change that could plausibly move the metric.

The math

Statistical significance, honestly.

A result is statistically significant when the difference between variants is unlikely under random chance given the sample size. The common bar is a 95 percent confidence interval, often stated as a p-value below 0.05. The number only means something if the sample size was chosen before the test started.

The control

Variant A is the version you ship.

The control is the current version, the thing already running in production. Variant B is the challenger. The test exists to decide whether the challenger beats the incumbent by enough to justify the swap. Running a test without a control is just comparing two guesses to each other.

The sample

Underpowered tests tell lies.

The sample size needed depends on the baseline rate, the lift you want to detect, and the confidence bar. A test with 400 sends and a 2 percent open-rate lift is not conclusive, it is noise. Calculate the required sample before the test runs, then let the test reach that size.

Why it matters

Compound gains across the funnel.

A 10 percent lift on open rate, a 10 percent lift on click, and a 10 percent lift on conversion compounds to a 33 percent lift on the end-to-end outcome. The teams that out-perform their category are usually the ones running a steady program of small, honest tests, not the ones chasing a single big redesign.

How an A/B test works

The six steps every test runs through.

The vocabulary around experimentation can sound heavier than the practice. In day-to-day marketing and sales, an A/B test is six steps. The teams that run tests reliably are the ones that write each step down before launching, so a test that fails to reach significance still produces a learning record, not an argument.

Pick the metric

One primary outcome, decided upfront.

Choose the single metric the test is designed to move: open rate for subject-line tests, click rate for CTA tests, booked meetings for cadence tests, demo requests for landing pages. Secondary metrics can be watched for context, but the test is called on the primary. Changing the primary metric after seeing the data invalidates the result.

Form a hypothesis

A specific, falsifiable prediction.

A hypothesis names the change, the direction, and the expected size. "A 4-word subject line will lift open rate by 10 percent over the current 9-word subject line." Vague goals like "improve engagement" are not testable. The hypothesis forces the team to specify what success looks like before the data arrives.

Change one thing

Isolate exactly one variable.

Variant A and variant B must be identical except for the variable under test. If the subject line and the send time both change, the result cannot be attributed to either. Multi-variable comparisons require a multivariate test or a factorial design, which need larger samples than two-arm A/B tests.

Size the sample

Calculate before launching, not after.

Use a sample-size calculator with three inputs: the baseline conversion rate, the minimum lift worth caring about, and the confidence level (usually 95 percent). The calculator returns the audience size per variant. If your list is smaller than that, the test will not reach significance and the result will be inconclusive by design.

Randomize and run

Random assignment, fixed window.

Split the audience 50 / 50 at random. Run the test for a predetermined period long enough to collect the sample. Do not peek at interim results and stop early when the challenger is briefly ahead. Early stopping on a running test inflates the false-positive rate well above the stated bar.

Call the result

Ship the winner, document the learning.

When the test hits its sample size, check significance. If the challenger wins by a significant margin, promote it to control. If it loses or is flat, keep the current version and log the hypothesis that failed. Either outcome is useful. A test that produces no clear winner is not a wasted test, it is a bounded answer.

Common tests

What sales and marketing teams actually test.

The headline-grabbing A/B tests are usually landing page redesigns, but the test program that moves pipeline most reliably is a steady run of small tests on the touchpoints that produce the most volume. The examples below are the tests the Strkr team has seen produce consistent, defensible lift in revenue programs.

Subject lines

Length, question marks, personalization.

Open rate is the easiest metric to move with the smallest change. Short vs long. Question vs statement. First name vs no first name. Specific number vs round number. The gotcha is that open rate overstates lift because the Apple Mail Privacy pre-fetches have blurred the signal. Pair open-rate tests with click-rate as a secondary metric.

Calls to action

Button copy and placement.

"Get started" vs "Start free." "Book a demo" vs "See it live." Above the fold vs below. Single CTA vs CTA plus secondary link. Click rate is the primary metric. CTA copy tests often look small but compound because every downstream page conversion is multiplied by the lift.

Send times

When the email actually goes out.

Tuesday 10 AM vs Thursday 2 PM. Local time vs sender time zone. Business hours vs off-hours for specific buyer personas. Open rate and click rate matter here, but reply rate is the more honest signal because an open at a bad time is a tab left unread, not an engaged read.

Landing pages

Headline, hero, form length.

The headline is the biggest lever and the one most teams underinvest in testing. Hero visual (screenshot vs illustration vs none) matters more on category pages than product pages. Form length almost always lifts conversion when shortened, at the cost of lead quality, so test both metrics together.

Cadence spacing

How many touches, how far apart.

Four touches over ten days vs eight touches over thirty days. First touch personalized vs templated. SMS inserted into an email cadence vs email-only. The primary metric is meetings booked per 100 prospects, not response rate. A cadence that gets more replies but produces fewer meetings is noise disguised as a win.

Pricing display

How the price gets shown.

Price above fold vs price below fold vs price behind a form. Annual pricing emphasized vs monthly pricing emphasized. Toggle default. The primary metric is signups or qualified demos, not the pricing page view count. This is one of the few tests where leadership often needs to be in the room when the result is called.

Common mistakes

The ways A/B tests go wrong.

Most "A/B tests" running in the wild are not actually A/B tests. They are splits of the audience that look like tests and produce numbers people quote in meetings. The six mistakes below are the ones that quietly turn a testing program into a source of confident wrong answers. Spotting them in your own program is the single biggest lift available.

Too many variables

When A and B differ on three things.

The challenger has a new headline, a new hero image, and a new CTA. The challenger wins. Which of the three caused the lift? Unknown. The result is not actionable because the next test cannot build on it. Multivariate testing exists for a reason, but it is a bigger lift, not a shortcut around single-variable discipline.

No control

Comparing two new versions to each other.

A test between two new headlines, with no run of the current headline in the window, cannot say whether either beats what is already shipped. Both could be worse than production and you would not know. The control group is the current version running concurrently against the challenger, in the same audience, in the same window.

Peeking

Calling the test when it is briefly ahead.

Checking a running test and stopping it as soon as the challenger pulls ahead inflates the false-positive rate well above the stated 5 percent bar. Early winners revert to the mean constantly. The discipline is to size the sample first, then look at the result only when the sample is reached, not before.

Underpowered

Not enough traffic to detect the lift.

A test on 400 sends looking for a 2 percent open-rate lift at 95 percent confidence is not under-powered because the test design was wrong, it is under-powered because the audience is too small for the lift being tested. The honest options are to run longer, test a bigger change, or accept the ambiguity.

Wrong metric

Optimizing the thing that is easy to see.

Open rate is easy to measure, so a lot of email programs optimize open rate exclusively and watch click, reply, and meeting rates drift down over time. Click is easy to measure, so landing-page tests optimize click and watch conversion drift. The primary metric needs to be the one closest to revenue the test can plausibly move.

Segmentation drift

The audience changed under the test.

A test that spans a product launch, a cold outreach push, or a major campaign is contaminated by the mix change. Variant B looks like it is winning, but the real cause is that half the sample came from a different list. Pause tests during big audience events, or stratify the sample so the mix is constant across variants.

When tests do not fit

A/B testing is not always the right tool.

For low-volume, long-cycle, or high-stakes decisions, A/B testing is often the wrong instrument. The teams that run the best test programs are also the ones that know when to put the testing framework down and use a different method. The cases below are the common ones where a classical A/B test will underperform other research methods.

Low volume

Not enough traffic for significance.

If the weekly audience is 500 and the test needs 4,000 per variant, the test will run for 16 weeks before anyone can read the result, by which point the market has moved. Low-volume teams get more mileage from qualitative research (user interviews, call recordings, win-loss) than from classical A/B testing.

Long cycles

The outcome lands months later.

In enterprise sales, the conversion signal (closed-won) can land six to nine months after the top-of-funnel touch under test. By then dozens of variables have changed and the attribution gets fuzzy. Test leading indicators that correlate with closed-won, not closed-won itself.

High-stakes

The downside is bigger than the lift.

A test on the pricing page that could depress signups by 20 percent on the losing variant costs more than the expected gain from the winning variant. For high-stakes surfaces, use smaller holdout groups, soft launches, or qualitative review with target customers before putting the test in front of the full audience.

Brand and positioning

The change is too strategic for A/B.

Positioning changes, brand voice shifts, and category framing cannot be tested in isolation because they interact with every other asset. Teams that A/B test a new brand voice on a single email get a result that does not generalize. Positioning is validated with customer research and sales win-rate over quarters, not with inbox-level splits.

Unmeasurable outcomes

The thing you want to move is unobservable.

If the goal is "trust" or "perceived quality" and no observable metric correlates reliably with it, an A/B test will measure a proxy and reward optimizing the proxy at the expense of the goal. Use interviews, panels, or surveys. The right tool is the one that measures the right thing.

Novelty effects

The winner only wins because it is new.

Short tests on repeat audiences capture a novelty bump that fades in weeks. The challenger looks like a winner, gets promoted, and quietly reverts. For audiences that see the same creative more than once, run tests long enough to let the novelty curve settle, or use a holdout to measure the sustained lift.

In the CRM

Where the CRM runs the test.

A CRM with marketing automation runs an A/B test as two coordinated jobs: variant assignment at the moment the audience is defined, and outcome tracking from that moment forward on the same contact record. When those two jobs live in the same tool, the result is defensible. When they live in two tools stitched together, the result is a reconciliation argument.

Variant assignment

Random split at campaign send.

The CRM picks a random 50 percent of the audience for variant A and 50 percent for variant B at the moment the campaign is scheduled. The assignment is written to the contact record, so the same contact always sees the same variant across the test window. Random-at-send prevents the drift that happens when the split is done by list order.

Outcome tracking

Every click, reply, and booked meeting.

Opens, clicks, form submissions, demo requests, and meetings land on the contact timeline with the variant tag attached. The reporting layer rolls the outcomes up by variant without a spreadsheet export. The question "did the challenger produce more meetings?" is answered in the dashboard, not in Excel.

Significance checks

Live confidence math on the dashboard.

Modern marketing automation surfaces a running confidence interval alongside the variant results so the team can see when the sample is large enough to call. The dashboard still enforces the "no peeking" discipline by not auto-promoting a winner until the predetermined sample size is reached.

Audience holdouts

A control group that sees nothing.

For high-stakes campaigns, the CRM carves out a holdout group that receives no touch at all. The lift is measured against the holdout, not against the losing variant. This is the only way to know whether the campaign produced incremental revenue versus pulling demand forward from contacts who would have converted anyway.

Promote the winner

The challenger becomes the new control.

When the test is called, the winning variant becomes the default for the next campaign. The losing variant moves to an archive with the test record so the next person asking "what did we try?" has an answer. Promotion is manual on purpose, so a human reads the result before it ships.

Repeatable program

A test cadence, not one-off tests.

The teams that compound gains run one or two tests per week across the funnel. The CRM holds the template, the audience definition, the variant library, and the result record. A testing program is a quarterly cadence of small lifts that stack, not a yearly redesign that gets tested with one big A/B.

Run A/B tests inside the CRM, not across five tools.

Strkr includes marketing automation with native variant assignment, outcome tracking, and significance math on the same contact records sales uses. The all-in price is published and the full feature list is on the site.

People also ask

Related questions.

What is the difference between A/B testing and split testing?

The two terms are used interchangeably in most contexts. A/B testing describes the general method of comparing two variants against each other on a chosen metric. Split testing is the colloquial name, usually applied to email and landing-page experiments inside marketing tools. Both refer to the same core practice: a randomized experiment with one variable and a predefined outcome.

How long should an A/B test run?

Long enough to reach the sample size you calculated upfront, and no longer than the audience and the business will tolerate. The common rule is at least one full business cycle (a week for most B2B programs) to avoid day-of-week bias, plus however much additional time the sample calculator says the test needs. Avoid both of the two failure modes: stopping too early at a false winner and letting a flat test run so long that the audience changes under it.

What sample size do I need for an A/B test?

The sample size depends on three inputs: the baseline conversion rate of the current version, the smallest lift you want to be able to detect, and the confidence level you require (typically 95 percent). A test on a 5 percent baseline looking for a 10 percent relative lift at 95 percent confidence needs roughly 30,000 contacts per variant. Smaller lifts and lower baselines require larger samples. Use a sample-size calculator before launching the test.

What is a p-value in A/B testing?

A p-value is the probability of seeing the observed difference (or a bigger one) between the two variants purely by random chance, assuming no true difference exists. A p-value below 0.05 is the common bar for calling a result statistically significant, meaning the difference would show up less than 5 percent of the time by chance alone. A p-value above 0.05 does not prove the variants are the same, it means the test did not produce enough evidence to call a winner.

Can I A/B test with a small email list?

You can run the mechanics, but the test will usually be underpowered. A list of 2,000 contacts and a 20 percent baseline open rate needs a very large relative lift (around 30 percent) before a two-variant test will reach significance in one send. Small-list teams get more value from testing high-leverage changes (the whole offer, the whole landing page) over multiple sends than from running statistically rigorous tests on single-variable changes.

What is the difference between A/B testing and multivariate testing?

A/B testing compares exactly two versions that differ on exactly one variable. Multivariate testing compares multiple combinations of multiple variables simultaneously (for example, three headlines paired with three button colors, yielding nine combinations). Multivariate tests surface interactions between variables that A/B tests cannot see, at the cost of needing substantially more traffic to reach significance in each combination.

How many variables can I change in an A/B test?

One. The entire point of A/B testing is that the result can be attributed to a specific cause, and that attribution only holds when every other variable is held constant between the two variants. If you need to change multiple things at once, you need a multivariate design, a factorial design, or sequential A/B tests that change one thing at a time in a known order.

What tools do I need to run an A/B test?

For email and landing pages, modern marketing automation and CRM platforms include native A/B testing with variant assignment, outcome tracking, and significance calculation in one tool. For broader web experiments, dedicated experimentation platforms handle traffic splitting across multiple surfaces. For sales outreach experiments, the CRM is usually the right home because the outcome (meetings booked, deals created) lives there anyway.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.