Answers

What is A/B testing?

The job of A/B testing is to replace opinion with evidence. Instead of arguing about which headline, button, or offer will perform better, the traffic itself decides, and the winning variant ships to everyone.

Short answer

A/B testing is the practice of splitting traffic between two or more variants of a page, email, ad, or flow to measure which one converts best. Each visitor sees one version at random, outcomes are tracked, and a statistical test decides whether the difference is real or noise. The winner replaces the control once the result reaches a confidence threshold, usually 95 percent, with enough sample size to trust it.

Key points

What matters most.

The six ideas that make A/B testing actually work, instead of producing confident-looking winners that quietly reverse the moment the test ends.

Definition

Two variants, random traffic, one winner.

A/B testing splits incoming traffic between a control (A) and one or more variants (B, C, D). Each visitor is assigned at random and sees only one version. Conversions are tracked per variant. After enough visitors have been exposed, a statistical test decides whether the gap between variants is real or random.

Why it matters

Decisions by evidence, not opinion.

Without A/B testing, teams ship the variant the loudest person in the room preferred. With A/B testing, the traffic picks the winner. The marketing team stops shipping based on taste and starts shipping based on measurable conversion lift that holds up under scrutiny.

Statistical significance

The 95 percent confidence bar.

A result is statistically significant when the odds that the observed difference is random noise drops below five percent. Most teams use 95 percent confidence as the ship bar. Calling a winner before that threshold is reached is how teams end up with reversals and false lifts in production.

Sample size

Enough visitors to trust the number.

A tiny sample can show a huge lift that is pure noise. Every A/B test needs a minimum sample size calculated from the baseline conversion rate, the minimum detectable effect, and the desired confidence. Running until the dashboard looks green is not a method. Running to the pre-computed sample size is.

One change at a time

A/B tests one element, cleanly.

A clean A/B test changes one variable between control and variant: the headline, the CTA, the hero image, the price point. If multiple elements change at once, the test cannot say which change produced the lift. Multivariate testing covers the multi-element case, but it needs far more traffic.

The system

Traffic, tracking, and context in one place.

A/B testing works when the variant assignment, the conversion event, and the downstream revenue all land on the same record. When the test lives in one tool and the CRM lives in another, the join is where accuracy quietly leaks and winning variants get credited to the wrong channel.

How it works

The six steps of a trustworthy A/B test.

Any A/B test that skips a step produces a result that cannot be trusted. The six steps below are the minimum a test has to clear before the winning variant ships to all traffic. Shortcuts here are why so many lifts fail to replicate when the full audience rolls out.

Hypothesis

A specific prediction, written down.

A good hypothesis names the change, the audience, and the expected effect. Not 'try a new headline' but 'a benefit-led headline will lift signup conversion on paid traffic from 3.2 percent to at least 3.6 percent.' A written hypothesis prevents moving the goalposts after the result is in.

Variants

Control plus one or more challengers.

Build the variant alongside the control, change only the one element being tested, and keep every other element identical. If the variant touches the headline and the button color at once, the test cannot attribute the lift to either one. One change per variant keeps the result clean.

Sample size

Compute before you start.

Use the baseline conversion rate, the minimum detectable lift, and the target confidence to compute the sample size required. Running past that number is fine. Stopping early because the numbers look good is not. Pre-committing to the sample size is the single biggest defense against false positives.

Random assignment

Every visitor, one variant, forever.

Each visitor is assigned to a variant at random and stays in that variant for the entire test. Flipping variants mid-session invalidates the comparison. The assignment should be sticky across visits, usually through a cookie or a server-side user identifier that persists.

Measurement

One primary conversion event.

Each test has one primary metric: the conversion event the hypothesis is predicting. Secondary metrics can be watched for guardrails like revenue per visitor or bounce rate, but the ship decision rests on the primary metric. Chasing whichever secondary metric happens to move invites a false winner.

Significance check

Hit the bar, then decide.

After the pre-computed sample size is reached, run the significance test. If the variant beats the control at 95 percent confidence or higher, ship it. If not, keep the control and start the next test. A null result is still a result. It saved the team from shipping a change that would not have moved anything.

What to test

The six surfaces that reward testing most.

A/B testing works everywhere a variant can be served and a conversion can be measured, but some surfaces reward testing far more than others. The six below are where most B2B marketing teams start, and where the payoff per test tends to be highest.

Landing pages

Headlines, hero shots, and CTAs.

Landing pages are the classic A/B testing surface because traffic is high, the conversion event is well-defined, and the elements are easy to swap. Testing the headline, the hero visual, the primary CTA label, and the social proof block tends to produce the biggest lifts per test on paid traffic.

Email campaigns

Subject lines, preview text, and sends.

Email A/B tests on subject line and preview text are cheap, fast, and move open rates meaningfully. Larger lists can also test the send time, the from name, and the body layout. Many email tools include native A/B send modes that split the list automatically and roll out the winner.

Signup flows

Form length, field order, and friction.

Signup and checkout flows reward testing because small friction removals compound. Fewer fields, smarter defaults, inline validation, and the order of the steps all swing completion rates. The sample size is often lower than ad traffic, so these tests run longer but the lift per visitor is large.

Pricing pages

Framing, order, and anchors.

Pricing pages test the order of plans, the recommended tier highlight, the feature matrix framing, and the primary CTA copy. Price changes themselves are usually tested through staged rollouts rather than classical A/B tests because of fairness and comparison-shopping concerns, but every other element is fair game.

Ads and audiences

Creative, copy, and targeting splits.

Ad platforms have A/B testing built in, with split-tested creative, headlines, and audience segments. The sample sizes come fast on paid traffic, which makes ad A/B tests one of the quickest sources of learning. The results feed landing page and email tests downstream.

Nurture sequences

Cadence, content, and offer timing.

Marketing automation nurture flows benefit from A/B testing on cadence (every two days versus weekly), content order, and the timing of the first sales-ready offer. Because nurtures run on a lagging conversion event, these tests need patience and a sample size calculator that respects the full cycle.

The pitfalls

Why most A/B tests ship false winners.

The classical A/B test is a disciplined statistical procedure. Most teams run it as a vibe check instead, and the result is a steady stream of winners that fail to replicate. The six pitfalls below are responsible for most of the gap between a test that looked great and a change that did nothing in production.

Peeking

Checking the dashboard too often.

Looking at the test every day and stopping the moment the variant pulls ahead inflates the false-positive rate dramatically. The statistical math assumes the test runs to a pre-set sample size. Peeking and stopping early is the single most common reason A/B test winners quietly reverse after launch.

Underpowered

Too few visitors to see the truth.

A test with 400 visitors per variant cannot reliably detect a two percent lift. It will either declare a false winner on noise or call every test inconclusive. Sample size calculators exist for a reason. Running a test without one is running a test blind, no matter how good the dashboard looks.

Novelty

The new-thing bump that fades.

Variants sometimes win because they look different, not because they are better. Returning visitors notice the change, click more, and the lift fades once the novelty wears off. Running tests long enough to cover a full behavior cycle, and watching the trend over time, catches this before the change ships.

Traffic quality drift

The audience changed mid-test.

If an email blast, a sales push, or a seasonal spike hits during the test, one variant may receive a very different audience than the other. The split is still random inside the test, but the mix changed. Pausing tests during big external events, or stratifying by source, keeps the comparison honest.

Multiple comparisons

Testing ten things and finding one.

Running one test produces a five percent false positive risk at 95 percent confidence. Running ten tests at once produces about a forty percent chance that at least one false winner slips through. Adjusting the confidence threshold or using correction methods for simultaneous tests is how disciplined programs stay honest.

No downstream check

The variant that won the click lost the deal.

A variant can win the primary conversion and still lose on the metric that actually matters. More signups, worse paid conversion. More clicks, worse revenue per visitor. Guardrail metrics further down the funnel catch this, which is why linking the test to the CRM and the revenue record is non-negotiable.

See A/B testing with the full journey in one record.

Strkr captures the variant assignment, the conversion event, and the closed deal on the same contact timeline, so A/B test results tie back to revenue instead of a separate dashboard. Pricing is published. The feature pages show exactly what ships today.

People also ask

Related questions.

What is A/B testing in simple terms?

A/B testing is a method for comparing two or more versions of something, like a landing page or email, by showing each version to a random slice of visitors and measuring which one converts best. The traffic itself picks the winner instead of a person guessing. Once the result is statistically significant, usually at 95 percent confidence, the winning variant replaces the original.

What is the difference between A/B testing and multivariate testing?

A/B testing compares one variable at a time (the headline, the CTA, the image), so the lift can be traced back to that single change. Multivariate testing changes several elements at once and uses more sophisticated math to isolate the effect of each. Multivariate testing needs far more traffic, so most teams start with A/B testing and only graduate when sample sizes support it.

What is statistical significance in A/B testing?

Statistical significance is the probability that the observed difference between two variants is not just random noise. Most teams use a 95 percent confidence threshold, which means there is less than a five percent chance that the winning variant beat the control by luck. Calling a winner before the test hits that threshold is a common reason A/B test lifts fail to replicate after launch.

How much traffic do you need for an A/B test?

It depends on three inputs: the baseline conversion rate, the minimum lift you want to detect, and the confidence you want in the result. A page converting at ten percent that needs to detect a one point lift at 95 percent confidence usually needs several thousand visitors per variant. A sample size calculator run before the test is the only honest way to answer this question for your specific case.

How long should an A/B test run?

Long enough to reach the pre-computed sample size and to cover a full behavior cycle, usually at least one full week to capture both weekday and weekend traffic, and ideally two weeks. Tests that end as soon as the dashboard turns green are peeking, which inflates false positives. The ship rule is sample size reached and significance threshold cleared, not the chart looking good.

What is split testing?

Split testing is another name for A/B testing. The two terms are used interchangeably. Some teams reserve 'split test' for the two-variant case and use 'A/B/n test' when more than two variants are being compared against the control, but the underlying method is the same: randomize visitors across versions, measure conversion, and let statistics pick the winner.

When should you not run an A/B test?

When the traffic is too low to reach significance in a reasonable window, when the change is a correction of something clearly broken, when the test would create an unfair pricing or legal experience, or when the primary metric is too lagging to measure within the test period. In those cases, teams rely on best-practice evidence, user research, or staged rollouts instead of a classical A/B test.

What are the most important A/B testing best practices?

Write the hypothesis before you start, compute the sample size before you start, change one element at a time, run tests to full duration without peeking, watch guardrail metrics further down the funnel, and link the test outcome back to revenue in the CRM. Programs that follow those six rules ship fewer winners but the winners that do ship actually hold up in production.

Try it free. Bring your team next week.

No sales call, no migration consultant, no four-month implementation. Enter your card, get 14 days of the full Pro tier, cancel any time before day 14 with zero charge. Spin up a workspace, import your CSV, and have something useful before lunch.