What is A/B Testing?
A/B testing, or split testing, randomly assigns users to product variants and compares a target metric. The control group sees the existing experience and the treatment group sees the proposed change. Statistical analysis helps assess the observed difference; a test may find an improvement, a decline, or an inconclusive result.
Formula
Relative Lift = (Treatment Rate - Control Rate) / Control Rate x 100%Lift = (Treatment Rate - Control Rate) / Control Rate x 100%. Example: Control conversion = 4%, Treatment conversion = 4.8%. Lift = (4.8 - 4) / 4 x 100% = 20% relative lift. Pre-specify the sample size, stopping rule, and statistical method. A p-value assesses how compatible the data are with a specified null model; it is not the probability that the observed difference happened by chance. Report the effect estimate and its confidence interval alongside the p-value.
Benchmarks and interpretation
- A 5% significance level is a common convention; choose and document the threshold before the test
- 80% power is a common planning choice: an 80% chance of detecting an improvement of the planned size, under the test assumptions
- Plan test duration around the required sample and relevant business cycles
- Include relevant business cycles in the stopping plan; do not stop a fixed-horizon test just because significance appears early
- Experiment count measures activity; assess improvement through the outcomes of those experiments
When to use A/B Testing
- Testing a new landing page headline or CTA before committing to a full redesign
- Validating that a new onboarding flow improves trial-to-paid conversion before full rollout
- Measuring the impact of a pricing change on signup rate and revenue per visitor
- Comparing two algorithm variants (recommendation, ranking, feed) on engagement metrics
- Stopping a fixed-horizon test early after seeing significance, which can increase false positives
- Running A/B tests without pre-calculating required sample size, leading to underpowered experiments with inconclusive results
- Testing too many variables simultaneously in a single experiment, making it impossible to attribute the effect to a specific change
- Pre-register your hypothesis and primary metric before launching the test to prevent p-hacking and post-hoc rationalisation of results
- Pre-specify important segment checks; treat unplanned subgroup findings as exploratory because repeated comparisons can produce false positives
- Establish a minimum detectable effect that is actually business-meaningful (e.g. 10% relative lift) rather than whatever sample size makes the test run quickly
Related terms
Free A/B Testing Calculator
Enter your inputs in the free A/B Test Calculator and review the result alongside its guidance.