Testing & Validation

What is A/B Testing?

A/B testing, or split testing, randomly assigns users to product variants and compares a target metric. The control group sees the existing experience and the treatment group sees the proposed change. Statistical analysis helps assess the observed difference; a test may find an improvement, a decline, or an inconclusive result.

Formula

Relative Lift = (Treatment Rate - Control Rate) / Control Rate x 100%

Lift = (Treatment Rate - Control Rate) / Control Rate x 100%. Example: Control conversion = 4%, Treatment conversion = 4.8%. Lift = (4.8 - 4) / 4 x 100% = 20% relative lift. Pre-specify the sample size, stopping rule, and statistical method. A p-value assesses how compatible the data are with a specified null model; it is not the probability that the observed difference happened by chance. Report the effect estimate and its confidence interval alongside the p-value.

Benchmarks and interpretation

  • A 5% significance level is a common convention; choose and document the threshold before the test
  • 80% power is a common planning choice: an 80% chance of detecting an improvement of the planned size, under the test assumptions
  • Plan test duration around the required sample and relevant business cycles
  • Include relevant business cycles in the stopping plan; do not stop a fixed-horizon test just because significance appears early
  • Experiment count measures activity; assess improvement through the outcomes of those experiments

When to use A/B Testing

  • Testing a new landing page headline or CTA before committing to a full redesign
  • Validating that a new onboarding flow improves trial-to-paid conversion before full rollout
  • Measuring the impact of a pricing change on signup rate and revenue per visitor
  • Comparing two algorithm variants (recommendation, ranking, feed) on engagement metrics
Common mistakes
  • Stopping a fixed-horizon test early after seeing significance, which can increase false positives
  • Running A/B tests without pre-calculating required sample size, leading to underpowered experiments with inconclusive results
  • Testing too many variables simultaneously in a single experiment, making it impossible to attribute the effect to a specific change
Practical tips
  • Pre-register your hypothesis and primary metric before launching the test to prevent p-hacking and post-hoc rationalisation of results
  • Pre-specify important segment checks; treat unplanned subgroup findings as exploratory because repeated comparisons can produce false positives
  • Establish a minimum detectable effect that is actually business-meaningful (e.g. 10% relative lift) rather than whatever sample size makes the test run quickly

Free A/B Testing Calculator

Enter your inputs in the free A/B Test Calculator and review the result alongside its guidance.

A/B Test Calculator