Complete Guide

How to Run an A/B Test

Follow these 6 steps to plan, run, and analyze an A/B test. Covers hypothesis formation, sample size, test design, peeking bias, and post-analysis for product managers.

1
Define Your Hypothesis and Success Metric

Write what you are changing, why you expect it to affect the metric, and how you will measure the result.

›Use the structure: "We believe that [change] will cause [outcome] because [reason], measured by [metric]."
›Example: "We believe that moving the CTA button above the fold will increase free trial sign-up rate because users will see the primary action before scrolling, measured by 7-day free trial conversion rate."
›Choose one primary metric. Secondary metrics are monitored but do not determine the winner.
›Primary metrics for product tests: conversion rate, activation rate, retention (Day 7, Day 30), revenue per user, task completion rate.
›Avoid vanity metrics like page views or time on page unless they directly tie to your business outcome.
›Specify the minimum detectable effect (MDE): the smallest improvement worth deploying. If a 1% lift is not worth the engineering cost, set your MDE at 5%. This also determines your required sample size.

Formula

Hypothesis: If [change], then [metric] will [increase/decrease] by [X]% because [reason]

Pro tip: Use historical data to develop the hypothesis, then record it before examining test results. Label ideas suggested by the results as exploratory and validate them in a new test.

2
Calculate Required Sample Size

Plan sample size before a fixed-sample test. Too little data reduces your chance of detecting an improvement of the size you care about.

›Set your statistical power at 80%. That means an 80% chance of detecting a real effect at your specified MDE.
›Set alpha at 5% for this example. Under the test assumptions, this allows a 5% false-positive rate when there is no real difference.
›You need your current baseline conversion rate (from historical data) and your minimum detectable effect (MDE).
›Sample-size inputs: baseline conversion p, absolute MDE, Z-score for 80% power (about 0.84), and Z-score for 95% two-sided confidence (about 1.96).
›Example: At a 3% baseline and 0.5-percentage-point MDE (about 17% relative lift), PM Toolkit estimates 18,708 users per variant with 95% two-sided confidence and 80% power.
›At 1,000 eligible new users daily split 50/50, 37,416 total users takes about 38 days. Allow time for the outcome window too.

Formula

Approximate N per variant = 2 x (Z_power + Z_alpha)^2 x p x (1-p) / absolute MDE^2

Pro tip: Use the pre-test planner for a sample-size estimate based on your settings. If the required duration is too long to keep conditions stable, reconsider the test or the effect size you need to detect.

3
Design the control and variant around your question

A randomized A/B test can compare one change or a bundle of changes. A single change makes the result easier to interpret; a bundle estimates the combined effect, not each component’s contribution.

›Control (A): The current experience, unchanged. This is your baseline.
›Variant (B): The proposed experience. Document each difference from the control.
›Resist the temptation to "while we are at it" add multiple changes. If you change the button color AND the headline AND add a trust badge, you cannot know which change drove the result.
›A redesign can be tested as a bundle against the existing experience. Use a factorial or multivariate design if you need to separate component effects and have enough traffic for that design.
›Verify assignment, exposure logging, and delivery of both experiences. Monitor load time so an implementation problem does not go unnoticed.
›Check for sample ratio mismatch using a statistical test. A 55/45 split against a 50/50 target may reflect chance in a small sample or a data-quality problem in a large one.

Formula

Variant = Control + documented changes being evaluated

Pro tip: An A/A test compares identical experiences and can help check assignment and logging. One significant result can occur by chance, so investigate it rather than treating it as proof of a bug.

4
Randomize Users and Launch the Test

Random assignment helps make the groups comparable so differences in outcomes can be attributed to the tested experience, subject to the design and measurement assumptions.

›For a user-level test, use a stable identifier so returning users keep their assignment. Logged-out users may need a persistent browser identifier, with its limitations documented.
›Use a 50/50 split for maximum statistical power. Unequal splits (e.g., 90/10) require significantly larger sample sizes to detect the same effect.
›For long-term effects, consider a separately planned holdout. The control group already provides the baseline for the A/B comparison.
›Log every exposure: user ID, variant assigned, timestamp, and session context. This data is needed for post-analysis and debugging.
›Start measuring after deployment and exposure logging are verified. Apply eligibility rules consistently; prior use of the existing product does not by itself disqualify a customer.
›Enter exposure and conversion counts into the PM Toolkit A/B test suite for analysis. It does not collect product exposure events for you.

Formula

Variant Assignment = hash(user_id + experiment_id) mod 2; stable for each user

Pro tip: If one user’s treatment affects another user’s outcome, consider randomizing by team or organization. That design needs analysis and sample-size planning at the group level.

5
Monitor the Test Without Peeking

Stopping a fixed-sample test early because it looks significant increases the false-positive rate. Follow the stopping rules agreed before launch.

›Repeatedly checking a fixed-sample test and stopping at the first significant result increases the false-positive rate. The increase depends on when and how often you check.
›Set a predetermined end date based on your sample size calculation. Do not stop until you have reached this date AND the required sample size.
›Monitor guardrail metrics (latency, error rate, revenue per user, core engagement), not your primary test metric. If guardrails degrade significantly, stopping is justified.
›If your business genuinely requires early stopping (time-sensitive launches, safety issues), use sequential testing methods like SPRT (Sequential Probability Ratio Test) or mSPRT, which control false positive rate while allowing valid early stopping.
›Run the test for at least one full business cycle (typically 7-14 days) to capture weekday/weekend behavioral differences.
›Document the planned duration and stopping rules in the experiment log before launch.

Formula

Estimated days = required users per variant / eligible new users per variant per day

Pro tip: Use dashboard alerting on guardrail metrics only during the test. Hide the primary metric from daily dashboards for experiment owners during the run to reduce the temptation to peek. Some teams adopt a strict "two-key" policy where two people must agree before a test is stopped early. Apply the same rigor you would to changing any high-stakes business process.

6
Analyze Results with Statistical Significance Testing

Estimate the effect and its uncertainty, then assess whether it is large enough to matter. Statistical significance is one part of that review.

›Calculate the p-value: the probability of observing your result (or more extreme) if the null hypothesis (no difference) were true. A p-value below 0.05 means you reject the null hypothesis.
›Report the confidence interval as the range of effect sizes consistent with the data under the method’s assumptions. For a matching two-sided test, excluding zero corresponds to statistical significance.
›Check practical significance alongside statistical significance: a 0.1% lift with p = 0.001 is statistically significant but may not be worth deploying if the implementation cost is high.
›Analyze secondary metrics and guardrails: did the winning variant improve your primary metric without degrading other important signals?
›Explore results by device, tenure, geography, or plan. Treat unplanned segment findings as exploratory and verify them before making separate claims of a win.
›Record the hypothesis, method, results, and decision, including inconclusive or negative results. This helps future teams understand what was tested.

Formula

p-value < 0.05 AND confidence interval excludes 0 → Statistically significant result

Pro tip: Use the PM Toolkit A/B Test Post-Analysis tool to calculate p-values, confidence intervals, and effect sizes from your raw counts. A statistically significant result is not a deployment mandate. Factor in implementation complexity, technical debt, downstream effects, and strategic alignment before shipping a variant. Statistical significance is one input to a business decision, not the decision itself.

Run Better A/B Tests with PM Toolkit

Our free 3-part A/B test suite covers pre-test sample size planning, live test monitoring, and post-analysis with statistical significance calculations.

Frequently Asked Questions

How long should an A/B test run?

For a fixed-sample test, plan both the required sample and a duration covering relevant usage cycles. A weekly product may need 7-14 days or longer. Wait for outcomes to mature; use a sequential design if you need valid early stopping.

What is statistical significance and why does it matter?

Statistical significance (typically p < 0.05 at 95% confidence) means that if there were truly no difference between your control and variant, the probability of observing a result as extreme as yours by random chance alone is less than 5%. It does not mean your result is correct with 95% probability. It means you have enough evidence to reject the null hypothesis of no difference. Always combine statistical significance with practical significance (is the effect size worth acting on?) before making a deployment decision.

What is the minimum sample size for an A/B test?

There is no universal minimum. At a 3% baseline, detecting a 0.5-percentage-point lift with 95% two-sided confidence and 80% power requires an estimated 18,708 users per variant in PM Toolkit. Smaller effects need larger samples. Other designs or formulas can give different estimates.

Can I run multiple A/B tests at the same time?

Yes, with caution. Simultaneous tests are valid as long as the tested changes are on different parts of the user flow and are unlikely to interact. If two tests could affect the same user behavior or metric, run them sequentially or use a factorial design. When running multiple tests, keep strict user-level assignment consistency: a user in experiment A should have a stable, random assignment in experiment B independent of their experiment A assignment.

What is peeking bias in A/B testing?

Peeking bias occurs when a fixed-sample test is stopped early because it looks significant. Repeated checks increase the false-positive rate by an amount that depends on the stopping rule. Follow the planned sample and duration, or use a sequential method designed for valid repeated checks.

Sources