Follow these 6 steps to plan, run, and analyze an A/B test. Covers hypothesis formation, sample size, test design, peeking bias, and post-analysis for product managers.
Write what you are changing, why you expect it to affect the metric, and how you will measure the result.
Formula
Hypothesis: If [change], then [metric] will [increase/decrease] by [X]% because [reason]Pro tip: Use historical data to develop the hypothesis, then record it before examining test results. Label ideas suggested by the results as exploratory and validate them in a new test.
Plan sample size before a fixed-sample test. Too little data reduces your chance of detecting an improvement of the size you care about.
Formula
Approximate N per variant = 2 x (Z_power + Z_alpha)^2 x p x (1-p) / absolute MDE^2Pro tip: Use the pre-test planner for a sample-size estimate based on your settings. If the required duration is too long to keep conditions stable, reconsider the test or the effect size you need to detect.
A randomized A/B test can compare one change or a bundle of changes. A single change makes the result easier to interpret; a bundle estimates the combined effect, not each component’s contribution.
Formula
Variant = Control + documented changes being evaluatedPro tip: An A/A test compares identical experiences and can help check assignment and logging. One significant result can occur by chance, so investigate it rather than treating it as proof of a bug.
Random assignment helps make the groups comparable so differences in outcomes can be attributed to the tested experience, subject to the design and measurement assumptions.
Formula
Variant Assignment = hash(user_id + experiment_id) mod 2; stable for each userPro tip: If one user’s treatment affects another user’s outcome, consider randomizing by team or organization. That design needs analysis and sample-size planning at the group level.
Stopping a fixed-sample test early because it looks significant increases the false-positive rate. Follow the stopping rules agreed before launch.
Formula
Estimated days = required users per variant / eligible new users per variant per dayPro tip: Use dashboard alerting on guardrail metrics only during the test. Hide the primary metric from daily dashboards for experiment owners during the run to reduce the temptation to peek. Some teams adopt a strict "two-key" policy where two people must agree before a test is stopped early. Apply the same rigor you would to changing any high-stakes business process.
Estimate the effect and its uncertainty, then assess whether it is large enough to matter. Statistical significance is one part of that review.
Formula
p-value < 0.05 AND confidence interval excludes 0 → Statistically significant resultPro tip: Use the PM Toolkit A/B Test Post-Analysis tool to calculate p-values, confidence intervals, and effect sizes from your raw counts. A statistically significant result is not a deployment mandate. Factor in implementation complexity, technical debt, downstream effects, and strategic alignment before shipping a variant. Statistical significance is one input to a business decision, not the decision itself.
Our free 3-part A/B test suite covers pre-test sample size planning, live test monitoring, and post-analysis with statistical significance calculations.
For a fixed-sample test, plan both the required sample and a duration covering relevant usage cycles. A weekly product may need 7-14 days or longer. Wait for outcomes to mature; use a sequential design if you need valid early stopping.
Statistical significance (typically p < 0.05 at 95% confidence) means that if there were truly no difference between your control and variant, the probability of observing a result as extreme as yours by random chance alone is less than 5%. It does not mean your result is correct with 95% probability. It means you have enough evidence to reject the null hypothesis of no difference. Always combine statistical significance with practical significance (is the effect size worth acting on?) before making a deployment decision.
There is no universal minimum. At a 3% baseline, detecting a 0.5-percentage-point lift with 95% two-sided confidence and 80% power requires an estimated 18,708 users per variant in PM Toolkit. Smaller effects need larger samples. Other designs or formulas can give different estimates.
Yes, with caution. Simultaneous tests are valid as long as the tested changes are on different parts of the user flow and are unlikely to interact. If two tests could affect the same user behavior or metric, run them sequentially or use a factorial design. When running multiple tests, keep strict user-level assignment consistency: a user in experiment A should have a stable, random assignment in experiment B independent of their experiment A assignment.
Peeking bias occurs when a fixed-sample test is stopped early because it looks significant. Repeated checks increase the false-positive rate by an amount that depends on the stopping rule. Follow the planned sample and duration, or use a sequential method designed for valid repeated checks.