Statistical significance helps assess evidence of a difference. Practical significance asks whether that difference is large enough to justify the change.
Last updated: 2026-04-01
Evidence that the observed difference is hard to explain by random variation alone if there is no real difference. Reported via a p-value or confidence interval. A common threshold is p < 0.05. See the NIST explanation of tests and intervals.
Best as a guard against acting on noise. It helps assess evidence against no difference; it does not give the probability that an effect is real.
A business judgment about whether the size of the effect is worth shipping, given costs, risk, and opportunity cost. Often expressed as a minimum effect of interest chosen before the test.
Best as a sanity check. A practically significant result tells you the effect is worth acting on, not just measurable.
z = (p1 - p2) / sqrt(p_pool x (1 - p_pool) x (1/n1 + 1/n2))Compare z against the critical value for your alpha. Or compare the resulting p-value against alpha directly. Default alpha = 0.05.
No formula. Set a minimum effect of interest (MEI) before the test starts.The MEI drives the MDE, which drives the sample size. Detecting half the effect requires roughly four times the sample.
| Criteria | Statistical Significance | Practical Significance |
|---|---|---|
| Question answered | How strong is the evidence of a difference? | Is the effect worth it? |
| Type | Statistical claim | Business judgment |
| Reported as | p-value, confidence interval | Effect size weighed against costs and risk |
| When small samples | Critical. Easy to mistake noise for signal | Still needed; a large estimate may have wide uncertainty |
| When huge samples | Can detect very small differences | Critical. Tiny effects pass p < 0.05 without mattering |
| Set when | Choose alpha before the test; calculate results afterward | Before the test, agreed by the team |
| The trap | Shipping anything with p < 0.05 | Ignoring uncertainty around the estimated benefit |
| Healthy practice | Always check both gates | Always check both gates |
Pros
Cons
Pros
Cons
Use your own inputs to explore the calculations and compare results.
The improvement size your test is planned to detect at the chosen power. Set it before the test. Detecting half the effect needs roughly four times the sample, with other assumptions unchanged.
No. It is evidence against no difference under the test assumptions, not the probability that the result is real. Also consider the estimated improvement, its uncertainty, costs, and guardrail metrics.
Yes. A 10% estimated lift might be worth acting on if real, but the data may still allow no improvement or harm. Review the uncertainty and test plan before deciding whether more data is worth collecting.
Start with the smallest improvement that would justify the change, then check whether you can collect enough data to detect it. There is no universal percentage for growth, pricing, or redesign tests.
Larger effects are easier to detect. Planning for a 10% lift usually needs fewer users than a 1% lift. The tradeoff is lower power to detect smaller effects, even when those effects would still be useful.