Statistical vs Practical Significance

Statistical significance helps assess evidence of a difference. Practical significance asks whether that difference is large enough to justify the change.

Last updated: 2026-04-01

Overview

Statistical Significance
Evidence of a Difference

Evidence that the observed difference is hard to explain by random variation alone if there is no real difference. Reported via a p-value or confidence interval. A common threshold is p < 0.05. See the NIST explanation of tests and intervals.

Best as a guard against acting on noise. It helps assess evidence against no difference; it does not give the probability that an effect is real.

Practical Significance
Is It Worth It?

A business judgment about whether the size of the effect is worth shipping, given costs, risk, and opportunity cost. Often expressed as a minimum effect of interest chosen before the test.

Best as a sanity check. A practically significant result tells you the effect is worth acting on, not just measurable.

Formula comparison

Statistical Significance

z = (p1 - p2) / sqrt(p_pool x (1 - p_pool) x (1/n1 + 1/n2))

Compare z against the critical value for your alpha. Or compare the resulting p-value against alpha directly. Default alpha = 0.05.

Practical Significance

No formula. Set a minimum effect of interest (MEI) before the test starts.

The MEI drives the MDE, which drives the sample size. Detecting half the effect requires roughly four times the sample.

Side-by-side comparison

CriteriaStatistical SignificancePractical Significance
Question answeredHow strong is the evidence of a difference?Is the effect worth it?
TypeStatistical claimBusiness judgment
Reported asp-value, confidence intervalEffect size weighed against costs and risk
When small samplesCritical. Easy to mistake noise for signalStill needed; a large estimate may have wide uncertainty
When huge samplesCan detect very small differencesCritical. Tiny effects pass p < 0.05 without mattering
Set whenChoose alpha before the test; calculate results afterwardBefore the test, agreed by the team
The trapShipping anything with p < 0.05Ignoring uncertainty around the estimated benefit
Healthy practiceAlways check both gatesAlways check both gates

When to use each

Choose Statistical Significance when
  • Sample size is small. Small samples are noisy and easy to misread
  • The cost of a false positive is high (a launch that breaks something)
  • Stakeholders will scrutinize the result
  • You're under pressure to ship and need to defend the decision
Choose Practical Significance when
  • Sample size is huge. Tiny differences become "significant" but not meaningful
  • The change has costs beyond the test (engineering, support, risk)
  • You're prioritizing a roadmap of experiments. Small wins eat capacity
  • The effect size is below your minimum effect of interest

Pros and cons

Statistical Significance

Pros

  • Quantitative gate against noise
  • Standard. Stakeholders recognize p-values and confidence intervals
  • Pairs naturally with sample-size planning

Cons

  • Large samples can make tiny differences statistically significant
  • A cutoff can distract from effect size and uncertainty
  • Doesn't speak to whether the effect matters

Practical Significance

Pros

  • Forces the business question before the test
  • Filters out small wins that aren't worth ship cost
  • Makes the test plan honest about the MDE

Cons

  • "Worth it" depends on the team and the moment
  • Easy to forget. Most A/B testing tools don't surface it
  • Planning for a large MDE can leave too little power for smaller useful effects

Try the related tools

Use your own inputs to explore the calculations and compare results.

Frequently asked questions

What is the minimum detectable effect (MDE)?

The improvement size your test is planned to detect at the chosen power. Set it before the test. Detecting half the effect needs roughly four times the sample, with other assumptions unchanged.

Is p < 0.05 enough to ship?

No. It is evidence against no difference under the test assumptions, not the probability that the result is real. Also consider the estimated improvement, its uncertainty, costs, and guardrail metrics.

Can a result look useful but not be statistically significant?

Yes. A 10% estimated lift might be worth acting on if real, but the data may still allow no improvement or harm. Review the uncertainty and test plan before deciding whether more data is worth collecting.

What is the right MDE for my test?

Start with the smallest improvement that would justify the change, then check whether you can collect enough data to detect it. There is no universal percentage for growth, pricing, or redesign tests.

Why does sample size shrink with a larger MDE?

Larger effects are easier to detect. Planning for a 10% lift usually needs fewer users than a 1% lift. The tradeoff is lower power to detect smaller effects, even when those effects would still be useful.