How Many Patients Do You Actually Need?

A Plain-Language Guide to Statistical Power

Does “no significant difference” mean the treatments are equal?

A trial compares two treatments and finds no significant difference. It’s tempting to conclude the treatments work equally well. But there’s another explanation that gets overlooked constantly: maybe the study simply didn’t enroll enough patients to detect a real difference, even if one exists. That’s the idea behind statistical power.

What Statistical Power Really Is

In earlier posts, we talked about the p-value and the risk of a false positive: concluding there’s an effect when there really isn’t one. Power is the flip side of that coin. It’s the probability that a study will detect a real effect, assuming one truly exists.

A study with low power can easily miss a genuine effect. A study with high power is far more likely to catch it. Power isn’t about whether your result is correct. It’s about whether your study was even capable of finding the answer in the first place.

The Four Ingredients That Drive Power

Four factors determine how much power a study has.

  • Effect size. Larger, more obvious differences are easier to detect than small, subtle ones.

  • Sample size. More patients give you a clearer picture and more power to detect real effects.

  • Variability. The more scattered your data is, the harder it is to spot a true signal underneath the noise.

  • Significance threshold. A stricter cutoff (like p < 0.01 instead of p < 0.05) makes it harder to find significance, which lowers power.

Researchers usually can’t control effect size or variability much. Sample size is the lever they pull most often to boost power.

Why “Not Significant” Doesn’t Mean “No Effect”

This is the most important lesson in this post: absence of evidence is not evidence of absence.

A study that fails to find a significant result may simply have been too small to detect an effect that’s actually there.

Before trusting a “no difference” conclusion, ask whether the study had enough patients to detect a meaningful effect in the first place. If it didn’t, the honest answer is “we don’t know,” not “there’s no effect.”

The 80% Rule of Thumb

Most studies are designed to target 80% power. In plain terms: if the same study were repeated 100 times, and a real effect truly exists, about 80 of those 100 attempts would successfully detect it.

That also means roughly 20 attempts would miss it, purely due to chance. Eighty percent isn’t a guarantee. It’s simply the balance most researchers accept between running a study that’s large enough to be reliable and one that’s still practical to conduct.

Key Takeaways

  • Statistical power is the probability of detecting a real effect, assuming one exists.

  • Power depends on effect size, sample size, variability, and the significance threshold.

  • A “not significant” result can mean there’s truly no effect, or it can mean the study was too small to find one.

  • 80% power is the common target: it means a real effect would be detected in about 80 out of 100 repeated studies.

Up Next:

Next, we’ll look at bias in research design, and why even a perfectly powered, perfectly analyzed study can still reach the wrong conclusion.

Medicine. Research. Analytics.

Medicine. Research. Analytics.