A statistical test can be wrong in two different ways. Good study design starts by understanding both.
A logistics company tests a new route-optimization system. The company could conclude that the system works when it does not, or fail to detect a real reduction in delivery time.
Hypothesis tests always carry both risks: detecting an effect that is not real and missing an effect that is.
Key question
How do false positives and false negatives differ, and how can a study improve its chance of detecting a real effect?
How can a statistical test be wrong?
A test either rejects the null hypothesis or does not. In reality, the effect either exists or it does not. Combining those possibilities produces four outcomes.
| Reality | Reject H₀ | Do not reject H₀ |
|---|---|---|
| No real effect | Type I error A false positive | Correct decision No effect detected |
| A real effect exists | Correct detection Statistical power | Type II error A false negative |
Two outcomes are correct decisions. The other two are Type I and Type II errors.
How do Type I and Type II errors differ?
The new system has no real effect, but sampling variation produces p < 0.05. The probability of this error is controlled by α.
Type I error
The company adopts an ineffective system and continues investing in it.
The system really does reduce delivery time, but the sample is too small or too variable to produce clear evidence. The probability of this error is β.
Type II error
The company abandons a useful system because the study did not collect enough information.
The two risks are connected
With the same data, making α more stringent reduces Type I error but can increase Type II error. Thresholds and sample size must be planned together.
What determines statistical power?
Statistical power is the probability of correctly rejecting the null hypothesis when the assumed effect is real.
Statistical power
1 − β
Power of 80% means that studies conducted under the same assumptions would detect the specified effect about 80% of the time.
Effect size
Larger effects are easier to detect.
Sample size
Larger samples reduce standard error.
Data spread
Less variable data make a signal easier to separate from noise.
Significance level
A less strict threshold raises power but also raises Type I error.
What 80% power does not mean
It does not mean that a significant result has an 80% chance of being true. Power is a long-run detection rate under a specified effect size and study design.
How should sample size be planned?
A power analysis combines the target effect size, significance level, and desired power to estimate the required sample size.
Power analysis
Independent-samples t-test · two-sided
Effect size
d = 0.50
Alpha
α = 0.05
Power
0.80
Required sample
64 per group
Assuming an unrealistically large effect produces an unrealistically small sample requirement. Use prior research, pilot data, or the smallest effect that would matter in practice.
| Input | Common choice |
|---|---|
| Significance level | α = 0.05 |
| Target power | 0.80 or 0.90 |
| Effect size | Prior research or minimum important difference |
| Expected attrition | Add to the required sample |
Reading a non-significant result
In a low-powered study, p > 0.05 often means “not enough information to separate the effect from noise,” not “the effect is zero.”
Change effect size, sample size, and alpha
Adjust the study conditions below and watch Type I error, Type II error, and power move together.
α
0.05
β
0.193
Power
80.7%
Signal / noise
0.50
Key lesson
Type I and Type II errors are not isolated settings. Effect size, sample size, variability, and alphadetermine them together.
A good test does more than calculate a p-value. It makes the acceptable risks explicit before the study begins.
Error criterion
Significance level α
Detection ability
Power 1 − β
Study planning
Effect size · sample size
Result interpretation
p-value · interval · effect size