A study tests a null hypothesis: that there is no real effect (no difference, no relationship). The graph shows two worlds. In the grey one the null hypothesis is true; in the amber one there is a real effect. Each curve shows how the study's result would vary if it were run many times.
- The threshold (the dashed line) is the significance level. A result to the right of it counts as significant: the null hypothesis is rejected.
- Type I error (α), red: rejecting the null hypothesis when it is true, a false positive. Its rate is the significance level, the share of the grey curve beyond the threshold, such as 5% (p < 0.05).
- Type II error (β), blue: failing to reject the null hypothesis when there is a real effect, a false negative: the share of the amber curve before the threshold.
- Power (1 − β) is the chance of detecting a real effect.
- Effect size is how far apart the two worlds are, measured in standard deviations.
The trade-off
Moving the threshold right (a stricter significance level, such as 1%) makes Type I errors rarer but Type II errors more common, and the reverse. With the same study you can't make both small. What does help: a larger sample (each curve gets narrower, so they overlap less) and a larger effect. Psychology usually uses 5%; a stricter 1% is used when a false positive would be costly, for example before a new treatment is used.
Common exam mistakes: mixing up the two errors (Type I is the false positive); saying a significant result proves the hypothesis; and thinking a non-significant result proves there is no effect, when it may be a Type II error.
Objective: explain Type I and Type II errors, significance levels and the trade-off between them (for example, Cambridge International AS & A Level Psychology 9990 and IB Psychology, research methods and inferential testing; any course's research skills).