Sample size and type 2 error move in opposite directions. As a study gets larger, the chance of a type 2 error goes down. A type 2 error is a false negative — the study misses a real effect that actually exists. Small studies are the main reason this happens, and it is one of the most common reasons a real finding never gets confirmed.
What Is a Type 2 Error?
A type 2 error is a false negative. The study concludes there is no effect when a real effect is there. It is different from a type 1 error, which is a false positive — claiming an effect that does not exist.
Say a drug truly lowers blood pressure by a small but real amount. If a study enrolls too few people, the results may scatter so widely that the real drop never rises above the noise. The researchers report “no significant difference.” The effect was real. The study just could not see it. That is a type 2 error.
Researchers describe the chance of avoiding this mistake as statistical power. Power is the probability of detecting an effect if one truly exists. A study with high power is unlikely to miss a real effect. A study with low power is likely to miss it. This is why power and type 2 error are two sides of the same coin.
Conventional practice sets power at 80% or higher when planning a study. At 80% power, the type 2 error rate is 20%. That number — 20% — surprises many people. It means a study can be considered adequately designed and still miss a real effect one time in five.
How Does Sample Size Affect Type 2 Error?
Sample size is the main lever that controls type 2 error. More participants means more information, and more information means the study can separate a real signal from random variation.
Here is the mechanism. Every measurement carries random variation. When you study only a few people, that random variation is large relative to any real effect. The effect gets buried. When you study many people, the random variation averages out, and a real effect becomes visible.
This relationship is not linear. Precision improves with the square root of sample size. Going from 100 to 400 participants does not double your precision — it roughly doubles it, but you had to add 300 people to get there. Going from 400 to 1,600 does the same again, and now you have added 1,200. Each additional gain in precision costs more participants than the last.
The practical takeaway: small studies are not just less precise. They are systematically more likely to produce false negatives. A trial with 30 people per group can easily miss a moderate effect that a trial with 300 would detect without difficulty.
Why Do Small Studies Produce So Many False Negatives?
Small studies have wide confidence intervals. A confidence interval is the range of values that plausibly contains the true effect. When that range is wide, it can include both “meaningful benefit” and “no benefit at all” — and the study cannot tell you which is correct.
Three things happen at once in a small study:
- The estimate of the effect is unstable. Run the study again with different people and you may get a very different answer.
- The confidence interval is wide, so the result is compatible with many different truths.
- The threshold for “statistical significance” becomes harder to cross, even when a real effect is present.
This is why a single small study rarely settles a question. It can suggest something worth investigating, but it usually cannot confirm or rule out an effect on its own.
There is a non-obvious point here that often gets lost. A small study that does find a significant result is not necessarily more trustworthy than a large one — it may be more likely to have overestimated the effect. The combination of small samples and selective reporting is a well-documented problem in medical research. Large studies tend to give estimates closer to the truth, in both directions.
What Else Affects Type 2 Error Besides Sample Size?
Sample size is the biggest factor, but it is not the only one. Four other things change the risk of a type 2 error:
- Effect size. Big effects are easier to detect than small ones. A drug that cuts mortality in half will show up in a modest study. A drug that reduces risk by 5% may need thousands of participants.
- Variability of the measurement. Noisy measurements hide real effects. A precise measurement tool lowers the sample size needed.
- The significance threshold. A stricter threshold (a smaller p-value cutoff) reduces false positives but raises the risk of false negatives. There is a trade-off between the two error types.
- Study design. Poor design, dropout, or measurement error all reduce effective power, even if the planned sample size looked adequate on paper.
These factors interact. A study looking for a small effect with a noisy measurement and a strict threshold needs a very large sample to have adequate power.
How Do Researchers Calculate the Sample Size They Need?
Researchers run a power calculation before the study begins. This calculation estimates how many participants are needed to detect a given effect size with a set level of power.
The calculation requires four inputs:
- The smallest effect size that would be clinically meaningful
- The expected variability in the measurement
- The chosen significance level (usually 0.05)
- The desired power (usually 80% or 90%)
The output is a target sample size. If the study cannot enroll that many people, the researchers know in advance that they are at higher risk of a type 2 error.
This is where a common problem appears. Many published studies are underpowered. They were designed with sample sizes too small to reliably detect the effects they were looking for. A null result from such a study does not mean the treatment does not work. It means the study was not capable of answering the question.
What Does a “Negative” Result Actually Mean?
A negative result does not prove absence of effect. It means the study did not detect an effect. Those are different statements.
This distinction matters for how you read health news. When a headline says a treatment “doesn’t work,” check the study size. A small trial with a null result is weak evidence of no effect. A large trial with a null result is strong evidence of no meaningful effect.
The confidence interval tells you more than the p-value. If a study reports no significant difference but the confidence interval is wide, the data are compatible with a real benefit. If the interval is narrow and centered near zero, the data suggest any real effect is probably small.
This is why systematic reviews and meta-analyses matter. By combining many studies, they increase the effective sample size and reduce the chance that a real effect gets lost in the noise of any single small trial.
Why Underpowered Studies Waste Time and Money
An underpowered study is not a neutral exercise. It carries real costs.
Participants take on risk and inconvenience for a question the study cannot answer. Funding gets spent. And the resulting null result may discourage further research into a treatment that actually works.
There is also an ethical dimension. Research ethics guidelines generally hold that studies should be designed with adequate power to produce meaningful answers. Enrolling people in a study that is too small to detect a plausible effect raises concerns about whether the research justifies the burden on participants.
The flip side is also true. Oversized studies can detect effects so small they have no practical meaning. A huge trial might find a statistically significant difference that no patient would ever notice. Good study design aims for the sample size that answers the question that matters — not the largest sample possible.
How to Judge a Study’s Risk of Type 2 Error
You do not need to run statistics to spot a study at risk of a type 2 error. A few questions help:
- How many participants were in each group?
- Did the researchers report a power calculation?
- How wide is the confidence interval around the main result?
- Was the effect size they were looking for large or small?
- Does the conclusion match what the study was actually powered to detect?
If a study reports “no significant difference” but enrolled few participants and had a wide confidence interval, be cautious about concluding the treatment does not work. The honest reading is that the study was not designed to settle the question.
Frequently Asked Questions
Does a larger sample size reduce type 2 error?
Yes. Increasing sample size raises statistical power, which lowers the probability of a type 2 error. The improvement follows the square root of the sample size, so each gain requires more participants than the last.
What is the relationship between sample size and statistical power?
Power rises as sample size grows, because more data make it easier to separate a real effect from random variation. A study with low power is more likely to miss an effect that truly exists.
Can a study with a small sample still detect a real effect?
Yes, but usually only if the effect is large. Small studies can detect big effects but often miss small or moderate ones, which is why they produce more false negatives.
Does a negative result from a small study mean the treatment does not work?
No. A negative result means the study did not detect an effect, not that no effect exists. Small studies with wide confidence intervals often cannot rule out a real benefit.

