Dimitri's Letters · Issue 4
The Five Statistical Mistakes That Kill Papers
Sent to subscribers on . Details may have changed since.
Most papers are not rejected because the science is wrong. They are rejected because the statistics are handled in a way that makes reviewers doubt the science, even when the underlying work is solid.
I have reviewed hundreds of manuscripts. The statistical errors I see are almost always the same ones, made by smart people who simply were not taught what they needed to know. None of them are exotic. All of them are avoidable.
Here are the five I see most often, with concrete examples of what they look like in practice.
The five statistical traps
-
Treating p=0.05 as the finish line
A p-value tells you the probability of seeing your result if the null hypothesis were true. It does not tell you the size or clinical importance of the effect. A study with 2,000 patients can produce p=0.001 for a difference that is statistically significant but clinically meaningless. Always report effect sizes and confidence intervals alongside p-values. A reviewer who sees only p=0.048 and no effect size will be skeptical, and rightly so.
-
Using the wrong test for the data type
Using a t-test on non-normally distributed data, or a parametric test on ordinal variables such as Likert scale responses, is one of the most common errors in surgical research. Continuous normally distributed data gets parametric tests. Skewed or ordinal data gets non-parametric alternatives such as Mann-Whitney. Before you run any test, state explicitly why it is appropriate for your data structure.
-
Multiple comparisons without correction
If you test twenty outcomes and use p=0.05 as your threshold, by chance alone you will find one significant result even if nothing is truly different. If your paper tests more than one primary outcome, you need a correction strategy such as Bonferroni, or a clear pre-specified hierarchy of primary and secondary endpoints. Reviewers who see a table of twenty p-values with two just below 0.05 and no correction will flag it immediately.
-
Confusing association with causation
An observational study can show that two things are associated. It cannot, on its own, prove that one caused the other. The language in your paper must reflect the study design. Phrases like "X led to" or "X caused" in an observational paper will draw immediate criticism. Use "X was associated with" or "X was independently predictive of" instead. This is not just semantics. It is the honest representation of what your data can and cannot show.
-
Underpowered studies reported as negative
A study that finds no significant difference is not necessarily a negative study. It may simply be too small to detect a real difference that exists. If you did not perform a power calculation before data collection, a non-significant result cannot be interpreted confidently either way. Always report your power calculation, your assumed effect size, and your alpha. A study powered to detect a 20% difference that found only 8% is not a negative study. It is an underpowered one.
The goal is not to become a statistician. The goal is to know enough to avoid the errors that cost papers their acceptance, and to know when the question requires a statistician's input before you begin.
That last point matters. There is a category of study design, including propensity score matching, survival analysis with competing risks, and mixed-effects models, where getting statistical input before data collection is not optional. If you are designing something in that territory and you have not spoken to a biostatistician, do it before you collect a single data point. Fixing the statistics after the fact is expensive. Sometimes it is impossible.
What to do right now
- Pull your last submitted or published paper. Check whether you reported effect sizes and confidence intervals alongside every p-value. If not, that is the first habit to build into your next submission.
- For your current project, write down your primary endpoint before you run any analysis. One primary endpoint. Everything else is secondary. This single discipline eliminates the multiple comparisons problem before it starts.
- Use AskDimitri to check your statistical approach. Describe your study design and your planned analysis. If there is a mismatch or a trap in the plan, it will surface before you collect data rather than after a reviewer finds it.
Check your statistical approach →
Statistics is not the enemy of good research. Bad statistics is. The difference is smaller than most people think, and fixable with less effort than most people assume.
Until next week,