Dimitri's Letters · Issue 9
AI Models Are Getting Smarter. Reporting Them Is Getting Worse.
Sent to subscribers on . Details may have changed since.
If you have built or plan to build a prediction model, this issue is directly for you. Not a general methodology point. A specific, checkable standard that your next submission will very likely be judged against, whether or not anyone tells you in advance.
It is called TRIPOD plus AI. It replaced the older TRIPOD checklist in 2024 and it is now the reporting standard for any study that develops or validates a clinical prediction model, whether the method is logistic regression, a random forest, or a neural network trained on electronic health records.
Here is the part that should genuinely surprise you.
28%
Read that again. The newer, more sophisticated modeling methods are being reported worse, not better, than the simpler methods they are replacing. Sophistication is outrunning rigor. In the same review, no study fully reported its sample size calculation, and none fully met the reporting requirements for the abstract.
The checklist is not bureaucracy. It is the difference between a reviewer trusting your model and a reviewer quietly assuming you do not know what you built.
Most researchers treat model reporting as an afterthought, something to fill in after the analysis is done. TRIPOD plus AI works the opposite way. It is a design document. The items it asks for, such as how missing data was handled, whether the model was internally and externally validated, how performance was measured beyond a single accuracy number, and whether fairness across subgroups was assessed, are questions you should be answering while you build the model, not while you write the paper.
Here are the five items reviewers flag most often, in plain language.
Where machine learning papers fail most
- No external validation. A model tested only on the data it was trained on tells you almost nothing about how it will perform elsewhere. State clearly whether validation used a separate cohort, a separate site, or a separate time period.
- Accuracy reported alone. A single accuracy figure hides how the model performs across calibration, discrimination, and clinically relevant thresholds. Report discrimination such as the C statistic, calibration, and a clinically meaningful metric together.
- Missing data left unexplained. Every model handles missing values one way or another. State the method explicitly, whether that is imputation, exclusion, or something else, rather than letting a reviewer guess.
- No subgroup or fairness assessment. A model that performs well overall can still fail badly for a specific age group, sex, or comorbidity profile. Reviewers increasingly expect this to be checked and reported, not assumed away.
- The model is a black box with no interpretation. Even complex models can report which features drive predictions. If you cannot explain in one paragraph what is driving your model's output, that paragraph is exactly what a reviewer will ask for.
The full checklist has 27 items and it is genuinely worth reading once, in full, before your next prediction model project rather than after. It takes about twenty minutes and it will save you a much longer revision cycle later.
If any of this feels relevant to something you are building right now, this is a good moment to check your plan against it before more time goes in. AskDimitri can walk through your design against the checklist directly.
Check your model against the checklist →
The gap between what a model can do and what a paper proves it does is exactly where reviewers live. Close that gap before they have to point it out.
Editor's note: in September 2026 the figures in this letter were checked against the sources listed below, and the text was corrected where the letter as sent differed from them.
Until next week,
Sources
- Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378. PMID: 38626948.
- Partheniadis I, Talimtzi P, Nikolakopoulou A, Haidich AB. Machine learning-based COVID-19 prognostic models lag behind in reporting quality: findings from a TRIPOD/TRIPOD + AI systematic review. Diagn Progn Res. 2026;10(1):3. doi:10.1186/s41512-026-00218-x. PMID: 41630071.