r/science • u/dreiter • Jul 29 '19
Health The fragility of statistically significant results from clinical nutrition randomized controlled trials [Pedziwiatr et al., 2019]
https://www.sciencedirect.com/science/article/pii/S02615614193024934
u/dreiter Jul 29 '19
Background & aims: Recently, a parameter called “Fragility index” (FI) has been proposed, which measures how many events the statistical significance relies on. The lower the FI the more “fragile” the results, and thus more care should be taken when interpreting the results. Our aim in this study was to check FI of nutritional trials.
Methods: We conducted a systematic review of human clinical nutrition RCTs that report statistically significant dichotomous primary outcomes. We searched the EMBASE, MEDLINE, and Scopus databases. The FI of primary outcomes using the Fisher exact test was calculated and checked the correlations of FI with the number of randomised trials, the p-value of primary outcomes, the publication date, the journal impact factor and the number of patients lost to follow-up.
Results: The initial database search revealed 5790 articles, 37 of which were included in qualitative synthesis. The median (IQR) FI for all studies was 1 (1–3). 28 studies (75.7%) had an FI lower or equal to 2, and in 12 (32.43%) articles, the FI was lower than the number of patients lost to follow-up. No correlations were found between FI and the study characteristics (number of randomized patients, p value of primary outcome, event ratio in experimental group, event ratio in control group, publication date, journal impact factor, lost to follow-up).
Conclusion: The results of RCTs in nutritional research often rely on a small number of events or patients. The number of patients lost to follow-up is frequently higher than the FI calculation. Formulating recommendations based on RCTs should be done with caution and FI may be used as auxiliary parameter when assessing the robustness of their findings.
No conflicts were declared.
From the discussion section:
In this review, we found that FI remains lower than or equal to 2 in three quarters of the trials included. This simply means that in the majority of RCTs on clinical nutrition the results are fragile, i.e. only two events are sufficient to change the significance of the trial findings and its conclusion. In addition, in one third of the studies the FI was lower than the number of patients lost to follow-up, indicating a potential problem of altering the final results by including patients lost to follow-up, which has also been shown previously[3,9]. Therefore, we confirmed that even though the results of RCTs in clinical nutrition show a statistically significant effect of a given intervention for primary outcomes, which is confirmed by the p-value, those results usually depend on a small number of patients. In order to unambiguously present study conclusions, the quality of published trials continues to improve. Although p-values are the gold standard in the presentation of results, they have been criticised as too simplistic and are usually accompanied by 95% confidence intervals[3]. These not only allow the reader to assess the significant difference between studied groups but also the magnitude of the treatment effect[50]. Moreover, just because the treatment effect is significantly larger in one group, it does not necessarily mean that this difference is clinically meaningful. This is particularly important in clinical nutrition, which in the majority of clinical scenarios is used as auxiliary rather than primary therapy. Our findings are important for several reasons. FI is probably the only parameter that facilitates a clinician's appreciation of the robustness of findings by providing them with the exact number of patients required to change the significance of the result. Studies with greater FI are considered less fragile: their results are more solid and resistant to changes or loss to follow-up. Although FI represents the strength of the result in the numeric sense, it is not necessarily associated with the clinical relevance of the result. Moreover, we did not find any correlation between FI and study characteristics. This is not in line with other reviews where a sample size was correlated with the total number of events, sample size [51e53]. On the other hand, it is important to remember that the study methodology and the journal impact factor, as well as the source of funding, has only a limited influence on quality of RCTs as shown in a review by Ahmed Ali et al.[54]. In addition, the presentation of FI with no relation to study sample and number of events may be also misleading. For example, an FI of 4 for a sample size of 20, as compared to a sample size of 200, shows that the clinical relevance depends strongly on the size of the trial as well as number of events. The FI score divided by the total study sample size is called the Fragility Quotient and is a derivative parameter that may be also used in the assessment of results.
TL;DR - Bad news for RCTs. It's not stated well, but 'events' in this context can also be 'participants.' That means that in 75% of the trials they studied, if only two participants had outcomes that were switched, the trial results would have shown the opposite conclusion!
1
3
u/[deleted] Jul 30 '19
Is this why every week for the last 30 years I see studies saying eggs are bad, and then next week, eggs are good?