Spring til hovednavigation Spring til søgning Spring til hovedindhold

Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials

Christian Møller Jensen*, Ramtin Zargari Marandi, Kasper Sommerlund Moestrup, Ahmad Mourad, Alfredo J. Mena Lora, Brad T. Sherman, David M. Vock, Jacqueline A. Nordwall, Joanne M. Carson, Nathan Peiffer-Smadja, Neil R. Aggarwal, Nicole E. Naiman, Samuel M. Brown, Thomas W. Barrett, Timothy Hatlen, Victoria S. Kjærgaard, Weizhong Chang, Matthew R. Sydes, Jens Lundgren, Tomas O. Jensenfor the STRIVE Network and ITAC and TICO Study Groups

*Corresponding author af dette arbejde

Abstract

Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort. Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC. Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (−15.8%), but performance remained above chance level. Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors. Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.

OriginalsprogEngelsk
Artikelnummerooag079
TidsskriftJAMIA Open
Vol/bind9
Udgave nummer3
DOI
StatusUdgivet - jun. 2026

Fingeraftryk

Dyk ned i forskningsemnerne om 'Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods: analysis from 2 international randomized controlled trials'. Sammen danner de et unikt fingeraftryk.

Citationsformater