TY - JOUR
T1 - Comparing risk factors in severe COVID-19 using machine learning and non-machine learning methods
T2 - analysis from 2 international randomized controlled trials
AU - Møller Jensen, Christian
AU - Zargari Marandi, Ramtin
AU - Moestrup, Kasper Sommerlund
AU - Mourad, Ahmad
AU - Mena Lora, Alfredo J.
AU - Sherman, Brad T.
AU - Vock, David M.
AU - Nordwall, Jacqueline A.
AU - Carson, Joanne M.
AU - Peiffer-Smadja, Nathan
AU - Aggarwal, Neil R.
AU - Naiman, Nicole E.
AU - Brown, Samuel M.
AU - Barrett, Thomas W.
AU - Hatlen, Timothy
AU - Kjærgaard, Victoria S.
AU - Chang, Weizhong
AU - Sydes, Matthew R.
AU - Lundgren, Jens
AU - Jensen, Tomas O.
AU - Polizzotto, Mark N.
AU - Nordwall, Jacqueline
AU - Babiker, Abdel G.
AU - Phillips, Andrew
AU - Vock, David M.
AU - Eriobu, Nnakelu
AU - Kwaghe, Vivian
AU - Paredes, Roger
AU - Mateu, Lourdes
AU - Ramachandruni, Srikanth
AU - Narang, Rajeev
AU - Jain, Mamta K.
AU - Lazarte, Susana M.
AU - Baker, Jason V.
AU - Frosch, Anne E.P.
AU - Poulakou, Garyfallia
AU - Syrigos, Konstantinos N.
AU - Arnoczy, Gretchen S.
AU - McBride, Natalie A.
AU - Robinson, Philip A.
AU - Sarafian, Farjad
AU - Bhagani, Sanjay
AU - Taha, Hassan S.
AU - Benfield, Thomas
AU - Liu, Sean T.H.
AU - Antoniadou, Anastasia
AU - Jensen, Jens Ulrik Stæhr
AU - Kalomenidis, Ioannis
AU - Susilo, Adityo
AU - Hariadi, Prasetyo
AU - Jensen, Tomas O.
AU - Morales-Rull, Jose Luis
AU - Helleberg, Marie
AU - Meegada, Sreenath
AU - Johansen, Isik S.
AU - Canario, Daniel
AU - Fernández-Cruz, Eduardo
AU - Metallidis, Simeon
AU - Shah, Amish
AU - Sakurai, Aki
AU - Koulouris, Nikolaos G.
AU - Trotman, Robin
AU - Weintrob, Amy C.
AU - Podlekareva, Daria
AU - Hadi, Usman
AU - Lloyd, Kathryn M.
AU - Røge, Birgit Thorup
AU - Saito, Sho
AU - Sweerus, Kelly
AU - Malin, Jakob J.
AU - Lübbert, Christoph
AU - Muñoz, Jose
AU - Cummings, Matthew J.
AU - Losso, Marcelo H.
AU - Turner, Dan
AU - Shaw-Saliba, Kathryn
AU - Dewar, Robin
AU - Highbarger, Helene
AU - Lallemand, Perrine
AU - Rehman, Tauseef
AU - Gerry, Norman
AU - Arlinda, Dona
AU - Chang, Christina C.
AU - Grund, Birgit
AU - Holbrook, Michael R.
AU - Holley, Horace P.
AU - Hudson, Fleur
AU - McNay, Laura A.
AU - Murray, Daniel D.
AU - Pett, Sarah L.
AU - Shaughnessy, Megan
AU - Smolskis, Mary C.
AU - Touloumi, Giota
AU - Wright, Mary E.
AU - Doyle, Mittie K.
AU - Popik, Sharon
AU - Hall, Christine
AU - Ramanathan, Roshan
AU - Cao, Huyen
AU - Mondou, Elsa
AU - Willis, Todd
AU - Thakuria, Joseph V.
AU - Yel, Leman
AU - Higgs, Elizabeth
AU - Kan, Virginia L.
AU - Lundgren, Jens
AU - Neaton, James D.
AU - Lane, H. Clifford
AU - for the STRIVE Network and ITAC and TICO Study Groups
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press on behalf of the American Medical Informatics Association. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
PY - 2026/6
Y1 - 2026/6
N2 - Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort. Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC. Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (−15.8%), but performance remained above chance level. Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors. Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
AB - Objective: To compare differences in risk factors and 90-day mortality prediction from 2 machine learning (ML) models with a previously published non-ML model and investigate their validity in an external cohort. Materials and methods: Prospectively collected data from 2 separate randomized controlled trial (RCT) cohorts from 2020 to 2021, the Therapeutics for Inpatients with COVID-19 (TICO/ACTIV-3) Trial (derivation and internal validation cohort) and the Inpatient Treatment with Anti-Coronavirus Immunoglobulin (ITAC) Trial (external validation cohort) were used. Data were collected from 114 sites in 10 countries (TICO/ACTIV-3) and 63 sites in 11 countries (ITAC). A ML pipeline including 5 classification models, and 1 survival model was used for risk factor identification and clinical outcome prediction. Risk factors were compared between a ML-based classification model, a ML-based survival model and a previously published Cox model. Performance of the ML-based classification model was compared across TICO/ACTIV-3 and ITAC. Results: A total of 2625 (TICO/ACTIV-3) and 579 (ITAC) adults hospitalized for COVID-19 were included. Some overlap of risk factors was identified across models. Five were identified in all models, 3 only in ML models, and 4 only in the non-ML model. The ML model showed good predictive performance in TICO/ACTIV-3. Internal validation showed no overfitting. Lower model performance was observed in ITAC (−15.8%), but performance remained above chance level. Discussion: Differences in methods for risk factor identification using ML and non-ML complicates the comparison of results derived from each approach, but using multiple approaches may unveil overlooked risk factors. Conclusion: Risk factor identification may benefit from integrating both ML and non-ML methods, but external validation is necessary, even in RCTs.
KW - biostatistics
KW - machine learning
KW - medical informatics
KW - randomized controlled trial
KW - SARS-CoV-2
UR - https://www.scopus.com/pages/publications/105042589154
U2 - 10.1093/jamiaopen/ooag079
DO - 10.1093/jamiaopen/ooag079
M3 - Journal article
C2 - 42344108
AN - SCOPUS:105042589154
SN - 2574-2531
VL - 9
JO - JAMIA Open
JF - JAMIA Open
IS - 3
M1 - ooag079
ER -