TY - JOUR
T1 - Machine Learning for Prediction of High-Risk Infections in Patients With Cancer
AU - Fuglkjær, Alexander Djupnes
AU - Eskesen, Mathias Holmsgaard
AU - Werling, Mikkel
AU - Simonsen, Mikkel Runason
AU - Poulsen, Laurids Østergaard
AU - Niemann, Carsten Utoft
AU - Jensen, Paw
AU - Søgaard, Kirstine Kobberøe
AU - Christensen, Frederik
AU - Nielsen, Izabela Ewa
AU - El-Galaly, Tarec Christoffer
N1 - Publisher Copyright:
© 2026 The Author(s). Cancer Medicine published by John Wiley & Sons Ltd.
PY - 2026/7
Y1 - 2026/7
N2 - Purpose: Infectious complications in patients with cancer are contributors to hospitalisations, treatment disruption, and mortality. This study aimed to develop machine learning (ML) models for risk stratification of infection-related hospitalisations (IRHs) among patients with cancer. Methods: Adult patients diagnosed with and treated for lymphoma, multiple myeloma (MM), chronic lymphocytic leukaemia (CLL), colorectal, or lung cancer from 2013 to 2023 in the North Denmark Region were included. Data from national cancer registries and patient health records were used. Serious infections were defined as sepsis, positive blood culture, intensive care admission, or death during hospitalisation. Models were developed using multiple ML methodologies across three different data setups: ML—using data up to hospital admission; ML48—using data up to 48 h after admission; and ML48,reduced—a pruned version of ML48 with fewer features. Data were split into training, validation and held-out test sets. Results: Among 9874 patients 2661 IRHs occurred, including 691 (21.5%) serious infections. On the held-out test set, ML achieved the best performance (ROC-AUC: 0.79) versus ML48 (0.76), with ML48 having higher specificity at matched sensitivity (50.9% vs 44.9% for sensitivity ~85%). ML48,reduced reached a ROC-AUC of 0.75 (specificity 46.5% at sensitivity ~85%). Subgroup analysis showed the best performance for lymphoma (ROC-AUC: 0.86) and lowest for lung cancer (ROC-AUC: 0.73). The model performed better in patients with ECOG score 0–1 compared to ≥ 2 (ROC-AUC 0.88 vs. 0.77). SHAP analysis identified previous blood cultures, carbamide, and neutrophil counts as key predictive features across all models. Conclusion: The ML models showed moderate performance in risk stratification of IRHs in patients with cancer. Such models could provide decision support in hospitalisation decisions and discharge.
AB - Purpose: Infectious complications in patients with cancer are contributors to hospitalisations, treatment disruption, and mortality. This study aimed to develop machine learning (ML) models for risk stratification of infection-related hospitalisations (IRHs) among patients with cancer. Methods: Adult patients diagnosed with and treated for lymphoma, multiple myeloma (MM), chronic lymphocytic leukaemia (CLL), colorectal, or lung cancer from 2013 to 2023 in the North Denmark Region were included. Data from national cancer registries and patient health records were used. Serious infections were defined as sepsis, positive blood culture, intensive care admission, or death during hospitalisation. Models were developed using multiple ML methodologies across three different data setups: ML—using data up to hospital admission; ML48—using data up to 48 h after admission; and ML48,reduced—a pruned version of ML48 with fewer features. Data were split into training, validation and held-out test sets. Results: Among 9874 patients 2661 IRHs occurred, including 691 (21.5%) serious infections. On the held-out test set, ML achieved the best performance (ROC-AUC: 0.79) versus ML48 (0.76), with ML48 having higher specificity at matched sensitivity (50.9% vs 44.9% for sensitivity ~85%). ML48,reduced reached a ROC-AUC of 0.75 (specificity 46.5% at sensitivity ~85%). Subgroup analysis showed the best performance for lymphoma (ROC-AUC: 0.86) and lowest for lung cancer (ROC-AUC: 0.73). The model performed better in patients with ECOG score 0–1 compared to ≥ 2 (ROC-AUC 0.88 vs. 0.77). SHAP analysis identified previous blood cultures, carbamide, and neutrophil counts as key predictive features across all models. Conclusion: The ML models showed moderate performance in risk stratification of IRHs in patients with cancer. Such models could provide decision support in hospitalisation decisions and discharge.
KW - haematological malignancies
KW - infections risk prediction
KW - machine learning
KW - precision oncology
KW - risk stratification
KW - solid tumours
UR - https://www.scopus.com/pages/publications/105045911236
U2 - 10.1002/cam4.72116
DO - 10.1002/cam4.72116
M3 - Journal article
C2 - 42503233
AN - SCOPUS:105045911236
SN - 2045-7634
VL - 15
JO - Cancer Medicine
JF - Cancer Medicine
IS - 7
M1 - e72116
ER -