TY - JOUR
T1 - Multi-centre generalisability of deep learning-based dose prediction for head and neck radiotherapy using DAHANCA real-world data
T2 - Deep learning dose prediction in DAHANCA HNC RT
AU - Nielsen, Camilla P.
AU - Huet-Dastarac, Margerie
AU - Jensen, Kenneth
AU - Brink, Carsten
AU - Smulders, Bob
AU - Holm, Anne I.S.
AU - Nielsen, Martin S.
AU - Sibolt, Patrik
AU - Lee, John A.
AU - Sterpin, Edmond
AU - Zukauskaite, Ruta
AU - Johansen, Jørgen
AU - Krogh, Simon L.
AU - Konrad, Maximilian L.
AU - Friborg, Jeppe
AU - Sommer, Jeanette F.A.
AU - Stougaard, Sarah W.
AU - Overgaard, Jens
AU - Toustrup, Kasper
AU - Lonkvist, Camilla K.
AU - Farhadi, Mohammad
AU - Kaplan, Laura P.
AU - Kjeldsen, Rasmus
AU - Lorenzen, Ebbe L.
AU - Barragán-Montero, Ana M.
AU - Hansen, Christian R.
N1 - Publisher Copyright:
© 2026 The Author(s)
PY - 2026/9
Y1 - 2026/9
N2 - Introduction: Deep learning dose prediction shows promise for automated radiotherapy planning and quality assurance in head and neck cancer (HNC). Clinical adoption requires validation of model generalisability using clinically relevant metrics. The study evaluated generalisability of a dose prediction model trained on single-centre data in a large multi-centre cohort. Materials and methods: A deep neural network was trained on 388 HNC treatment plans from one Danish centre and evaluated on an internal test dataset of 42 patients and on a national DAHANCA cohort of 560 plans from six institutions. Performance was evaluated using dose metrics and compared with a median model assigning each OAR the median Dmean from the training data. Expected toxicity was explored using NTCP and the normalised toxicity index (NTI). Results: Predictions closely matched clinically applied dose distributions with median Dmean differences of −0.1 Gy [interquartile range −0.5, 0.3] for PTVs and 1.1 Gy [-0.6, 3.7] for OARs across cohorts. The median-based model showed larger interquartile ranges, with Dmean differences of 0.7 Gy [0.0, 2.1] for PTVs and −4.6 Gy [-14.6, 3.9] for OARs. Comparison of expected toxicity between model-predicted and clinically applied dose distributions showed median NTI differences of 4.4% [ 0.4, 9.4] and median NTCP differences of about 1–3% for xerostomia and dysphagia grade 2+. Conclusions: The dose prediction model outperformed the median-based model, with prediction-plan differences within interquartile ranges observed across cohorts, supporting generalisability. Discrepancies between model-predicted and clinically planned toxicity may indicate suboptimality in clinical plans and could inform quality assurance.
AB - Introduction: Deep learning dose prediction shows promise for automated radiotherapy planning and quality assurance in head and neck cancer (HNC). Clinical adoption requires validation of model generalisability using clinically relevant metrics. The study evaluated generalisability of a dose prediction model trained on single-centre data in a large multi-centre cohort. Materials and methods: A deep neural network was trained on 388 HNC treatment plans from one Danish centre and evaluated on an internal test dataset of 42 patients and on a national DAHANCA cohort of 560 plans from six institutions. Performance was evaluated using dose metrics and compared with a median model assigning each OAR the median Dmean from the training data. Expected toxicity was explored using NTCP and the normalised toxicity index (NTI). Results: Predictions closely matched clinically applied dose distributions with median Dmean differences of −0.1 Gy [interquartile range −0.5, 0.3] for PTVs and 1.1 Gy [-0.6, 3.7] for OARs across cohorts. The median-based model showed larger interquartile ranges, with Dmean differences of 0.7 Gy [0.0, 2.1] for PTVs and −4.6 Gy [-14.6, 3.9] for OARs. Comparison of expected toxicity between model-predicted and clinically applied dose distributions showed median NTI differences of 4.4% [ 0.4, 9.4] and median NTCP differences of about 1–3% for xerostomia and dysphagia grade 2+. Conclusions: The dose prediction model outperformed the median-based model, with prediction-plan differences within interquartile ranges observed across cohorts, supporting generalisability. Discrepancies between model-predicted and clinically planned toxicity may indicate suboptimality in clinical plans and could inform quality assurance.
KW - DAHANCA
KW - Dose prediction
KW - Head and neck cancer
KW - Multi-centre
KW - Quality assurance
KW - Radiotherapy planning
UR - https://www.scopus.com/pages/publications/105044802129
U2 - 10.1016/j.radonc.2026.111702
DO - 10.1016/j.radonc.2026.111702
M3 - Journal article
C2 - 42468605
AN - SCOPUS:105044802129
SN - 0167-8140
VL - 222
JO - Radiotherapy and Oncology
JF - Radiotherapy and Oncology
M1 - 111702
ER -