Glioblastoma is the most aggressive primary malignant tumour of the central nervous system, with a median Overall Survival (OS) that typically remains limited to approximately 12–15 months despite intensive standard-of-care treatment.The marked heterogeneity of the disease motivates patient-specific risk stratification, for which a noninvasive and reproducible prognostic tool would be clinically valuable. This thesis investigates whether multiparametric brain MRI can support deep learning–based prediction of overall survival (OS) in glioblastoma patients, formulated as a threeclass classification task. Beyond model development, the study aims to evaluate the potential clinical applicability of this approach and to unveil the ability of predictive models to generalize across heterogeneous real-world data. The study relies on four publicly available datasets: UPENN-GBM, the largest one, used for training and internal validation, and BraTS2020, UCSF-PDGM, and RHUH-GBM, reserved exclusively to assess the generalizability of the models across independent and heterogeneous datasets. Each of them provides T1, T1ce, T2, and FLAIR sequences. The methodological narrative develops along three blocks. First, a systematic benchmark compares several 3D CNN architectures adapted from established 2D convolutional models (ResNet, DenseNet, GoogleNet, EfficientNet) under a common stratified 5-fold cross-validation protocol; monomodal models are combined into multimodal predictions through prediction-level Softmax ensembling, and predictive performance is interpreted jointly with model complexity. Based on the benchmark results, GoogleNet is selected as the reference architecture for its efficiency–performance trade-off; an ordinal error analysis is subsequently applied to characterise the severity of its misclassifications, distinguishing Long Errors, between non-adjacent classes, from Short Errors, between adjacent ones. Second, two strategies address the limitations identified during benchmarking: an error-sensitive loss extending weighted cross-entropy with an ordinal cost term that weights errors by their severity, and a modality-specific expert ensemble assigning dedicated architectures, loss functions, and augmentation settings to individual MRI sequences; the two are subsequently combined into a unified errorsensitive expert framework. Third, Grad-CAM is used as an explainability tool to inspect the spatial attention regions associated with the model predictions. External testing revealed modest overall performance, consistent with the difficulty of the task and the strict evaluation protocol adopted. The high-level expert ensemble achieved the most competitive result, reaching an accuracy of 0.61 on BraTS2020. This performance was in line with the reported state-of-the-art range. At the same time, the error-sensitive loss consistently reduced Long Errors across all external cohorts and modality combinations. A decrease in performance was observed on datasets processed through different preprocessing pipelines, highlighting the impact of domain shift. Moreover, Grad-CAM revealed that correct predictions were not always associated with attention patterns focused on meaningful anatomical or tumour-related regions. In conclusion, this thesis tackles two key challenges in deep learning–based survival prediction for glioblastoma: robustness to missing MRI modalities and the formulation of survival estimation as an ordinal clinical task. By adopting a sequenceagnostic approach, the proposed framework remains applicable when different combinations of MRI modalities are available, increasing its potential usability in realworld clinical settings. At the same time, the use of ordinal-aware optimisation reduces the impact of clinically severe errors. While the issue of generalisation across heterogeneous external cohorts is only partially resolved, the results provide evidence that the proposed design choices represent a step towards more robust and adaptable models. In this sense, the thesis contributes to the broader integration of deep learning methods in neuro-oncology, with the aim of supporting more flexible and clinically relevant prognostic tools.
Il glioblastoma è il tumore maligno primario più aggressivo del sistema nervoso centrale, con una sopravvivenza globale (Overall Survival, OS) mediana che rimane limitata a circa 12-15 mesi, nonostante il trattamento standard intensivo. La marcata eterogeneità della malattia motiva una stratificazione del rischio paziente-specifica, per la quale uno strumento prognostico non invasivo e riproducibile risulterebbe clinicamente rilevante. Questa tesi indaga se la risonanza magnetica cerebrale multiparametrica possa essere utilizzata per predire l’OS nei pazienti affetti da glioblastoma tramite deep learning, formulando il problema come un compito di classificazione a tre classi. Lo studio si basa su quattro coorti pubblicamente disponibili: UPENN-GBM, impiegata per l’addestramento e la validazione interna, e BraTS2020, UCSF-PDGM e RHUH-GBM, riservate esclusivamente al test esterno, ciascuna fornita delle sequenze T1, T1ce, T2 e FLAIR. La narrazione metodologica si sviluppa lungo tre blocchi. In primo luogo, un benchmark sistematico confronta diverse architetture di rete neurale convoluzionale 3D (3D CNN) adattate da modelli convoluzionali 2D consolidati (ResNet, DenseNet, GoogleNet, EfficientNet) sotto un comune protocollo di cross-validation stratificata a 5 fold; i modelli monomodali sono combinati in predizioni multimodali tramite Softmax ensembling, e le prestazioni di predizione sono interpretate congiuntamente alla complessità dei modelli. Sulla base del benchmark, GoogleNet viene selezionata come architettura di riferimento per il suo compromesso tra efficienza e prestazioni; un’analisi degli errori ordinali viene successivamente condotta su di essa per caratterizzare la severità delle classificazioni errate, distinguendo i Long Errors, tra classi non adiacenti, dagli Short Errors, tra classi adiacenti. In secondo luogo, due strategie affrontano le limitazioni individuate durante il benchmarking: una loss error-sensitive che estende la weighted cross-entropy con un termine di costo ordinale che pondera gli errori in base alla loro severità, e un ensemble di esperti specifici per modalità che assegna architetture, funzioni di perdita e tecniche di augmentation dedicate alle singole sequenze MRI; le due strategie vengono successivamente combinate in un unico framework di esperti error sensitive. In terzo luogo, Grad-CAM viene impiegato per ispezionare le regioni di attenzione spaziale associate alle predizioni dei modelli. La validazione esterna ha rivelato prestazioni complessive modeste, coerenti con la difficoltà del compito e con il rigoroso protocollo di valutazione adottato. La strategia high-level expert ensemble ha ottenuto il risultato più competitivo, raggiungendo un’accuratezza di 0,61 su BraTS2020 ed entrando nella porzione inferiore dell’intervallo stato dell’arte riportato; mentre la loss error-sensitive ha prodotto una riduzione consistente dei Long Errors in tutte le coorti esterne e combinazioni di modalità. Le prestazioni sono tuttavia diminuite sulle restanti coorti, riflettendo una sostanziale variabilità tra i domini, e Grad-CAM ha mostrato come predizioni corrette non sempre coincidano con un’attenzione focalizzata su regioni plausibilmente significative. Nel complesso, questi risultati confermano la fattibilità della formulazione ordinale proposta e mostrano come la specializzazione per modalità e l’ottimizzazione sensibile alla struttura ordinale del problema rappresentino strategie progettuali efficaci per questo compito. Emerge inoltre che il principale ostacolo all’applicabilità clinica non è la formulazione del problema, quanto piuttosto la difficoltà di ottenere una generalizzazione stabile su coorti eterogenee per protocolli di acquisizione e popolazioni di pazienti.
Multi-cohort benchmark of 3D Deep networks for glioblastoma Overall Survival prediction
Pau, Nicolò
2025/2026
Abstract
Glioblastoma is the most aggressive primary malignant tumour of the central nervous system, with a median Overall Survival (OS) that typically remains limited to approximately 12–15 months despite intensive standard-of-care treatment.The marked heterogeneity of the disease motivates patient-specific risk stratification, for which a noninvasive and reproducible prognostic tool would be clinically valuable. This thesis investigates whether multiparametric brain MRI can support deep learning–based prediction of overall survival (OS) in glioblastoma patients, formulated as a threeclass classification task. Beyond model development, the study aims to evaluate the potential clinical applicability of this approach and to unveil the ability of predictive models to generalize across heterogeneous real-world data. The study relies on four publicly available datasets: UPENN-GBM, the largest one, used for training and internal validation, and BraTS2020, UCSF-PDGM, and RHUH-GBM, reserved exclusively to assess the generalizability of the models across independent and heterogeneous datasets. Each of them provides T1, T1ce, T2, and FLAIR sequences. The methodological narrative develops along three blocks. First, a systematic benchmark compares several 3D CNN architectures adapted from established 2D convolutional models (ResNet, DenseNet, GoogleNet, EfficientNet) under a common stratified 5-fold cross-validation protocol; monomodal models are combined into multimodal predictions through prediction-level Softmax ensembling, and predictive performance is interpreted jointly with model complexity. Based on the benchmark results, GoogleNet is selected as the reference architecture for its efficiency–performance trade-off; an ordinal error analysis is subsequently applied to characterise the severity of its misclassifications, distinguishing Long Errors, between non-adjacent classes, from Short Errors, between adjacent ones. Second, two strategies address the limitations identified during benchmarking: an error-sensitive loss extending weighted cross-entropy with an ordinal cost term that weights errors by their severity, and a modality-specific expert ensemble assigning dedicated architectures, loss functions, and augmentation settings to individual MRI sequences; the two are subsequently combined into a unified errorsensitive expert framework. Third, Grad-CAM is used as an explainability tool to inspect the spatial attention regions associated with the model predictions. External testing revealed modest overall performance, consistent with the difficulty of the task and the strict evaluation protocol adopted. The high-level expert ensemble achieved the most competitive result, reaching an accuracy of 0.61 on BraTS2020. This performance was in line with the reported state-of-the-art range. At the same time, the error-sensitive loss consistently reduced Long Errors across all external cohorts and modality combinations. A decrease in performance was observed on datasets processed through different preprocessing pipelines, highlighting the impact of domain shift. Moreover, Grad-CAM revealed that correct predictions were not always associated with attention patterns focused on meaningful anatomical or tumour-related regions. In conclusion, this thesis tackles two key challenges in deep learning–based survival prediction for glioblastoma: robustness to missing MRI modalities and the formulation of survival estimation as an ordinal clinical task. By adopting a sequenceagnostic approach, the proposed framework remains applicable when different combinations of MRI modalities are available, increasing its potential usability in realworld clinical settings. At the same time, the use of ordinal-aware optimisation reduces the impact of clinically severe errors. While the issue of generalisation across heterogeneous external cohorts is only partially resolved, the results provide evidence that the proposed design choices represent a step towards more robust and adaptable models. In this sense, the thesis contributes to the broader integration of deep learning methods in neuro-oncology, with the aim of supporting more flexible and clinically relevant prognostic tools.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Pau_Tesi.pdf
non accessibile
Descrizione: Testo della tesi
Dimensione
7.75 MB
Formato
Adobe PDF
|
7.75 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Pau_ExecutiveSummary.pdf
non accessibile
Descrizione: Testo executive summary
Dimensione
2.27 MB
Formato
Adobe PDF
|
2.27 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261237