Colorectal cancer (CRC) remains a leading cause of cancer-related mortality, and its incidence is increasing among individuals under 50, motivating the identification of biomarkers that are accurate, scalable, and compatible with routine clinical workflows. Hematoxylin and eosin (H&E) whole-slide images (WSIs) are widely available and retrospective, but learning from WSIs is computationally demanding and weakly supervised due to gigapixel resolution, substantial heterogeneity, and the lack of dense annotations. This thesis investigates whether routine H&E WSIs contain sufficient morphological signal to predict a neutrophil-related label derived from immunophenotyping: High Neutrophil (HN) if the percentage of tumor-associated neutrophils expressing high levels of CD15 (CD15high) exceeds 50, and Low Neutrophils (LN) otherwise. We developed a modular computational pathology pipeline that indexes slides through tissue-aware tiling and quality control, extracts tile embeddings using self-supervised foundation models pretrained on histopathology, and aggregates tile level evidence into patient-level predictions via Multiple Instance Learning (MIL), with transformer-based aggregation as the final configuration. We then analyzed model failure modes such as inter-fold logit shift and bag-sampling variance that can produce optimistic pooled out- of-fold rankings yet brittle single-patient inference. To mitigate these effects, we adopted deterministic multi-bag validation and stability-oriented regularization during training and model selection. Attention-derived heatmaps are used as an engineering tool for qualitative inspection and confounder detection. Finally, after training on an internal single-center cohort from the European Institute of Oncology (IEO) biobank, we applied the model in inference mode to an independent external public cohort, namely The Cancer Genome Atlas (TCGA) Colon Adenocarcinoma (COAD) and Rectal Adenocarcinoma (READ) projects. Using the disease-free survival (DFS) annotations available for this external cohort, we have then assessed whether the model’s predicted HN-risk score is associated with prognosis.
Il cancro colorettale (CRC) rimane una delle principali cause di mortalità oncologica e la sua incidenza è in aumento tra gli individui al di sotto dei 50 anni, motivando l’identificazione di biomarcatori accurati, scalabili e compatibili con i flussi di lavoro clinici di routine. Le immagini whole-slide (WSI) colorate con ematossilina ed eosina (H&E) sono ampiamente disponibili anche in retrospettiva, ma l’apprendimento dalle WSI è oneroso e debolmente supervisionato per la risoluzione gigapixel, l’eterogeneità e la mancanza di annotazioni dense. Questa tesi indaga se le WSI H&E di routine contengano un segnale morfologico sufficiente a predire un’etichetta correlata ai neutrofili derivata dall’immunofenotipizzazione: Neutrofili Alti (HN) se la percentuale di neutrofili associati al tumore con alti livelli di CD15 (CD15high) supera il 50, e Neutrofili Bassi (LN) altrimenti. Abbiamo sviluppato una pipeline di patologia computazionale modulare che indicizza i vetrini tramite tassellatura basata sul tessuto, estrae embeddings con modelli di fondazione auto-supervisionati pre-addestrati su istopatologia e aggrega le evidenze a livello di tasselli in predizioni a livello di paziente tramite Apprendimento a Istanze Multiple (MIL). Abbiamo inoltre analizzato modalità di fallimento come lo shift dei logit tra i fold e la varianza di campionamento delle bag, che possono rendere ottimistica la validazione aggregata e fragile l’inferenza sul singolo paziente; per mitigare, adottiamo validazione deterministica multi-bag e regolarizzazione orientata alla stabilità durante addestramento e selezione del modello. Le mappe di calore derivate dall’attenzione sono utilizzate per l’ispezione qualitativa e la rilevazione di fattori confondenti. Infine, dopo l’addestramento su una coorte interna monocentrica della biobanca dell’Istituto Europeo di Oncologia (IEO), applichiamo il modello in inferenza a una coorte pubblica esterna indipendente (TCGA COAD/READ) e, usando le annotazioni di sopravvivenza libera da malattia (DFS), valutiamo se il punteggio di rischio HN predetto sia associato alla prognosi.
Weakly supervised multiple instance learning for prediction of neutrophil phenotypes in Colorectal Cancer
Raganato, Nicolò
2024/2025
Abstract
Colorectal cancer (CRC) remains a leading cause of cancer-related mortality, and its incidence is increasing among individuals under 50, motivating the identification of biomarkers that are accurate, scalable, and compatible with routine clinical workflows. Hematoxylin and eosin (H&E) whole-slide images (WSIs) are widely available and retrospective, but learning from WSIs is computationally demanding and weakly supervised due to gigapixel resolution, substantial heterogeneity, and the lack of dense annotations. This thesis investigates whether routine H&E WSIs contain sufficient morphological signal to predict a neutrophil-related label derived from immunophenotyping: High Neutrophil (HN) if the percentage of tumor-associated neutrophils expressing high levels of CD15 (CD15high) exceeds 50, and Low Neutrophils (LN) otherwise. We developed a modular computational pathology pipeline that indexes slides through tissue-aware tiling and quality control, extracts tile embeddings using self-supervised foundation models pretrained on histopathology, and aggregates tile level evidence into patient-level predictions via Multiple Instance Learning (MIL), with transformer-based aggregation as the final configuration. We then analyzed model failure modes such as inter-fold logit shift and bag-sampling variance that can produce optimistic pooled out- of-fold rankings yet brittle single-patient inference. To mitigate these effects, we adopted deterministic multi-bag validation and stability-oriented regularization during training and model selection. Attention-derived heatmaps are used as an engineering tool for qualitative inspection and confounder detection. Finally, after training on an internal single-center cohort from the European Institute of Oncology (IEO) biobank, we applied the model in inference mode to an independent external public cohort, namely The Cancer Genome Atlas (TCGA) Colon Adenocarcinoma (COAD) and Rectal Adenocarcinoma (READ) projects. Using the disease-free survival (DFS) annotations available for this external cohort, we have then assessed whether the model’s predicted HN-risk score is associated with prognosis.| File | Dimensione | Formato | |
|---|---|---|---|
|
Thesis_Nicolo_Raganato.pdf
non accessibile
Descrizione: Thesis
Dimensione
20.97 MB
Formato
Adobe PDF
|
20.97 MB | Adobe PDF | Visualizza/Apri |
|
Executive_Summary_Nicolo_Raganato.pdf
non accessibile
Descrizione: Executive Summary
Dimensione
2.33 MB
Formato
Adobe PDF
|
2.33 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/251726