Deep learning models for medical image segmentation must provide not only accurate predictions but also calibrated confidence and clinically meaningful uncertainty estimates. In full-body CT segmentation of small and low-contrast structures such as the lymphatic system, predictive probabilities are often miscalibrated and voxel-wise uncertainty maps are dominated by trivial boundary effects. This thesis investigates calibration and uncertainty estimation within the nnU-Net v2 framework on 45 annotated CT volumes. Three configurations are compared: a single baseline model, a deep ensemble (5-fold cross-validation), and a checkpoint ensemble based on cyclical learning rates. Confidence calibration is performed via post-hoc temperature scaling, with the temperature parameter estimated by minimizing negative log-likelihood (NLL) on a held-out region of interest. On the test set, temperature scaling consistently reduces miscalibration across all models, with relative improvements up to 41% in NLL and 56% in Expected Calibration Error, without altering segmentation decisions. Voxel-wise uncertainty maps derived from single or ensemble predictions (entropy, variance, mutual information, disagreement and anti-confidence metrics) are quantitatively evaluated as error localization tools through threshold calibration based on Dice overlap with segmentation errors. Within a 15 mm ROI, ensemble-based measures achieve the highest overall error recall (recall_FP ∼ 0.81 and recall_total ∼ 0.64), with F1-scores up to 0.44. When excluding a 2 mm boundary band, standard maps show a recall reduction of up to 37%, whereas distance-aware entropy reaches recall_total ∼ 0.56, outperforming conventional measures in non-boundary regions, especially Distance Expected Uncertainty (DEU). Exceedance-Based Contextual Uncertainty (EBCU) further suppresses trivial boundary uncertainty and increases relative recall in the no-border setting by more than 40% compared to the standard ROI. An interactive visualization tool integrates calibrated predictions and thresholded uncertainty maps to support structured qualitative inspection. Overall, the results show that post-hoc calibration improves the quantitative alignment between uncertainty and actual segmentation errors. Voxel-wise uncertainty maps can be used as error detection tool enhancing clinical interpretability without architectural modifications or substantial computational overhead.
I modelli di deep learning per la segmentazione di immagini mediche devono fornire non solo predizioni accurate, ma anche stime di confidenza calibrate e misure di incertezza clinicamente significative. Nella segmentazione full-body in TC di strutture piccole e a basso contrasto, come il sistema linfatico, le probabilità predittive risultano spesso mal calibrate e le mappe di incertezza voxel-wise sono dominate da effetti banali di bordo. Questa tesi indaga la calibrazione e la stima dell’incertezza all’interno del framework nnU-Net v2 su 45 volumi TC annotati. Vengono confrontate tre configurazioni: un modello base singolo, un deep ensemble (5-fold cross-validation) e un checkpoint ensemble basato su learning rate ciclici. La calibrazione della confidenza è effettuata tramite temperature scaling post-hoc, con il parametro di temperatura stimato minimizzando la negative log-likelihood (NLL) su una regione di interesse separata. Sul test set, il temperature scaling riduce in modo consistente la miscalibrazione in tutti i modelli, con miglioramenti relativi fino al 41% in termini di NLL e fino al 56% in termini di Expected Calibration Error, senza alterare le decisioni di segmentazione. Le mappe di incertezza voxel-wise derivate da predizioni singole o ensemble (entropia, varianza, mutual information, disaccordo e metriche di anti-confidence) sono valutate quantitativamente come strumenti di localizzazione dell’errore attraverso una calibrazione della soglia basata sull’overlap di Dice con gli errori di segmentazione. All’interno di una ROI di 15 mm, le misure basate su ensemble raggiungono il valore più elevato di richiamo complessivo dell’errore (recall_FP circa 0,81 e recall_total circa 0,64), con F1-score fino a 0,44. Escludendo una banda di 2 mm dal bordo, le mappe standard mostrano una riduzione del recall fino al 37%, mentre l’entropia distance-aware raggiunge un recall_total di circa 0,56, superando le misure convenzionali nelle regioni non di bordo, in particolare la Distance Expected Uncertainty (DEU). La Exceedance-Based Contextual Uncertainty (EBCU) riduce ulteriormente l’incertezza banale di bordo e aumenta il richiamo relativo nel setting senza bordo di oltre il 40% rispetto alla ROI standard. Uno strumento di visualizzazione interattivo integra predizioni calibrate e mappe di incertezza sogliate per supportare un’ispezione qualitativa strutturata. Nel complesso, i risultati mostrano che la calibrazione post-hoc migliora l’allineamento quantitativo tra incertezza e reali errori di segmentazione. Le mappe di incertezza voxel-wise possono essere utilizzate come strumenti di rilevazione dell’errore, migliorando l’interpretabilità clinica senza modifiche architetturali o un significativo aumento del costo computazionale.
Calibration and uncertainty visualization in medical image segmentation
REITANI, LORENZO
2025/2026
Abstract
Deep learning models for medical image segmentation must provide not only accurate predictions but also calibrated confidence and clinically meaningful uncertainty estimates. In full-body CT segmentation of small and low-contrast structures such as the lymphatic system, predictive probabilities are often miscalibrated and voxel-wise uncertainty maps are dominated by trivial boundary effects. This thesis investigates calibration and uncertainty estimation within the nnU-Net v2 framework on 45 annotated CT volumes. Three configurations are compared: a single baseline model, a deep ensemble (5-fold cross-validation), and a checkpoint ensemble based on cyclical learning rates. Confidence calibration is performed via post-hoc temperature scaling, with the temperature parameter estimated by minimizing negative log-likelihood (NLL) on a held-out region of interest. On the test set, temperature scaling consistently reduces miscalibration across all models, with relative improvements up to 41% in NLL and 56% in Expected Calibration Error, without altering segmentation decisions. Voxel-wise uncertainty maps derived from single or ensemble predictions (entropy, variance, mutual information, disagreement and anti-confidence metrics) are quantitatively evaluated as error localization tools through threshold calibration based on Dice overlap with segmentation errors. Within a 15 mm ROI, ensemble-based measures achieve the highest overall error recall (recall_FP ∼ 0.81 and recall_total ∼ 0.64), with F1-scores up to 0.44. When excluding a 2 mm boundary band, standard maps show a recall reduction of up to 37%, whereas distance-aware entropy reaches recall_total ∼ 0.56, outperforming conventional measures in non-boundary regions, especially Distance Expected Uncertainty (DEU). Exceedance-Based Contextual Uncertainty (EBCU) further suppresses trivial boundary uncertainty and increases relative recall in the no-border setting by more than 40% compared to the standard ROI. An interactive visualization tool integrates calibrated predictions and thresholded uncertainty maps to support structured qualitative inspection. Overall, the results show that post-hoc calibration improves the quantitative alignment between uncertainty and actual segmentation errors. Voxel-wise uncertainty maps can be used as error detection tool enhancing clinical interpretability without architectural modifications or substantial computational overhead.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Reitani_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
1.92 MB
Formato
Adobe PDF
|
1.92 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Reitani.pdf
accessibile in internet per tutti
Descrizione: Tesi Magistrale
Dimensione
15.42 MB
Formato
Adobe PDF
|
15.42 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253581