Volume electron microscopy resolves the internal architecture of cells at nanometre scale, but turning the resulting gigapixel volumes into quantitative biology requires semantic segmentation of every pixel into its organelle class — a step that remains the field’s central bottleneck. The obstacle is twofold: dense pixel-level annotation of terabyte volumes demands scarce expert time and must be repeated for each new sample, and accurate segmentation depends on relational context spanning multiple spatial scales rather than on local texture alone, since the identity of an organelle is defined as much by its neighbourhood as by its own appearance. These two pressures motivate ε-SegFormer, a sparsely-supervised segmentation framework built on ε-Seg [1] that replaces its hierarchical-VAE backbone with a Vision Transformer encoder and a convolutional decoder. The choice of transformer is deliberate: unlike CNNs, whose receptive field grows only through depth and forces a trade-off between local precision and global context, self-attention operates globally over the entire input at every layer. To preserve local patch fidelity while still supplying wide-field cellular context, the model is fed smaller native-resolution views alongside coarser co-centred crops at wider physical scales — inspired by MuViT [2] — so that the transformer attends jointly over multiple resolution levels in a single forward pass. Sparse supervision is achieved through a joint objective that combines a self-supervised reconstruction loss on every input with a cross-entropy loss applied only at a small number of labelled centre pixels. We evaluate on the BetaSeg dataset of FIB-SEM-imaged mouse pancreatic islet β cells, using fewer than ten thousand labelled coordinates to segment nucleus, granules, mitochondria, and a residual “other” class. The best configuration reaches a mean Dice of 0.8531 on a fully held-out test cell, exceeding ε-Seg by +0.073 at matched label sparsity — evidence that the transformer reformulation, multi-resolution context, and a decreasing mask-ratio schedule together deliver meaningful gains under a fixed, minimal annotation budget.
La microscopia elettronica volumetrica risolve l'architettura interna delle cellule alla scala del nanometro, ma trasformare i volumi gigapixel risultanti in biologia quantitativa richiede la segmentazione semantica di ogni pixel nella sua classe di organello — un passaggio che rimane il principale collo di bottiglia del settore. L'ostacolo è duplice: l'annotazione densa a livello di pixel di volumi in terabyte richiede tempo di esperti scarso e deve essere ripetuta per ogni nuovo campione; inoltre, una segmentazione accurata dipende da un contesto relazionale che abbraccia molteplici scale spaziali, piuttosto che dalla sola texture locale, poiché l'identità di un organello è definita tanto dal suo vicinato quanto dalla propria apparenza. Queste due pressioni motivano ε-SegFormer, un framework di segmentazione con supervisione sparsa costruito su ε-Seg [1], che sostituisce il suo backbone a VAE gerarchico con un encoder Vision Transformer e un decoder convoluzionale. La scelta del transformer è deliberata: a differenza delle CNN, il cui campo recettivo cresce solo attraverso la profondità e impone un compromesso tra precisione locale e contesto globale, la self-attention opera globalmente sull'intero input a ogni livello. Per preservare la fedeltà locale delle patch pur fornendo un ampio contesto cellulare, il modello riceve viste più piccole a risoluzione nativa insieme a crop co-centrati più grossolani a scale fisiche più ampie — ispirandosi a MuViT [2] — così che il transformer attende congiuntamente su più livelli di risoluzione in un singolo passaggio in avanti. La supervisione sparsa è ottenuta tramite un obiettivo congiunto che combina una perdita di ricostruzione auto-supervisionata su ogni input con una perdita di cross-entropia applicata solo a un numero ridotto di pixel centrali etichettati. Valutiamo sul dataset BetaSeg di cellule β delle isole pancreatiche di topo acquisite con FIB-SEM, utilizzando meno di diecimila coordinate etichettate per segmentare nucleo, granuli, mitocondri e una classe residua "altro". La configurazione migliore raggiunge un Dice medio di 0,8531 su una cellula di test completamente separata, superando ε-Seg di +0,073 alla stessa sparsità di etichette — a dimostrazione che la riformulazione basata su transformer, il contesto multi-risoluzione e uno schedule decrescente del mask-ratio producono insieme guadagni significativi con un budget di annotazione fisso e minimo.
Eps-SegFormer: sparsely supervised semantic segmentation of microscopy data
AMIDI, ERFAN
2025/2026
Abstract
Volume electron microscopy resolves the internal architecture of cells at nanometre scale, but turning the resulting gigapixel volumes into quantitative biology requires semantic segmentation of every pixel into its organelle class — a step that remains the field’s central bottleneck. The obstacle is twofold: dense pixel-level annotation of terabyte volumes demands scarce expert time and must be repeated for each new sample, and accurate segmentation depends on relational context spanning multiple spatial scales rather than on local texture alone, since the identity of an organelle is defined as much by its neighbourhood as by its own appearance. These two pressures motivate ε-SegFormer, a sparsely-supervised segmentation framework built on ε-Seg [1] that replaces its hierarchical-VAE backbone with a Vision Transformer encoder and a convolutional decoder. The choice of transformer is deliberate: unlike CNNs, whose receptive field grows only through depth and forces a trade-off between local precision and global context, self-attention operates globally over the entire input at every layer. To preserve local patch fidelity while still supplying wide-field cellular context, the model is fed smaller native-resolution views alongside coarser co-centred crops at wider physical scales — inspired by MuViT [2] — so that the transformer attends jointly over multiple resolution levels in a single forward pass. Sparse supervision is achieved through a joint objective that combines a self-supervised reconstruction loss on every input with a cross-entropy loss applied only at a small number of labelled centre pixels. We evaluate on the BetaSeg dataset of FIB-SEM-imaged mouse pancreatic islet β cells, using fewer than ten thousand labelled coordinates to segment nucleus, granules, mitochondria, and a residual “other” class. The best configuration reaches a mean Dice of 0.8531 on a fully held-out test cell, exceeding ε-Seg by +0.073 at matched label sparsity — evidence that the transformer reformulation, multi-resolution context, and a decreasing mask-ratio schedule together deliver meaningful gains under a fixed, minimal annotation budget.| File | Dimensione | Formato | |
|---|---|---|---|
|
Amidi_Erfan_214166_Thesis.pdf
accessibile in internet per tutti
Descrizione: Thesis
Dimensione
11.51 MB
Formato
Adobe PDF
|
11.51 MB | Adobe PDF | Visualizza/Apri |
|
Amidi_Erfan_214166_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
3.22 MB
Formato
Adobe PDF
|
3.22 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260242