Fourier Light Field Microscopy (FLFM) records volumetric fluorescence in one camera exposure by multiplexing spatial and angular information onto one sensor. With few light-field views available, 3D recovery from this measurement is ill posed. This thesis examines whether the local light-field views and their epipolar structure help a network recover a 3D fluorescence volume from three extracted views, and whether cross-view attention offers an advantage over a convolutional baseline that treats the views as image channels. Three light-field views are extracted from each FLFM measurement and assembled into local multi-view patches. From these, several transformer configurations are built that differ in token formation and attention axis: spatial attention over patch-grid tokens, attention restricted to epipolar-direction neighbors, local attention across the view axis within each patch, together with linear-attention and convolutional-refinement variants. Each model predicts 41 axial (depth) slices and is trained on a common mitochondrial FLFM dataset against a Fourier view-channel-depth (FVCD)-style U-Net baseline. Geometry-guided attention reduces voxel-wise reconstruction error, with the strongest configuration (rope_disp_linear) reaching 36.40 dB validation peak signal-to-noise ratio (PSNR) compared with 34.24 dB for the U-Net baseline. The ablations show that rotary position embedding (RoPE) alone does not explain the gain. The improvement appears when positional encoding is combined with the local multi-view patch construction, showing that how attention is organized over the light field is the useful design choice. The U-Net baseline retains the highest all-slice structural similarity (SSIM, 0.957) and a more favorable speed profile, while several more accurate variants remain memory-bound or rely on region-of-interest (ROI) inference. Forward-model consistency, at its present weighting, does not yet improve supervised training. The results identify attention construction, namely the selected tokens and comparison axis, as the operative variable in limited-view FLFM reconstruction.
La microscopia a campo di luce di Fourier (FLFM) registra la fluorescenza volumetrica con una singola esposizione della camera, multiplexando informazione spaziale e angolare su un unico sensore. Con poche viste disponibili, la ricostruzione 3D risulta mal posta. Questa tesi verifica se viste locali e struttura epipolare aiutino a ricostruire un volume 3D da tre viste estratte, e se l’attenzione incrociata tra viste superi una baseline convoluzionale che usa le viste come canali immagine. Da ogni misura FLFM si estraggono tre viste e si formano patch locali a viste multiple. Sono valutate varianti transformer con diversa formazione dei token e diverso asse di attenzione: griglia spaziale di patch, vicini epipolari, asse delle viste nella patch, attenzione lineare e raffinamento convoluzionale. Tutti i modelli predicono 41 sezioni assiali e sono addestrati sullo stesso dataset FLFM mitocondriale, con confronto rispetto a una baseline U-Net in stile Fourier view-channel-depth (FVCD). L’attenzione guidata dalla geometria riduce l’errore di ricostruzione per voxel. Il modello migliore, rope_disp_linear, raggiunge 36.40 dB di PSNR di validazione contro 34.24 dB della U-Net. Le ablazioni mostrano che RoPE da solo non spiega il guadagno: il miglioramento compare quando la codifica posizionale viene combinata con patch locali a viste multiple. La U-Net mantiene il miglior SSIM su tutte le sezioni (0.957) e una maggiore velocità, mentre varianti più accurate sono limitate dalla memoria o richiedono inferenza su ROI. La coerenza con il modello diretto, con la pesatura attuale, non migliora ancora l’addestramento supervisionato. I risultati indicano la costruzione dell’attenzione, cioè token scelti e asse di confronto, come variabile chiave nella ricostruzione FLFM a viste limitate.
Transformer-based volumetric reconstruction for Fourier light field microscopy
ROJAS CASADIEGO, DAVID FELIPE
2025/2026
Abstract
Fourier Light Field Microscopy (FLFM) records volumetric fluorescence in one camera exposure by multiplexing spatial and angular information onto one sensor. With few light-field views available, 3D recovery from this measurement is ill posed. This thesis examines whether the local light-field views and their epipolar structure help a network recover a 3D fluorescence volume from three extracted views, and whether cross-view attention offers an advantage over a convolutional baseline that treats the views as image channels. Three light-field views are extracted from each FLFM measurement and assembled into local multi-view patches. From these, several transformer configurations are built that differ in token formation and attention axis: spatial attention over patch-grid tokens, attention restricted to epipolar-direction neighbors, local attention across the view axis within each patch, together with linear-attention and convolutional-refinement variants. Each model predicts 41 axial (depth) slices and is trained on a common mitochondrial FLFM dataset against a Fourier view-channel-depth (FVCD)-style U-Net baseline. Geometry-guided attention reduces voxel-wise reconstruction error, with the strongest configuration (rope_disp_linear) reaching 36.40 dB validation peak signal-to-noise ratio (PSNR) compared with 34.24 dB for the U-Net baseline. The ablations show that rotary position embedding (RoPE) alone does not explain the gain. The improvement appears when positional encoding is combined with the local multi-view patch construction, showing that how attention is organized over the light field is the useful design choice. The U-Net baseline retains the highest all-slice structural similarity (SSIM, 0.957) and a more favorable speed profile, while several more accurate variants remain memory-bound or rely on region-of-interest (ROI) inference. Forward-model consistency, at its present weighting, does not yet improve supervised training. The results identify attention construction, namely the selected tokens and comparison axis, as the operative variable in limited-view FLFM reconstruction.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Rojas_Casadiego_Executive_Summary_02.pdf
accessibile in internet per tutti
Descrizione: Executive_Summary
Dimensione
731.62 kB
Formato
Adobe PDF
|
731.62 kB | Adobe PDF | Visualizza/Apri |
|
2026_07_Rojas_Casadiego_Thesis_01.pdf
accessibile in internet per tutti
Descrizione: Thesis
Dimensione
2.46 MB
Formato
Adobe PDF
|
2.46 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261527