Recent feed-forward 3D Gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single feed-forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose-free multi-view methods address only static scenes; and per-scene optimization methods bridge some of these gaps but at minutes-to-hours cost per scene. This thesis introduces NoPo4D, the first feed-forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per-pixel image-plane shifts and depth changes, allowing direct supervision from pseudo ground-truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose-free methods require. The system is rounded out by a bidirectional motion encoder for cross-view and cross-frame feature aggregation, and view-dependent opacity that mitigates cross-view and cross-timestep Gaussian misalignments. On four multi-view dynamic benchmarks, NoPo4D consistently outperforms prior feed-forward baselines, and with an optional post-optimization stage surpasses per-scene optimization methods, while running orders of magnitude faster.
I recenti metodi feed-forward per il 3D Gaussian Splatting hanno compiuto progressi notevoli su singoli aspetti della ricostruzione di scene tridimensionali, ma nessun metodo esistente affronta congiuntamente contenuti dinamici, input multi-vista e pose delle telecamere sconosciute in un unico passaggio feed-forward. I metodi che gestiscono la dinamica richiedono pose accurate delle telecamere o accettano solo input monoculare; i metodi multi-vista senza pose affrontano unicamente scene statiche; i metodi di ottimizzazione per-scena colmano alcune di queste lacune, ma con un costo da minuti a ore per scena. Questa tesi introduce NoPo4D, il primo sistema feed-forward che affronta questo quadrante inesplorato. Basandosi su un backbone geometrico pre-addestrato e sui recenti framework 4D Gaussian, NoPo4D introduce una scomposizione della velocità che separa il moto delle Gaussiane in spostamenti nel piano immagine per pixel e variazioni di profondità, consentendo la supervisione diretta della componente 2D tramite flusso ottico pseudo-ground-truth. Questo approccio evita sia la dipendenza dal rendering differenziabile sia la necessità di annotazioni di moto 3D. Il sistema è completato da un encoder di moto bidirezionale per l'aggregazione di feature cross-vista e cross-frame, e da un'opacità dipendente dalla vista che mitiga i disallineamenti delle Gaussiane tra viste e istanti temporali diversi. Su quattro benchmark multi-vista dinamici, NoPo4D supera costantemente i baseline feed-forward precedenti e, con una fase opzionale di post-ottimizzazione, raggiunge o supera i metodi per-scena pur essendo significativamente più veloce.
No pose, no problem in 4D: feed-forward dynamic Gaussians from unposed multi-view videos
Balice, Matteo
2025/2026
Abstract
Recent feed-forward 3D Gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single feed-forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose-free multi-view methods address only static scenes; and per-scene optimization methods bridge some of these gaps but at minutes-to-hours cost per scene. This thesis introduces NoPo4D, the first feed-forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per-pixel image-plane shifts and depth changes, allowing direct supervision from pseudo ground-truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose-free methods require. The system is rounded out by a bidirectional motion encoder for cross-view and cross-frame feature aggregation, and view-dependent opacity that mitigates cross-view and cross-timestep Gaussian misalignments. On four multi-view dynamic benchmarks, NoPo4D consistently outperforms prior feed-forward baselines, and with an optional post-optimization stage surpasses per-scene optimization methods, while running orders of magnitude faster.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Balice_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive summary
Dimensione
925.63 kB
Formato
Adobe PDF
|
925.63 kB | Adobe PDF | Visualizza/Apri |
|
2026_07_Balice_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
4.21 MB
Formato
Adobe PDF
|
4.21 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/258157