Object segmentation is a foundational task in computer vision, yet achieving high accuracy with intuitive, minimal user input remains a significant challenge. Noun-phrase-based segmentation, while user-friendly, is frequently undermined by both visual and semantic ambiguity. This thesis investigates methods for enhancing SAM3's native zero-shot capabilities while also exploring the feasibility of its adaptation to few-shot segmentation, aiming to resolve these ambiguities through a combination of language-based guidance and visual exemplars within a strictly training-free framework. We evaluate SAM3’s generalization capabilities and its performance under domain shift, exploring whether few-shot adaptation can improve its robustness to varying data distributions and synonym-based prompt phrasing. Our methodology employs a two-step prompting strategy, systematically integrating text labels with sampled points or bounding boxes derived from support images. Our experimental results reveal an unexpected performance dynamic: while semantic information (specifically well-chosen text labels) is highly effective, the integration of visual exemplars paradoxically degraded performance in our tests. Furthermore, we demonstrate that vanilla SAM3 exhibits extreme sensitivity to synonymous phrasing, a critical limitation we mitigate through our proposed prompting strategies. We observe modest but meaningful improvements in addressing domain shift, with consistent metric improvements at both the synset and class levels. Ultimately, this work suggests that while adapting large foundation models to few-shot tasks is feasible, it remains sensitive to the alignment between a model's original training assumptions and the constraints of the deployment environment. While our training-free experiments found that semantic information currently provides more stable guidance than visual exemplars, this does not preclude the effectiveness of visual grounding. Rather, it highlights that without specialized architectural components to bridge the gap between support and query features, the model’s internal priors may conflict with external visual inputs.
La segmentazione di oggetti è un compito fondamentale nella computer vision, tuttavia raggiungere un'elevata accuratezza con un input minimo e intuitivo rimane una sfida significativa. La segmentazione basata su parole chiave o piccole frasi, sebbene facile da usare, è frequentemente compromessa da ambiguità sia visiva che semantica. Questa tesi indaga metodi per potenziare le capacità zero-shot native di SAM3 ed esplora al contempo la fattibilità di un suo adattamento alla few-shot segmentation, con l'obiettivo di risolvere queste ambiguità attraverso una combinazione di guide basate su linguaggio ed esempi visivi all'interno di un framework rigorosamente training-free. Valutiamo le capacità di generalizzazione di SAM3 e la sua performance sotto domain shift, esplorando se l'adattamento few-shot possa migliorare la sua robustezza a distribuzioni dati variabili e formulazioni di prompt basate su sinonimi. La nostra metodologia impiega una strategia di prompting a due step, integrando sistematicamente label testuali con punti campionati o bounding box derivati dalle immagini di supporto. I nostri risultati sperimentali rivelano una dinamica di performance inaspettata: mentre l'informazione semantica (specificamente label testuali ben scelte) si dimostra altamente efficace, l'integrazione di esempi visivi esterni ha paradossalmente ridotto le performance nei nostri test. Inoltre, dimostriamo che la versione originale di SAM3 esibisce una sensibilità estrema all'uso di sinonimi, una limitazione critica che mitighiamo attraverso le nostre strategie di prompting proposte. Osserviamo miglioramenti modesti ma significativi nell'affrontare il domain shift, con miglioramenti consistenti delle metriche sia a livello di concetto generale che di classe specifica. In conclusione, questo lavoro suggerisce che mentre l'adatammento di large foundation models a compiti few-shot è fattibile, rimane sensibile all'allineamento tra le assunzioni originali di training del modello e i vincoli dell'ambiente di deployment. Sebbene i nostri esperimenti training-free abbiano rivelato che l'informazione semantica fornisce attualmente una guidance più stabile rispetto ad esempi visivi esterni, ciò non preclude l'efficacia di diversi metodi che invece ne fanno uso. Piuttosto, evidenzia che senza componenti architetturali specializzate per colmare il gap tra feature di supporto e query, le priors interne del modello possono entrare in conflitto con input visivi esterni.
Enhancing SAM3 robustness against semantic and domain shift via dual-prompting
CAVICCHIOLI, MICHELE
2025/2026
Abstract
Object segmentation is a foundational task in computer vision, yet achieving high accuracy with intuitive, minimal user input remains a significant challenge. Noun-phrase-based segmentation, while user-friendly, is frequently undermined by both visual and semantic ambiguity. This thesis investigates methods for enhancing SAM3's native zero-shot capabilities while also exploring the feasibility of its adaptation to few-shot segmentation, aiming to resolve these ambiguities through a combination of language-based guidance and visual exemplars within a strictly training-free framework. We evaluate SAM3’s generalization capabilities and its performance under domain shift, exploring whether few-shot adaptation can improve its robustness to varying data distributions and synonym-based prompt phrasing. Our methodology employs a two-step prompting strategy, systematically integrating text labels with sampled points or bounding boxes derived from support images. Our experimental results reveal an unexpected performance dynamic: while semantic information (specifically well-chosen text labels) is highly effective, the integration of visual exemplars paradoxically degraded performance in our tests. Furthermore, we demonstrate that vanilla SAM3 exhibits extreme sensitivity to synonymous phrasing, a critical limitation we mitigate through our proposed prompting strategies. We observe modest but meaningful improvements in addressing domain shift, with consistent metric improvements at both the synset and class levels. Ultimately, this work suggests that while adapting large foundation models to few-shot tasks is feasible, it remains sensitive to the alignment between a model's original training assumptions and the constraints of the deployment environment. While our training-free experiments found that semantic information currently provides more stable guidance than visual exemplars, this does not preclude the effectiveness of visual grounding. Rather, it highlights that without specialized architectural components to bridge the gap between support and query features, the model’s internal priors may conflict with external visual inputs.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Cavicchioli_Tesi.pdf
accessibile in internet per tutti
Descrizione: Testo della tesi
Dimensione
34.59 MB
Formato
Adobe PDF
|
34.59 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Cavicchioli_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Testo dell'executive summary
Dimensione
4.44 MB
Formato
Adobe PDF
|
4.44 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260630