This thesis explores the use of lightweight Vision-Language Models (VLMs) to generate behavior trees for robots from a visual observation of the scene and a natural language instruction. Motivated by recent advances in vision-language models, we investigate whether visual perception of the scene can be exploited to generate robotic plans that account for the observed state of the environment. While recent works have explored vision-language models for robotic task planning, none has targeted the generation of behavior trees from compact, open-source multimodal models. Since no existing dataset links visual observations and instructions to executable behavior trees, we propose a method to construct one starting from existing robotic datasets. A large-scale model serves as a teacher in a multi-stage generation pipeline guided by prompt engineering. The resulting dataset is used to fine-tune VLMs ranging from 500\,M to 4\,B parameters via Parameter-Efficient Fine-Tuning (PEFT). We compare multiple VLM architectures, both fine-tuned and in their base form, against significantly larger closed-source models. Our results show that this approach effectively closes the performance gap while requiring only a fraction of the computational resources. The quality of the generated behavior trees, compatible with BehaviorTree.CPP, is assessed both offline, through structural, action-level, and lexical-overlap metrics, and online, through execution of household tasks in a state-of-the-art embodied simulator.
Questa tesi esplora l'impiego di Vision-Language Model (VLM) leggeri per generare behavior tree destinati a robot, a partire da un'osservazione visiva della scena e da un'istruzione in linguaggio naturale. Alla luce dei recenti progressi nei modelli vision-language, si indaga se la percezione visiva della scena possa essere sfruttata per produrre piani robotici che tengano conto dello stato osservato dell'ambiente. Sebbene diversi lavori recenti abbiano impiegato modelli vision-language per la pianificazione robotica, nessuno ha affrontato la generazione di behavior tree a partire da modelli multimodali compatti e open-source. Poiché non esiste un dataset che colleghi osservazioni visive e istruzioni a behavior tree eseguibili, viene proposto un metodo per costruirne uno a partire da dataset robotici esistenti. Un modello di grandi dimensioni svolge il ruolo di insegnante in una pipeline di generazione multi-stadio guidata da prompt engineering. Il dataset risultante viene utilizzato per il fine-tuning di VLM con un numero di parametri compreso tra 500 milioni e 4 miliardi, tramite Parameter-Efficient Fine-Tuning (PEFT). Vengono confrontate più architetture VLM, sia nella versione base sia dopo fine-tuning, con modelli closed-source di scala molto superiore. I risultati mostrano che l'approccio proposto colma efficacemente il divario prestazionale richiedendo solo una frazione delle risorse computazionali. La qualità dei behavior tree generati, compatibili con BehaviorTree.CPP, è valutata sia offline, mediante metriche strutturali, di correttezza delle azioni e di sovrapposizione lessicale, sia online, tramite l'esecuzione di attività domestiche in un simulatore embodied di riferimento.
Multimodal behavior tree generation: a lightweight vision-language approach to robotic task planning
Battistini, Cristiano
2025/2026
Abstract
This thesis explores the use of lightweight Vision-Language Models (VLMs) to generate behavior trees for robots from a visual observation of the scene and a natural language instruction. Motivated by recent advances in vision-language models, we investigate whether visual perception of the scene can be exploited to generate robotic plans that account for the observed state of the environment. While recent works have explored vision-language models for robotic task planning, none has targeted the generation of behavior trees from compact, open-source multimodal models. Since no existing dataset links visual observations and instructions to executable behavior trees, we propose a method to construct one starting from existing robotic datasets. A large-scale model serves as a teacher in a multi-stage generation pipeline guided by prompt engineering. The resulting dataset is used to fine-tune VLMs ranging from 500\,M to 4\,B parameters via Parameter-Efficient Fine-Tuning (PEFT). We compare multiple VLM architectures, both fine-tuned and in their base form, against significantly larger closed-source models. Our results show that this approach effectively closes the performance gap while requiring only a fraction of the computational resources. The quality of the generated behavior trees, compatible with BehaviorTree.CPP, is assessed both offline, through structural, action-level, and lexical-overlap metrics, and online, through execution of household tasks in a state-of-the-art embodied simulator.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Battistini_Tesi.pdf
accessibile in internet per tutti
Descrizione: Testo della tesi
Dimensione
6.88 MB
Formato
Adobe PDF
|
6.88 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Battistini_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
1.25 MB
Formato
Adobe PDF
|
1.25 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/252766