Modern robotics is increasingly based on blending state-of-the-art foundation models with edge architectures and real-time understanding pipelines. People interacting with robots may expect the embodied agent to provide information about locations, events, or objects, which requires the agent to return precise answers on context-dependent information, within human-like inference times. Allowing embodied agents to operate in dynamic and heterogeneous environments requires a generalization of the mechanisms for perception and semantic scene understanding, and a consequent memory structure tailored for long time horizons. This prospect is still unfulfilled in modern literature, tailored for constrained deployment conditions, and failing at generalizing such tasks. While prior work has addressed dense spatial memory-building, little work has optimized memory representations to enable agents based on a Large Language Model (LLM) to quickly build, traverse, and efficiently retrieve information in real-time robot scenarios. To address this gap, we propose a hierarchical graph representation and memory-building framework for LLM-driven robotic agents. This approach is encompassed by two architectural steps that respectively address entity information redundancy in Retrieval-Augmented Generation (RAG) agents, and propose a novel memory architecture extending the use of semantic graphs at a hierarchical level, splitting information into different spatial-semantic tiers, while allowing for spatial grounding of atomic entities. The former is evaluated on the NaVQA dataset, showing how structuring semantic information retains real-time performance in inference and querying times, while showing competitive accuracy on the global task relative to the baseline; the latter is validated through a direct comparison of memory building and querying phases in a simulation setting, addressing the memory structure limitations of the former. We also test the core agentic pipeline on a physical robot, showing practical utility in real-world contexts through human-robot interaction, while running the visual-language model and the architecture locally.
La robotica moderna è sempre più basata sull'integrazione di modelli fondazionali all'avanguardia con architetture edge e comprensione in tempo reale. Le persone che interagiscono con i robot si aspettano che l'agente in essi fornisca informazioni su luoghi, eventi o oggetti, il che richiede di restituire risposte precise su informazioni relative al contesto in tempo reale. Consentire agli agenti embodied di operare in ambienti dinamici ed eterogenei richiede una generalizzazione dei meccanismi di percezione e di comprensione semantica delle scene visive, e una conseguente struttura di memoria concepita per applicazioni di lungo periodo. Questa prospettiva è spesso ingegnerizzata per ambienti e contesti altamente vincolati e strutturali, fallendo nel generalizzare tali compiti. Sebbene lavori di ricerca precedenti abbiano affrontato la costruzione di memorie spaziali, pochi studi hanno ottimizzato le rappresentazioni di memoria per consentire ad agenti basati su modelli di linguaggio di grandi dimensioni (LLM) di costruire e recuperare efficientemente le informazioni in scenari robotici in tempo reale. Questo lavoro propone una memoria a grafo gerarchico e un framework di costruzione della stessa per agenti robotici controllati da LLM. Il lavoro si articola in due step architetturali che rispettivamente affrontano la ridondanza delle informazioni sulle entità negli agenti basati su Retrieval-Augmented Generation (RAG), e propongono una nuova architettura di memoria che estende l'uso dei grafi semantici a livello gerarchico, suddividendo le informazioni in diversi livelli spaziali-semantici. Il primo step è valutato sul dataset NaVQA, mostrando come strutturare le informazioni semantiche mantenga prestazioni in contesti in tempo reale nei tempi di interrogazione dell'agente, pur dimostrando un'accuratezza competitiva rispetto alla baseline; il secondo è validato attraverso un confronto diretto delle fasi di costruzione e interrogazione della memoria in un ambiente di simulazione, risolvendo i limiti della struttura di memoria del primo. La pipeline agentica di base è inoltre testata su un robot reale, mostrandone l'utilità pratica in contesti fisici.
Language-driven robot spatial understanding with long-horizon semantic graph memory
Riva, Paolo
2025/2026
Abstract
Modern robotics is increasingly based on blending state-of-the-art foundation models with edge architectures and real-time understanding pipelines. People interacting with robots may expect the embodied agent to provide information about locations, events, or objects, which requires the agent to return precise answers on context-dependent information, within human-like inference times. Allowing embodied agents to operate in dynamic and heterogeneous environments requires a generalization of the mechanisms for perception and semantic scene understanding, and a consequent memory structure tailored for long time horizons. This prospect is still unfulfilled in modern literature, tailored for constrained deployment conditions, and failing at generalizing such tasks. While prior work has addressed dense spatial memory-building, little work has optimized memory representations to enable agents based on a Large Language Model (LLM) to quickly build, traverse, and efficiently retrieve information in real-time robot scenarios. To address this gap, we propose a hierarchical graph representation and memory-building framework for LLM-driven robotic agents. This approach is encompassed by two architectural steps that respectively address entity information redundancy in Retrieval-Augmented Generation (RAG) agents, and propose a novel memory architecture extending the use of semantic graphs at a hierarchical level, splitting information into different spatial-semantic tiers, while allowing for spatial grounding of atomic entities. The former is evaluated on the NaVQA dataset, showing how structuring semantic information retains real-time performance in inference and querying times, while showing competitive accuracy on the global task relative to the baseline; the latter is validated through a direct comparison of memory building and querying phases in a simulation setting, addressing the memory structure limitations of the former. We also test the core agentic pipeline on a physical robot, showing practical utility in real-world contexts through human-robot interaction, while running the visual-language model and the architecture locally.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Riva_Tesi.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Testo della Tesi in formato Journal Paper
Dimensione
39.89 MB
Formato
Adobe PDF
|
39.89 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Riva_ExecutiveSummary.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Executive Summary della Tesi
Dimensione
2.38 MB
Formato
Adobe PDF
|
2.38 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261268