The transition to autonomous driving is fundamentally reshaping modern transportation, yet ensuring the safety of Vulnerable Road Users (VRUs) remains a critical bottleneck. Predicting pedestrian behavior is exceptionally challenging due to their ability to suddenly change speed and direction. While human drivers intuitively interpret subtle gestures and scene interactions to gauge intent, current autonomous systems still struggle to replicate this level of understanding. To address this limitation and enable safer, smoother automated maneuvers, this thesis presents the design and implementation of a map-aware pedestrian intention prediction algorithm. The proposed approach leverages a hybrid architecture that combines deterministic feature engineering with deep learning. Specifically, a carefully selected set of kinematic and map-aware spatial features is manually extracted for each pedestrian in the scene to capture their relationship with the surrounding road topology. These explicit feature vectors are subsequently fed into a Stacked Long Short-Term Memory (LSTM) network, which calculates the probability of a pedestrian crossing within a predefined future time-horizon. By relying on hand-crafted feature extraction prior to the neural network phase, this methodology guarantees a significantly higher degree of interpretability compared to traditional, fully "black-box" end-to-end deep learning models. A key contribution of this work is the model’s flexibility and efficiency, enabling it to handle varying numbers of pedestrians and track lengths within a computationally efficient architecture. Two variants are considered: a Full Model using the complete feature set, and a Reduced Model retaining the most relevant features identified through ablation and tailored to the proprietary dataset, where map information is less reliable. The model is trained on a custom-annotated version of the Argoverse 2 dataset and validated on a proprietary real-world dataset, demonstrating the robustness and practical applicability of the proposed approach.
La transizione verso la guida autonoma sta ridefinendo radicalmente il trasporto moderno, tuttavia garantire la sicurezza degli utenti vulnerabili della strada rimane un nodo critico. Prevedere il comportamento dei pedoni è estremamente complesso a causa della loro capacità di cambiare repentinamente velocità e direzione. Mentre i conducenti umani interpretano intuitivamente gesti sottili e interazioni ambientali per stimarne l'intento, gli attuali sistemi autonomi faticano ancora a replicare tale capacità di anticipazione. Per superare questo limite e consentire manovre automatizzate più sicure e fluide, questa tesi presenta la progettazione e l'implementazione di un algoritmo di previsione dell'intenzione dei pedoni basato sulla conoscenza della mappa (map-aware). L'approccio proposto sfrutta un'architettura ibrida che combina il feature extraction deterministico con il deep learning. Nello specifico, per ogni pedone nella scena viene estratto manualmente un set accuratamente selezionato di caratteristiche cinematiche e spaziali legate alla mappa, al fine di catturare la loro relazione con la topologia stradale circostante. Questi vettori di caratteristiche esplicite vengono successivamente forniti in input a una rete Stacked Long Short-Term Memory (LSTM), che calcola la probabilità di attraversamento del pedone entro un orizzonte temporale futuro predefinito. Affidandosi a un'estrazione delle caratteristiche eseguita prima della fase di rete neurale, questa metodologia garantisce un grado di interpretabilità significativamente superiore rispetto ai tradizionali modelli end-to-end completamente "black-box". %Un contributo chiave di questo lavoro risiede nella flessibilità ed efficienza del modello: esso gestisce agevolmente un numero variabile di pedoni nella scena, si adatta a diverse lunghezze delle tracce temporali e si basa su un'architettura di rete neurale leggera. Il modello è stato addestrato sul dataset Argoverse 2 Motion Forecasting, appositamente annotato per la previsione dell'intenzione dei pedoni, e le sue capacità di generalizzazione sono state validate su un dataset proprietario reale. Un contributo chiave di questo lavoro è la flessibilità ed efficienza del modello, che permette di gestire un numero variabile di pedoni e lunghezze delle traiettorie in modo computazionalmente efficiente. Sono considerate due varianti: un modello completo, che utilizza tutte le feature, e uno ridotto, che mantiene solo quelle più rilevanti identificate tramite analisi di ablazione e adattate al dataset proprietario, dove le informazioni di mappa sono meno affidabili. Il modello è addestrato su una versione annotata di Argoverse 2 e validato su un dataset proprietario reale, dimostrandone robustezza e applicabilità.
From motion cues to intent: predicting pedestrian crossing probability with stacked LSTM
Dametto, Fabio
2024/2025
Abstract
The transition to autonomous driving is fundamentally reshaping modern transportation, yet ensuring the safety of Vulnerable Road Users (VRUs) remains a critical bottleneck. Predicting pedestrian behavior is exceptionally challenging due to their ability to suddenly change speed and direction. While human drivers intuitively interpret subtle gestures and scene interactions to gauge intent, current autonomous systems still struggle to replicate this level of understanding. To address this limitation and enable safer, smoother automated maneuvers, this thesis presents the design and implementation of a map-aware pedestrian intention prediction algorithm. The proposed approach leverages a hybrid architecture that combines deterministic feature engineering with deep learning. Specifically, a carefully selected set of kinematic and map-aware spatial features is manually extracted for each pedestrian in the scene to capture their relationship with the surrounding road topology. These explicit feature vectors are subsequently fed into a Stacked Long Short-Term Memory (LSTM) network, which calculates the probability of a pedestrian crossing within a predefined future time-horizon. By relying on hand-crafted feature extraction prior to the neural network phase, this methodology guarantees a significantly higher degree of interpretability compared to traditional, fully "black-box" end-to-end deep learning models. A key contribution of this work is the model’s flexibility and efficiency, enabling it to handle varying numbers of pedestrians and track lengths within a computationally efficient architecture. Two variants are considered: a Full Model using the complete feature set, and a Reduced Model retaining the most relevant features identified through ablation and tailored to the proprietary dataset, where map information is less reliable. The model is trained on a custom-annotated version of the Argoverse 2 dataset and validated on a proprietary real-world dataset, demonstrating the robustness and practical applicability of the proposed approach.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Dametto_Executive_Summary_02.pdf
non accessibile
Descrizione: executive summary
Dimensione
2.76 MB
Formato
Adobe PDF
|
2.76 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Dametto_Tesi_01.pdf
non accessibile
Descrizione: thesis
Dimensione
21.05 MB
Formato
Adobe PDF
|
21.05 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/252636