Continuous multi-agent fleet coordination in modern automated warehouses is frequently modeled as the Multi-Agent Pickup and Delivery (MAPD) problem. While recent predictive frameworks, such as MAPD with Task Probability Distributions (MAPD-P), significantly minimize average service times by speculatively routing unassigned agents to high-probability task generation zones, their efficiency remains closely tied to the accuracy of the underlying forecast. Under noisy or imperfect probability distributions, independent local decision-making can prompt agents to over-commit to highly speculative regions, inflating physical travel costs and reducing response efficiency when real demand emerges elsewhere. To address this sensitivity to predictive inaccuracies, this thesis introduces a cooperative, spatial-awareness framework denoted as Spatio-Temporal Regret (STR). Embedded within the Token Passing paradigm, the framework introduces two distinct cooperative mechanisms. For the positioning of idle agents, the TP-m2_STR algorithm evaluates an agent's proximity to a predicted zone against the effective temporal distance of its peers, scaling the baseline task probability via a continuous piecewise logarithmic regret function to maintain a balanced spatial fleet distribution. For the assignment of tasks to agents, the TP-m1_STR algorithm applies a dual-regret formulation to balance local speculative anticipation against global service coverage, mitigating localized task omissions. The proposed mechanisms are evaluated via dynamic simulation across varying grid complexities and operating conditions. Experimental results demonstrate that under standard predictive noise injections, the piecewise logarithmic idle positioning mechanism yields a measurable reduction in the cost of the solution and effectively reduces redundant, unproductive travel without compromising system throughput. Conversely, the application of regret logic within the task assignment phase yielded negligible performance variations, highlighting the structural limitations of local decision rules under high task saturation. Collectively, these findings indicate that incorporating spatio-temporal regret provides a viable behavioral safety net, enhancing the spatial stability and operational robustness of predictive multi-agent systems operating under uncertain forecasts.
Il coordinamento continuo di flotte multi-agente nei moderni magazzini automatizzati è frequentemente modellato come il problema del Multi-Agent Pickup and Delivery (MAPD). Sebbene recenti framework predittivi, come il MAPD con Distribuzioni di Probabilità dei Compiti (MAPD-P), riducano significativamente i tempi medi di servizio indirizzando speculativamente gli agenti non assegnati verso “zone calde” nelle quali i task potrebbero apparire con alta probabilità, la loro efficienza rimane strettamente legata all'accuratezza delle previsioni sottostanti. In presenza di distribuzioni di probabilità imprecise o rumorose, i processi decisionali locali e indipendenti possono spingere gli agenti a impegnarsi eccessivamente in regioni altamente speculative, incrementando i costi di percorrenza fisica e riducendo la reattività del sistema quando la domanda reale emerge altrove. Per mitigare questa sensibilità alle imprecisioni predittive, questa tesi introduce un framework cooperativo di consapevolezza spaziale denominato Spatio-Temporal Regret (STR). Integrato all'interno del paradigma del Token Passing, il framework introduce due distinti meccanismi cooperativi. Per la fase di posizionamento degli agenti inattivi, l'algoritmo TP-m2_STR valuta la prossimità di un agente a una zona “calda” rispetto alla distanza temporale effettiva dei suoi pari, scalando la probabilità di base tramite una funzione continua piecewise logarithmic regret per mantenere una distribuzione spaziale bilanciata della flotta. Per la fase di assegnazione dei task agli agenti, l'algoritmo TP-m1_STR applica una formulazione a doppio regret per bilanciare l'anticipazione speculativa locale con la copertura globale del servizio, mitigando le omissioni localizzate dei compiti. I meccanismi proposti sono valutati tramite simulazioni dinamiche su griglie di varia complessità e diverse condizioni operative. I risultati sperimentali dimostrano che, in presenza di rumore predittivo, il meccanismo di posizionamento logaritmico a tratti per gli agenti inattivi genera una riduzione misurabile del costo della soluzione e limita efficacemente i viaggi ridondanti e improduttivi senza compromettere la produttività complessiva del sistema. Al contrario, l'applicazione della logica di regret nella fase di assegnazione dei compiti ha mostrato variazioni prestazionali trascurabili, evidenziando i limiti strutturali delle regole decisionali locali in scenari ad alta saturazione. Complessivamente, questi risultati indicano che l'integrazione del regret spazio-temporale fornisce una valida rete di sicurezza comportamentale, migliorando la stabilità spaziale e la robustezza operativa dei sistemi multi-agente predittivi che operano con previsioni imperfette.
Spatio-temporal regret mechanisms for multi-agent pickup and delivery with task probability distributions
NÚÑEZ MILANÉS, RICARDO ARTURO
2025/2026
Abstract
Continuous multi-agent fleet coordination in modern automated warehouses is frequently modeled as the Multi-Agent Pickup and Delivery (MAPD) problem. While recent predictive frameworks, such as MAPD with Task Probability Distributions (MAPD-P), significantly minimize average service times by speculatively routing unassigned agents to high-probability task generation zones, their efficiency remains closely tied to the accuracy of the underlying forecast. Under noisy or imperfect probability distributions, independent local decision-making can prompt agents to over-commit to highly speculative regions, inflating physical travel costs and reducing response efficiency when real demand emerges elsewhere. To address this sensitivity to predictive inaccuracies, this thesis introduces a cooperative, spatial-awareness framework denoted as Spatio-Temporal Regret (STR). Embedded within the Token Passing paradigm, the framework introduces two distinct cooperative mechanisms. For the positioning of idle agents, the TP-m2_STR algorithm evaluates an agent's proximity to a predicted zone against the effective temporal distance of its peers, scaling the baseline task probability via a continuous piecewise logarithmic regret function to maintain a balanced spatial fleet distribution. For the assignment of tasks to agents, the TP-m1_STR algorithm applies a dual-regret formulation to balance local speculative anticipation against global service coverage, mitigating localized task omissions. The proposed mechanisms are evaluated via dynamic simulation across varying grid complexities and operating conditions. Experimental results demonstrate that under standard predictive noise injections, the piecewise logarithmic idle positioning mechanism yields a measurable reduction in the cost of the solution and effectively reduces redundant, unproductive travel without compromising system throughput. Conversely, the application of regret logic within the task assignment phase yielded negligible performance variations, highlighting the structural limitations of local decision rules under high task saturation. Collectively, these findings indicate that incorporating spatio-temporal regret provides a viable behavioral safety net, enhancing the spatial stability and operational robustness of predictive multi-agent systems operating under uncertain forecasts.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Nunez.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: complete thesis
Dimensione
3.31 MB
Formato
Adobe PDF
|
3.31 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260969