The 3D Bin Packing Problem (3D-BPP) is a well-known combinatorial optimization chal- lenge with significant implications for operational efficiency in the logistics and shipping industries. Effective palletization requires rapid configuration, high volumetric utilization, and strict adherence to structural constraints. While Deep Reinforcement Learning (DRL) has emerged as a viable approach to solving such tasks, integrating it with rigorous industrial and physical constraints remains a persistent challenge. This thesis addresses this gap by first providing a systematic review and taxonomy of existing DRL applications within the 3D-BPP domain. The literature is categorized based on state and action space representations, reward shaping techniques, algorithmic frameworks, stability assessments, and the overall scope of applicability regarding grid resolution. Building on this analysis, we propose a novel architecture employing Hierarchical Rein- forcement Learning (HRL) to decompose the 3D problem into layer-building subtasks. To ensure practical applicability, we introduce the Cornered Anchored Regions (CAR) method, which efficiently reduces the action space while maintaining sufficient exploration capabilities and ensuring the static stability of each sequential loading position. The model is trained exclusively to construct a single layer of items. When tested on a real-world industrial dataset, the model successfully generates a complete layer in 79.22% of the instances. Furthermore, the model is tested in a full-bin packing scenario, where it achieves physical stability in 99.9% of the cases, with an average volumetric utilization of 64.21%, surpassing traditional heuristic methods by more than 10%, although a minor deficit is observed relative to an established Genetic Algorithm (GA) baseline deployed in real industrial palletization. As a competitive advantage, the model generates a loading sequence in 2.5 seconds per order, an acceleration of nearly 300 compared to the GA. These results demonstrate that the proposed approach offers a robust alternative for industrial automated palletization capable of achieving comparable performance at a fraction of the inference computational cost.
Il problema della pallettizzazione tridimensionale (3D-BPP) rappresenta una sfida di ottimizzazione combinatoria di rilievo per l’efficienza operativa nei settori della logistica e delle spedizioni. Un’efficace pallettizzazione richiede una configurazione rapida, un’elevata utilizzazione volumetrica e una rigorosa aderenza ai vincoli strutturali. Deep Reinforcement Learning (DRL) è emerso come un approccio valido per tali compiti; tuttavia, la sua integrazione con i vincoli industriali resta una sfida persistente. Tale lacuna viene affrontata nella presente tesi fornendo, in primo luogo, una revisione sistematica e una tassonomia delle applicazioni DRL esistenti nel dominio del 3D-BPP. La letteratura è categorizzata in base alle rappresentazioni degli spazi di stato e di azione, alle tecniche di reward shaping, ai framework algoritmici, alle valutazioni di stabilità e all’ambito generale di applicabilità relativo alla risoluzione della griglia. Sulla base di tale analisi, viene proposta una nuova architettura che impiega Hierarchical Reinforcement Learning (HRL) per scomporre il problema 3D in sottocompiti di costruzione di strati (layer-building). Al fine di garantire l’applicabilità pratica, si introduce il metodo delle Cornered Anchored Regions (CAR), mediante il quale lo spazio delle azioni viene ridotto, mantenendo capacità di esplorazione e garantendo la stabilità statica del pallet. L’addestramento è eseguito esclusivamente per la costruzione di un singolo strato di colli. A seguito dei test su un dataset industriale reale, uno strato completo viene generato con successo nel 79,22% dei casi. Inoltre, la valutazione viene estesa a uno scenario di imballag- gio completo, in cui la stabilità fisica è mantenuta nel 99,9% dei casi, con un’utilizzazione volumetrica media del 64,21%, superando l’efficacia di un metodo euristico di confronto di oltre il 10%. Sebbene sia osservato un lieve deficit rispetto a una consolidata baseline basata su Algoritmo Genetico (GA) impiegata in ambito industriale reale, il modello HRL mostra un forte vantaggio competitivo nella velocità di generazione delle sequenze di carico (circa 2,5 secondi per ordine), con un’accelerazione di quasi un fattore 300 rispetto al GA. Tali risultati dimostrano che l’approccio proposto è una solida alternativa per la pallettiz- zazione automatizzata industriale.
Hierarchical option-based reinfocement learning for efficient three-dimensional bin packing
ALLEGRINI, FAUSTO
2024/2025
Abstract
The 3D Bin Packing Problem (3D-BPP) is a well-known combinatorial optimization chal- lenge with significant implications for operational efficiency in the logistics and shipping industries. Effective palletization requires rapid configuration, high volumetric utilization, and strict adherence to structural constraints. While Deep Reinforcement Learning (DRL) has emerged as a viable approach to solving such tasks, integrating it with rigorous industrial and physical constraints remains a persistent challenge. This thesis addresses this gap by first providing a systematic review and taxonomy of existing DRL applications within the 3D-BPP domain. The literature is categorized based on state and action space representations, reward shaping techniques, algorithmic frameworks, stability assessments, and the overall scope of applicability regarding grid resolution. Building on this analysis, we propose a novel architecture employing Hierarchical Rein- forcement Learning (HRL) to decompose the 3D problem into layer-building subtasks. To ensure practical applicability, we introduce the Cornered Anchored Regions (CAR) method, which efficiently reduces the action space while maintaining sufficient exploration capabilities and ensuring the static stability of each sequential loading position. The model is trained exclusively to construct a single layer of items. When tested on a real-world industrial dataset, the model successfully generates a complete layer in 79.22% of the instances. Furthermore, the model is tested in a full-bin packing scenario, where it achieves physical stability in 99.9% of the cases, with an average volumetric utilization of 64.21%, surpassing traditional heuristic methods by more than 10%, although a minor deficit is observed relative to an established Genetic Algorithm (GA) baseline deployed in real industrial palletization. As a competitive advantage, the model generates a loading sequence in 2.5 seconds per order, an acceleration of nearly 300 compared to the GA. These results demonstrate that the proposed approach offers a robust alternative for industrial automated palletization capable of achieving comparable performance at a fraction of the inference computational cost.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Allegrini_Executive_Summary_02.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
525.74 kB
Formato
Adobe PDF
|
525.74 kB | Adobe PDF | Visualizza/Apri |
|
2026_03_Allegrini_Thesis.pdf
accessibile in internet per tutti
Descrizione: Thesis
Dimensione
5.12 MB
Formato
Adobe PDF
|
5.12 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253593