Deep-space exploration with CubeSats is gaining momentum, yet operating in such perturbed environments remains a significant challenge. In the context of the Engineering Extremely Rare Events in Astrodynamics for Deep-Space Missions in Autonomy (EXTREMA) project, the reliance on fixed duty cycles for low-thrust Solar Electric Propulsion (SEP) leads to increased state uncertainty and propellant consumption, often necessitating costly ground-based interventions. This thesis investigates the application of Deep Reinforcement Learning (DRL) to enable fully autonomous on-board guidance. By modelling the duty cycle supervision as a Markov Decision Process (MDP), the spacecraft is designed to autonomously decide in real-time whether to proceed with the nominal schedule or trigger a trajectory re-optimization, basing its decisions on the state estimates provided by an on-board Unscented Kalman Filter (UKF). Using the Proximal Policy Optimization (PPO) algorithm, a neural network policy was trained to perform these supervisory tasks. However, preliminary training analysis on asteroid transfer scenarios revealed critical bottlenecks in the interaction between the DRL agent and the underlying trajectory optimizer. Due to the narrow convergence radius of the indirect FOCP solver, the agent learned a sub-optimal survival strategy, systematically avoiding re-optimizations to prevent solver divergence and the associated severe terminal penalties. While these results underscore the current limitations of coupling DRL with highly sensitive numerical solvers, this work successfully establishes a rigorous autonomous framework. The findings provide fundamental insights into reward engineering dynamics and highlight the necessity for robust convex optimizers and hierarchical control architectures, setting a clear trajectory for the development of the next generation of self-driving interplanetary CubeSats.
L’esplorazione dello spazio profondo con i CubeSat sta vivendo una fase di forte slancio, sebbene operare in ambienti così perturbati rappresenti ancora una sfida ingegneristica notevole. Nel contesto del progetto Engineering Extremely Rare Events in Astrodynamics for Deep-Space Missions in Autonomy (EXTREMA), la dipendenza da duty cycle prefissati per la propulsione solare elettrica (Solar Electric Propulsion (SEP)) a bassa spinta comporta un aumento dell’incertezza sullo stato e del consumo di propellente, rendendo spesso necessari costosi interventi da terra. Questa tesi indaga l’applicazione del Deep Reinforcement Learning (DRL) per abilitare un sistema di guida di bordo completamente autonomo. Modellando la supervisione del duty cycle come un Markov Decision Process (MDP), il veicolo spaziale è progettato per decidere autonomamente e in tempo reale se procedere con la pianificazione nominale o innescare una ri-ottimizzazione della traiettoria, basando le proprie decisioni sulle stime di stato fornite da un Unscented Kalman Filter (UKF) di bordo. Utilizzando l’algoritmo Proximal Policy Optimization (PPO), è stata addestrata una policy basata su reti neurali per eseguire questi compiti di supervisione. Tuttavia, l’analisi preliminare dell’addestramento su scenari di trasferimento verso asteroidi ha rivelato delle criticità nell’interazione tra l’agente DRL e l’ottimizzatore di traiettoria sottostante. A causa del ristretto bacino di convergenza del solutore indiretto FOCP, l’agente ha appreso una strategia di sopravvivenza sub-ottimale, evitando sistematicamente le riottimizzazioni per prevenire la divergenza del solutore e le conseguenti, severe penalità terminali. Sebbene questi risultati evidenzino le attuali limitazioni nell’accoppiare il DRL con solutori numerici altamente sensibili, questo lavoro definisce con successo un rigoroso framework autonomo. I risultati ottenuti forniscono indicazioni fondamentali sulle dinamiche di reward engineering e sottolineano la necessità di adottare solutori convessi più robusti e architetture di controllo gerarchiche, tracciando una rotta chiara per lo sviluppo della futura generazione di CubeSat interplanetari a guida autonoma.
Exploring a deep reinforcement Learning approach for on-board duty cycle control of low-thrust CubeSats
Puzzolante, Michele
2024/2025
Abstract
Deep-space exploration with CubeSats is gaining momentum, yet operating in such perturbed environments remains a significant challenge. In the context of the Engineering Extremely Rare Events in Astrodynamics for Deep-Space Missions in Autonomy (EXTREMA) project, the reliance on fixed duty cycles for low-thrust Solar Electric Propulsion (SEP) leads to increased state uncertainty and propellant consumption, often necessitating costly ground-based interventions. This thesis investigates the application of Deep Reinforcement Learning (DRL) to enable fully autonomous on-board guidance. By modelling the duty cycle supervision as a Markov Decision Process (MDP), the spacecraft is designed to autonomously decide in real-time whether to proceed with the nominal schedule or trigger a trajectory re-optimization, basing its decisions on the state estimates provided by an on-board Unscented Kalman Filter (UKF). Using the Proximal Policy Optimization (PPO) algorithm, a neural network policy was trained to perform these supervisory tasks. However, preliminary training analysis on asteroid transfer scenarios revealed critical bottlenecks in the interaction between the DRL agent and the underlying trajectory optimizer. Due to the narrow convergence radius of the indirect FOCP solver, the agent learned a sub-optimal survival strategy, systematically avoiding re-optimizations to prevent solver divergence and the associated severe terminal penalties. While these results underscore the current limitations of coupling DRL with highly sensitive numerical solvers, this work successfully establishes a rigorous autonomous framework. The findings provide fundamental insights into reward engineering dynamics and highlight the necessity for robust convex optimizers and hierarchical control architectures, setting a clear trajectory for the development of the next generation of self-driving interplanetary CubeSats.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Puzzolante.pdf
accessibile in internet solo dagli utenti autorizzati
Dimensione
5.61 MB
Formato
Adobe PDF
|
5.61 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253291