In this thesis, we investigate the application of Reinforcement Learning to the control of parametrized dynamical systems, which commonly arise in scientific and engineering applications. These systems are often characterized by high dimensionality, expensive numerical simulations, and partial knowledge of the underlying dynamics. In practice, their behavior is typically described by approximate models and depends on physical parameters that may be uncertain or difficult to estimate accurately. In these settings, standard reinforcement learning methods based on Q-learning often suffer from poor robustness and limited data efficiency. In particular, errors introduced in the Bellman backup can propagate through successive updates, leading to inaccurate value estimates and convergence to sub-optimal control policies. To address these challenges, we propose a novel reinforcement learning framework, named HyperSUNRISE. The proposed method combines two complementary ideas. First, an ensemble of policy-value functions is used to explicitly estimate epistemic uncertainty and improving the balance between exploration and exploitation during training. Second, both policies and value functions are parameterized using hypernetworks, which condition the model parameters on physical system parameters and enable effective generalization across different dynamical regimes. The proposed framework is evaluated on the control of two parametrized dynamical systems, namely a 1D Kuramoto-Sivashinsky equation and a particle-navigation problem in a 2D gyre flow, with a particular focus on robustness to sensor noise and model-parameter misspecification. Numerical results show that HyperSUNRISE achieves improved training stability, higher sample efficiency, and increased robustness to uncertainty and noise affecting both the system dynamics and the available measurements.
In questa tesi, studiamo l’applicazione del reinforcement learning (RL) al controllo di sistemi dinamici parametrizzati, che emergono comunemente in applicazioni scientifiche e ingegneristiche. Tali sistemi sono spesso caratterizzati da elevata dimensionalità, simulazioni numeriche computazionalmente costose e da una conoscenza parziale delle dinamiche sottostanti. In pratica, il loro comportamento è generalmente descritto da modelli approssimati e dipende da parametri fisici che possono essere incerti o difficili da stimare con precisione. In questi contesti, i metodi standard di RL basati su Q-learning soffrono spesso di scarsa robustezza e limitata efficienza di apprendimento dai dati. In particolare, gli errori introdotti nel Bellman backup possono propagarsi attraverso aggiornamenti successivi, portando a stime inaccurate della value function e alla convergenza verso politiche di controllo subottimali. Per affrontare queste sfide, proponiamo un nuovo framework di reinforcement learning, denominato HyperSUNRISE. Il metodo proposto combina due idee complementari. In primo luogo, un ensemble di policy-value functions è utilizzato per stimare esplicitamente l’incertezza epistemica e migliorare il bilanciamento tra exploration ed exploitation durante il training. In secondo luogo, sia le policies sia le value functions sono parametrizzate tramite hypernetworks, che condizionano i parametri del modello sui parametri fisici del sistema e consentono una generalizzazione efficace tra diversi regimi dinamici. Il framework proposto è valutato sul controllo di due sistemi dinamici parametrizzati, ovvero l’equazione di Kuramoto-Sivashinsky in 1D e un problema di navigazione di una particella in un flusso di gyre bidimensionale, con particolare attenzione alla robustezza rispetto al rumore sui sensori e all’errata specificazione dei parametri del modello. I risultati numerici mostrano che HyperSUNRISE ottiene una maggiore stabilità in fase di training, una maggiore efficienza campionaria e una maggiore robustezza rispetto all'incertezza e al rumore che influenzano sia le dinamiche del sistema sia le misurazioni disponibili.
HyperSUNRISE: ensemble-based Reinforcement Learning with hypernetworks for robust control of parametrized systems under noisy measurements and model uncertainty
PASCALI, GABRIELE
2024/2025
Abstract
In this thesis, we investigate the application of Reinforcement Learning to the control of parametrized dynamical systems, which commonly arise in scientific and engineering applications. These systems are often characterized by high dimensionality, expensive numerical simulations, and partial knowledge of the underlying dynamics. In practice, their behavior is typically described by approximate models and depends on physical parameters that may be uncertain or difficult to estimate accurately. In these settings, standard reinforcement learning methods based on Q-learning often suffer from poor robustness and limited data efficiency. In particular, errors introduced in the Bellman backup can propagate through successive updates, leading to inaccurate value estimates and convergence to sub-optimal control policies. To address these challenges, we propose a novel reinforcement learning framework, named HyperSUNRISE. The proposed method combines two complementary ideas. First, an ensemble of policy-value functions is used to explicitly estimate epistemic uncertainty and improving the balance between exploration and exploitation during training. Second, both policies and value functions are parameterized using hypernetworks, which condition the model parameters on physical system parameters and enable effective generalization across different dynamical regimes. The proposed framework is evaluated on the control of two parametrized dynamical systems, namely a 1D Kuramoto-Sivashinsky equation and a particle-navigation problem in a 2D gyre flow, with a particular focus on robustness to sensor noise and model-parameter misspecification. Numerical results show that HyperSUNRISE achieves improved training stability, higher sample efficiency, and increased robustness to uncertainty and noise affecting both the system dynamics and the available measurements.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Pascali_Tesi.pdf
non accessibile
Descrizione: Tesi
Dimensione
10.15 MB
Formato
Adobe PDF
|
10.15 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Pascali_Executive_Summary.pdf
non accessibile
Descrizione: Executive summary
Dimensione
5.17 MB
Formato
Adobe PDF
|
5.17 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253415