The remarkable success of Reinforcement Learning in solving complex decision-making problems is built upon the foundational assumption of environmental stationarity. Although this simplification has enabled the development of algorithms with strong convergence guarantees, capable of learning stable, optimal policies through interaction with fixed environments, it rarely holds in real-world applications where system dynamics evolve over time. This challenge is further compounded by partial observability, as the underlying processes driving these environmental shifts are often unobservable, making them subtle and difficult for the agent to detect. This work specifically addresses non-stationary reinforcement learning under partial observability, modelling the environment as a Hidden Markov Model where a latent context governs the transition dynamics. We propose a decoupled framework that separates latent mode estimation from policy optimisation. The latent estimator is learned offline from a static dataset of transitions, using an extended version of the classical Baum-Welch algorithm to account for conditioning on the state-action tuples. The learned model is then deployed online to maintain a belief over the currently active latent mode, using Bayesian filtering to augment the agent's observation space, thus reducing the non-stationary, partially observable problem to a stationary belief Markov Decision Process, solvable with standard reinforcement learning algorithms. We evaluate this framework on discrete and continuous non-stationary variants of the Cliff Walking environment, under both deterministic and stochastic latent dynamics. Our results demonstrate that the proposed latent estimator consistently bridges the performance gap with the oracle baseline, achieving convergence in scenarios where the naive agent systematically fails.
Il notevole successo dell'apprendimento per rinforzo (Reinforcement Learning) nella risoluzione di problemi decisionali complessi si basa su un'assunzione fondamentale: la stazionarietà dell'ambiente. Sebbene tale semplificazione abbia consentito lo sviluppo di algoritmi dotati di forti garanzie di convergenza, capaci di apprendere politiche stabili e ottimali in contesti fissi, essa raramente trova riscontro nelle applicazioni del mondo reale, dove la dinamica del sistema evolve nel tempo. Questa sfida è ulteriormente aggravata dalla parziale osservabilità, poiché i processi sottostanti che guidano tali cambiamenti ambientali sono spesso non osservabili, subdoli e difficili da rilevare per l'agente. Questo lavoro si occupa esattamente di apprendimento per rinforzo in contesti non stazionari di osservabilità parziale, modellando l'ambiente come un modello Markoviano nascosto in cui un contesto latente ne governa le dinamiche di transizione. Proponiamo un'architettura disaccoppiata che separa la stima dello stato latente dall'ottimizzazione della politica. Lo stimatore latente è appreso offline da un dataset statico di transizioni tramite una versione estesa dell'algoritmo classico di Baum-Welch che condiziona le emissioni sulle tuple stato-azione. Questo modello viene poi impiegato online, tramite un filtraggio Bayesiano, per mantenere un'opinione (belief) sul contesto latente attualmente attivo. In questo modo, il problema non stazionario e parzialmente osservabile viene ridotto a un classico processo decisionale Markoviamo, risolvibile tramite algoritmi standard di Reinforcement Learning. Abbiamo validato questo sistema su varianti non stazionarie, sia discrete che continue, dell'ambiente Cliff Walking, sia sotto dinamiche latenti deterministiche che stocastiche. I nostri risultati dimostrano che lo stimatore latente proposto colma il divario di prestazioni rispetto all'oracolo, raggiungendo la convergenza in scenari in cui l'agente ingenuo fallisce sistematicamente.
Handling partial observability in non-stationary reinforcement learning via hidden Markov models
Motti, Caterina
2025/2026
Abstract
The remarkable success of Reinforcement Learning in solving complex decision-making problems is built upon the foundational assumption of environmental stationarity. Although this simplification has enabled the development of algorithms with strong convergence guarantees, capable of learning stable, optimal policies through interaction with fixed environments, it rarely holds in real-world applications where system dynamics evolve over time. This challenge is further compounded by partial observability, as the underlying processes driving these environmental shifts are often unobservable, making them subtle and difficult for the agent to detect. This work specifically addresses non-stationary reinforcement learning under partial observability, modelling the environment as a Hidden Markov Model where a latent context governs the transition dynamics. We propose a decoupled framework that separates latent mode estimation from policy optimisation. The latent estimator is learned offline from a static dataset of transitions, using an extended version of the classical Baum-Welch algorithm to account for conditioning on the state-action tuples. The learned model is then deployed online to maintain a belief over the currently active latent mode, using Bayesian filtering to augment the agent's observation space, thus reducing the non-stationary, partially observable problem to a stationary belief Markov Decision Process, solvable with standard reinforcement learning algorithms. We evaluate this framework on discrete and continuous non-stationary variants of the Cliff Walking environment, under both deterministic and stochastic latent dynamics. Our results demonstrate that the proposed latent estimator consistently bridges the performance gap with the oracle baseline, achieving convergence in scenarios where the naive agent systematically fails.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Motti_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
1.19 MB
Formato
Adobe PDF
|
1.19 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Motti_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
792.57 kB
Formato
Adobe PDF
|
792.57 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260408