Opponent Modeling (OM) is a powerful framework in Multi-Agent Reinforcement Learning (MARL) to anticipate and adapt to the strategies of other agents. However, its success is highly dependent on the assumption of high-quality observations. In many real-world applications, agents must operate under imperfect information that can lead to inaccurate model representations. In our work, we investigate the drawbacks of agents conditioning their policies on flawed opponent models that may cause significant performance degradation compared to model-agnostic baselines. We analyze two state-of-the-art models that rely on perfect-information assumptions and demonstrate how their performance and robustness deteriorate when these assumptions do not hold. To address these limitations, we introduce Strategy Weighting for Adaptive Policies (SWAP), a novel adaptive framework that treats strategy selection as an online learning problem. Employing the EXP4 algorithm, our agent treats a predictive OM-based policy and a robust conservative policy as competing experts, dynamically switching between them based on their observed performance. Our experimental results demonstrate the advantages of adopting a conservative approach when information is flawed and using predictive modeling when information is reliable, outperforming state-of-the-art methods in these critical scenarios.
La Modellazione dell'Avversario (OM) è un approccio potente nell'ambito dell'Apprendimento per Rinforzo Multi-Agente (MARL) che consente di anticipare le strategie degli altri agenti e di adattarsi ad esse. Tuttavia, il suo successo dipende in larga misura dalla disponibilità di osservazioni di alta qualità. In molte applicazioni del mondo reale, gli agenti devono operare in condizioni di informazione imperfetta, il che può portare a rappresentazioni del modello imprecise. Nel nostro lavoro, esaminiamo gli svantaggi derivanti dal fatto che gli agenti basino le proprie politiche su modelli dell'avversario imperfetti, che possono causare un calo significativo delle prestazioni rispetto ai modelli di riferimento indipendenti dal modello dell'avversario. Analizziamo due modelli allo stato dell'arte che si basano su ipotesi di informazione perfetta e dimostriamo come le loro prestazioni e la loro robustezza peggiorino quando tali ipotesi non sono valide. Per ovviare a tali limiti, introduciamo SWAP (Selezione Ponderata di Strategie per Politiche Adattive), un nuovo quadro adattivo che considera la selezione della strategia come un problema di apprendimento in tempo reale. Impiegando l'algoritmo EXP4, il nostro agente considera una politica predittiva basata sul modello dell'avversario e una politica conservativa robusta come esperti in competizione tra loro, passando dinamicamente dall'una all'altra in base alle prestazioni osservate. I risultati sperimentali dimostrano i vantaggi di adottare un approccio prudente quando le informazioni sono incomplete e di ricorrere alla modellizzazione predittiva quando le informazioni sono affidabili, superando così i metodi all'avanguardia in questi scenari critici.
Opponent Modeling uncertainty in Multi-Agent Reinforcement learning
Maifredi, Francesca
2025/2026
Abstract
Opponent Modeling (OM) is a powerful framework in Multi-Agent Reinforcement Learning (MARL) to anticipate and adapt to the strategies of other agents. However, its success is highly dependent on the assumption of high-quality observations. In many real-world applications, agents must operate under imperfect information that can lead to inaccurate model representations. In our work, we investigate the drawbacks of agents conditioning their policies on flawed opponent models that may cause significant performance degradation compared to model-agnostic baselines. We analyze two state-of-the-art models that rely on perfect-information assumptions and demonstrate how their performance and robustness deteriorate when these assumptions do not hold. To address these limitations, we introduce Strategy Weighting for Adaptive Policies (SWAP), a novel adaptive framework that treats strategy selection as an online learning problem. Employing the EXP4 algorithm, our agent treats a predictive OM-based policy and a robust conservative policy as competing experts, dynamically switching between them based on their observed performance. Our experimental results demonstrate the advantages of adopting a conservative approach when information is flawed and using predictive modeling when information is reliable, outperforming state-of-the-art methods in these critical scenarios.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Maifredi_Tesi.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Testo della tesi
Dimensione
6.84 MB
Formato
Adobe PDF
|
6.84 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Maifredi_Executive Summary.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Executive summary
Dimensione
455.86 kB
Formato
Adobe PDF
|
455.86 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261481