Modern Large Language Models (LLMs) can generate coherent and fluent text. However, the techniques used to make them more helpful and aligned with human preferences, such as Reinforcement Learning with Human Feedback (RLHF), can sometimes have unwanted effects. One of these is sycophancy, where a model tends to agree with the user's assumptions even when they are incorrect, instead of relying on factual information. To study this behavior, this thesis proposes a framework for measuring how different prompt formulations affect the objectivity of Large Language Models. The framework is based on a triple-prompt strategy, where the same claim is presented in three different ways: neutral (no additional context), leading (assumes it is true) and contradictory (claims it is false). By comparing the responses to these prompts, it is possible to measure the presence of confirmation bias. The evaluation covers multiple domains, including factual knowledge, common misconceptions, and complex reasoning tasks. A significative aspect of this work is the analysis of different evaluation methods for detecting and measuring confirmation bias. The results show that semantic similarity metrics become less reliable when applied to long LLM outputs, due to a phenomenon known as semantic flattening. To address this limitation, the framework also uses an LLM-as-a-Judge component, which provides a more reliable measure of bias. Overall, that the level of sycophancy depends on both the model and the type of task. Models mainly designed for conversation were more likely to follow misleading prompts, while reasoning models were generally more resistant, especially when the task required careful thinking before reaching a conclusion.

Moderni Large Language Models (LLMs) sono in grado di generare testi coerenti e fluenti. Tuttavia, le tecniche utilizzate per renderli più utili e allineati alle preferenze umane, come il Reinforcement Learning with Human Feedback (RLHF), possono a volte avere effetti indesiderati. Uno di questi è la sycophancy, in cui il modello tende ad accettare le assunzioni dell'utente anche quando sono errate, invece di basarsi su informazioni fattuali. Per studiare questo comportamento, questa tesi propone un framework per misurare come diverse formulazioni dei prompt influenzano l'oggettività dei Large Language Models. Il framework si basa su una triple-prompt strategy, in cui la stessa affermazione viene presentata in tre modi diversi: neutral (senza contesto aggiuntivo), leading (assume che sia vera), contradictory (afferma che sia falsa). Confrontando le risposte a questi prompt, è possibile misurare la presenza di confirmation bias. La valutazione copre diversi domini, tra cui domande fattuali, credenze comuni e problemi di ragionamento più complessi. Un aspetto significativo di questo lavoro è l'analisi di diversi metodi di valutazione per rilevare e misurare il confirmation bias. I risultati mostrano che le metriche di similarità semantica diventano meno affidabili quando applicate a output lunghi dei LLM, a causa di un fenomeno noto come semantic flattening. Per superare questa limitazione, il framework utilizza anche un componente LLM-as-a-Judge, che fornisce una misura più robusta del bias. Nel complesso, il livello di sycophancy dipende sia dal modello sia dal tipo di task. I modelli principalmente progettati per la conversazione tendono più spesso a seguire prompt ingannevoli, mentre i modelli ottimizzati per il ragionamento risultano generalmente più resistenti, soprattutto quando il compito richiede un processo più attento prima di arrivare alla conclusione.

Sycophancy vs. factuality: detecting confirmation bias in Large Language Models

FONTANA, FABRIZIO
2025/2026

Abstract

Modern Large Language Models (LLMs) can generate coherent and fluent text. However, the techniques used to make them more helpful and aligned with human preferences, such as Reinforcement Learning with Human Feedback (RLHF), can sometimes have unwanted effects. One of these is sycophancy, where a model tends to agree with the user's assumptions even when they are incorrect, instead of relying on factual information. To study this behavior, this thesis proposes a framework for measuring how different prompt formulations affect the objectivity of Large Language Models. The framework is based on a triple-prompt strategy, where the same claim is presented in three different ways: neutral (no additional context), leading (assumes it is true) and contradictory (claims it is false). By comparing the responses to these prompts, it is possible to measure the presence of confirmation bias. The evaluation covers multiple domains, including factual knowledge, common misconceptions, and complex reasoning tasks. A significative aspect of this work is the analysis of different evaluation methods for detecting and measuring confirmation bias. The results show that semantic similarity metrics become less reliable when applied to long LLM outputs, due to a phenomenon known as semantic flattening. To address this limitation, the framework also uses an LLM-as-a-Judge component, which provides a more reliable measure of bias. Overall, that the level of sycophancy depends on both the model and the type of task. Models mainly designed for conversation were more likely to follow misleading prompts, while reasoning models were generally more resistant, especially when the task required careful thinking before reaching a conclusion.
ING - Scuola di Ingegneria Industriale e dell'Informazione
22-lug-2026
2025/2026
Moderni Large Language Models (LLMs) sono in grado di generare testi coerenti e fluenti. Tuttavia, le tecniche utilizzate per renderli più utili e allineati alle preferenze umane, come il Reinforcement Learning with Human Feedback (RLHF), possono a volte avere effetti indesiderati. Uno di questi è la sycophancy, in cui il modello tende ad accettare le assunzioni dell'utente anche quando sono errate, invece di basarsi su informazioni fattuali. Per studiare questo comportamento, questa tesi propone un framework per misurare come diverse formulazioni dei prompt influenzano l'oggettività dei Large Language Models. Il framework si basa su una triple-prompt strategy, in cui la stessa affermazione viene presentata in tre modi diversi: neutral (senza contesto aggiuntivo), leading (assume che sia vera), contradictory (afferma che sia falsa). Confrontando le risposte a questi prompt, è possibile misurare la presenza di confirmation bias. La valutazione copre diversi domini, tra cui domande fattuali, credenze comuni e problemi di ragionamento più complessi. Un aspetto significativo di questo lavoro è l'analisi di diversi metodi di valutazione per rilevare e misurare il confirmation bias. I risultati mostrano che le metriche di similarità semantica diventano meno affidabili quando applicate a output lunghi dei LLM, a causa di un fenomeno noto come semantic flattening. Per superare questa limitazione, il framework utilizza anche un componente LLM-as-a-Judge, che fornisce una misura più robusta del bias. Nel complesso, il livello di sycophancy dipende sia dal modello sia dal tipo di task. I modelli principalmente progettati per la conversazione tendono più spesso a seguire prompt ingannevoli, mentre i modelli ottimizzati per il ragionamento risultano generalmente più resistenti, soprattutto quando il compito richiede un processo più attento prima di arrivare alla conclusione.
File allegati
File Dimensione Formato  
2026_07_Fontana.pdf

accessibile in internet per tutti

Descrizione: Testo della tesi
Dimensione 1.62 MB
Formato Adobe PDF
1.62 MB Adobe PDF Visualizza/Apri

I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10589/261118