Large language models (LLMs) are increasingly deployed as global advisory systems, yet the moral reasoning patterns embedded through their training remain poorly understood, particularly regarding how these systems adapt (or fail to adapt) their ethical responses across diverse cultural contexts. This thesis develops a novel framework for quantifying how LLMs transform moral content from input queries to generated responses, using Moral Foundations Theory (MFT) as the analytical lens. We introduce a two-layer hierarchical Bayesian architecture that addresses fundamental measurement challenges in computational moral assessment. Layer 1 fuses annotations from four LLM-based MFT extraction methods (BHJ Binary, MoVA, Direct Continuous, Dual Extraction), explicitly modeling method-specific biases and noise to recover latent “true” divergence estimates with calibrated uncertainty. Dual Extraction achieves 2.3× lower measurement noise than binary methods, validating its use as anchor method. Layer 2 models these fused divergences as functions of cultural region, model architecture, and query characteristics, enabling detection of systematic moral transformation patterns. Analyzing 9,400 query-response pairs from the SafeWORLD benchmark across four LLMs (Llama 3.1 8B/70B, Qwen 2.5 14B, GPT-5 mini) spanning 50 countries, we find that models exhibit foundation-specific transformation patterns: Care/Harm is systemati- cally amplified (μ∆ = +0.226), while Loyalty/Betrayal is suppressed (μ∆ = −0.070). Critically, Region×Foundation interactions emerge as the dominant variance component (20.2× model effects), revealing that models do respond differently to cultural context, but in potentially maladaptive ways. Care amplification is strongest for MENA and Asia- Pacific queries compared to Europe, while Loyalty suppression is particularly intense in Americas and Europe. These patterns are remarkably invariant across model architec- tures (model effects negligible at 1.0× baseline), suggesting moral transformation patterns may be determined by shared training data distributions rather than architectural choices.
I modelli linguistici di grandi dimensioni (LLM) vengono sempre più utilizzati come sis- temi di consulenza globali, ma i pattern di ragionamento morale intrinseci rimangono poco compresi, in particolare riguardo a come questi sistemi adattano le risposte etiche in con- testi culturali diversi. Questa tesi sviluppa un framework per quantificare come gli LLM trasformano il contenuto morale dalle query di input alle risposte generate, utilizzando la Teoria dei Fondamenti Morali (MFT) come lente analitica. Introduciamo un’architettura bayesiana gerarchica a due livelli. Il livello 1 fonde le anno- tazioni di quattro metodi di estrazione MFT basati su LLM, modellando esplicitamente bias e rumore specifici del metodo per recuperare stime delle divergenze latenti con in- certezza calibrata. Il livello 2 modella queste divergenze come funzioni della regione culturale, dell’architettura del modello e delle caratteristiche della query, consentendo il rilevamento di pattern sistematici di trasformazione morale. Analizzando 9.400 coppie query-risposta dal benchmark SafeWORLD su quattro LLM (Llama 3.1 8B/70B, Qwen 2.5 14B, GPT-5 mini) in 50 paesi, riscontriamo pattern specifici per fondamento: Care/Harm viene sistematicamente amplificato, mentre Loyalty/Betrayal viene soppresso. Le interazioni Regione×Fondamento emergono come componente domi- nante della varianza (20,2× gli effetti modello), rivelando un adattamento culturale poten- zialmente disadattivo: l’amplificazione di Care è più forte per MENA e Asia-Pacifico, la soppressione di Loyalty è intensa nelle Americhe e in Europa. Questi pattern sono in- varianti tra i modelli (effetti modello: 1,0× baseline), suggerendo che la trasformazione morale dipende dai dati di addestramento condivisi più che dall’architettura.
A hierarchical bayesian framework for cross-cultural divergence analysis - quantifying biases, suppressions and amplifications of moral foundations in Large Language Models
BELLINI, EMANUELE
2024/2025
Abstract
Large language models (LLMs) are increasingly deployed as global advisory systems, yet the moral reasoning patterns embedded through their training remain poorly understood, particularly regarding how these systems adapt (or fail to adapt) their ethical responses across diverse cultural contexts. This thesis develops a novel framework for quantifying how LLMs transform moral content from input queries to generated responses, using Moral Foundations Theory (MFT) as the analytical lens. We introduce a two-layer hierarchical Bayesian architecture that addresses fundamental measurement challenges in computational moral assessment. Layer 1 fuses annotations from four LLM-based MFT extraction methods (BHJ Binary, MoVA, Direct Continuous, Dual Extraction), explicitly modeling method-specific biases and noise to recover latent “true” divergence estimates with calibrated uncertainty. Dual Extraction achieves 2.3× lower measurement noise than binary methods, validating its use as anchor method. Layer 2 models these fused divergences as functions of cultural region, model architecture, and query characteristics, enabling detection of systematic moral transformation patterns. Analyzing 9,400 query-response pairs from the SafeWORLD benchmark across four LLMs (Llama 3.1 8B/70B, Qwen 2.5 14B, GPT-5 mini) spanning 50 countries, we find that models exhibit foundation-specific transformation patterns: Care/Harm is systemati- cally amplified (μ∆ = +0.226), while Loyalty/Betrayal is suppressed (μ∆ = −0.070). Critically, Region×Foundation interactions emerge as the dominant variance component (20.2× model effects), revealing that models do respond differently to cultural context, but in potentially maladaptive ways. Care amplification is strongest for MENA and Asia- Pacific queries compared to Europe, while Loyalty suppression is particularly intense in Americas and Europe. These patterns are remarkably invariant across model architec- tures (model effects negligible at 1.0× baseline), suggesting moral transformation patterns may be determined by shared training data distributions rather than architectural choices.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Bellini_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
15.4 MB
Formato
Adobe PDF
|
15.4 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Bellini_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
457.55 kB
Formato
Adobe PDF
|
457.55 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/251541