In recent years, Retrieval-Augmented Generation (RAG) has become a common way to ground large language models (LLMs) on external knowledge. However, RAG pipelines introduce a clear security risk: an attacker can poison the retrieval corpus so that malicious documents are retrieved and then used to drive the final answer. This thesis addresses this problem by designing and implementing a practical, system-level defense framework that connects two complementary ideas: online protection with offline forensic attribution in a single workflow. We propose a merged pipeline that integrates TrustRAG as a real-time guardrail and RAGOrigin as a responsibility attribution module. The key contribution is not only to use both methods, but to make them operationally compatible: TrustRAG filters retrieved documents, the system then logs the trigger event and the related retrieval artifacts, and uses this signal to automatically start RAGOrigin offline attribution run. This enables a closed-loop “defend-then-attribute” process where TrustRAG reduces exposure at query time, and RAGOrigin narrows down which retrieved documents most likely caused a mis- generation and produces an attribution report that supports human remediation decisions, reviewing and removing suspect documents. We evaluate the framework on the BEIR Natural Questions benchmark and simulate retrieval poisoning using PoisonedRAG. Experimental results show that, under attack, Vanilla RAG achieves ACC = 55% with ASR = 48%, while enabling TrustRAG Stage 1 improves robustness to ACC = 73% and reduces ASR to 15During the RAGOrigin offline attribution, at document level, the integrated pipeline achieves DACC = 0.78, improving over running attribution on the full attacked scope reducing noise in the analyzed events. Overall, the proposed integration provides a first step toward a concrete and reproducible approach to defend RAG systems while also enabling post-incident investigation and re- mediation in realistic deployments.
Negli ultimi anni, Retrieval-Augmented Generation (RAG) è diventato uno strumento molto usato per aumentare il livello di precisione dei Large language models (LLM) utilizzando conoscenza esterna. Tuttavia, I RAG introducono una nuova superficie di rischio: un agente malevolo può avvelenare la base di dati esterna in modo da indurre l’estrazione e quindi l’utilizzo di documenti malevoli per influenzare la risposta finale del LLM. Questa tesi affronta questo problema progettando e implementando un sistema di difesa pratico che unisce due idee complementari: protezione online e attribuzione forense offline. Proponiamo una pipeline che integra TrustRAG come primo sistema di difesa in tempo reale e RAGOrigin come modulo di attribuzione della responsabilità. La contribuzione chiave non è solo quella di usare entrambi i metodi, ma di riuscire a farli lavorare insieme: TrustRAG filtra i documenti estratti, il sistema raccoglie le informazioni degli eventi che hanno fatto scattare il filtro e usa questo segnale per fare partire automaticamente il meccanismo di attribuzione di RAGOrigin. Tutto ciò porta a un processo unificato di "difesa e attribuzione" dove TrustRAG riduce l’esposizione a documenti malevoli durante l’utilizzo dell’utente, mentre RAGOrigin restringe il campo dei documenti estratti più probabilmente responsabili dell’evento di misgeneration producendo un report necessario a supportare una sanificazione della base di dati esterna da parte di un umano rimuovendo i documenti sospetti. Valutiamo il framework sul benchmark BEIR Natural Questions e simuliamo l’attaco usando PoisonedRAG. I risultati sperimentali mostrano che, sotto attacco, il RAG semplice raggiunge un ACC = 55% e un ASR = 48%, mentre attivando TrustRAG Stage 1 viene migliorata la robustezza del sistema con un ACC = 73% riducendo l’ASR al 15%. Durante l’attribuzione offline di RAGOrigin la nostra pipeline raggiunge un DACC = 0.78, migliorando leggermente le prestazioni rispetto al caso in cui venga eseguito RAGOrigin sull’intero set attaccato, riducendo il numero di eventi non di missgeneration analizzati. Nel complesso, il nostro sistema apre la strada verso un approccio concreto e riproducibile per difendere i sistemi RAG e al contempo permette un investigazione e sanificazione postincidente in situazioni reali.
Trust and trace: an ensemble approach against poisoning attacks in RAG-based LLMs
Fiano, Michael
2024/2025
Abstract
In recent years, Retrieval-Augmented Generation (RAG) has become a common way to ground large language models (LLMs) on external knowledge. However, RAG pipelines introduce a clear security risk: an attacker can poison the retrieval corpus so that malicious documents are retrieved and then used to drive the final answer. This thesis addresses this problem by designing and implementing a practical, system-level defense framework that connects two complementary ideas: online protection with offline forensic attribution in a single workflow. We propose a merged pipeline that integrates TrustRAG as a real-time guardrail and RAGOrigin as a responsibility attribution module. The key contribution is not only to use both methods, but to make them operationally compatible: TrustRAG filters retrieved documents, the system then logs the trigger event and the related retrieval artifacts, and uses this signal to automatically start RAGOrigin offline attribution run. This enables a closed-loop “defend-then-attribute” process where TrustRAG reduces exposure at query time, and RAGOrigin narrows down which retrieved documents most likely caused a mis- generation and produces an attribution report that supports human remediation decisions, reviewing and removing suspect documents. We evaluate the framework on the BEIR Natural Questions benchmark and simulate retrieval poisoning using PoisonedRAG. Experimental results show that, under attack, Vanilla RAG achieves ACC = 55% with ASR = 48%, while enabling TrustRAG Stage 1 improves robustness to ACC = 73% and reduces ASR to 15During the RAGOrigin offline attribution, at document level, the integrated pipeline achieves DACC = 0.78, improving over running attribution on the full attacked scope reducing noise in the analyzed events. Overall, the proposed integration provides a first step toward a concrete and reproducible approach to defend RAG systems while also enabling post-incident investigation and re- mediation in realistic deployments.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Fiano_Executive_Summary.pdf
accessibile in internet solo dagli utenti autorizzati
Dimensione
4.9 MB
Formato
Adobe PDF
|
4.9 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Fiano_Tesi.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Testo Tesi Aggiornato
Dimensione
5.49 MB
Formato
Adobe PDF
|
5.49 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/250587