This thesis presents the design, implementation, and evaluation of a Retrieval-Augmented Generation (RAG) system for question answering over football-oriented technical, regula- tory, and documentary sources. The work was developed in an industrial context, where users needed a more flexible way to access large collections of heterogeneous documents through natural language queries. The proposed system follows a modular RAG pipeline composed of document parsing, chunking, embedding generation, vector database indexing, retrieval, and answer gener- ation. Different configurations were evaluated in order to select a final setup that bal- ances retrieval quality, computational efficiency, and grounded answer generation. The evaluation considered both traditional Information Retrieval metrics and LLM-as-a-judge metrics, allowing the retrieval and generation stages to be analysed from complementary perspectives. The results show that the quality of a RAG system depends not only on the generative model, but also on earlier pipeline components such as parsing, chunking, indexed repre- sentation, retrievalstrategy, andindexconfiguration. Thefinalsystemprioritizesevidence coverage, low retrieval latency, and answer faithfulness, while avoiding components whose computational cost is not justified by consistent quality improvements. Since no manually annotated ground-truth dataset was available, part of the evaluation relied on synthetically generated reference answers. For this reason, the reported scores should be interpreted mainly as comparative indicators between configurations rather than as absolute measures of system performance. Overall, the thesis demonstrates how a domain-specific RAG system can be built and evaluated through a structured experimental methodology for reliable document-based question answering.
Questa tesi presenta la progettazione, l’implementazione e la valutazione di un sistema di Retrieval-Augmented Generation (RAG) per il question answering su fonti tecniche, normative e documentali orientate al mondo del calcio. Il lavoro è stato sviluppato in un contesto industriale, in cui gli utenti necessitavano di un modo più flessibile per accedere a grandi raccolte di documenti eterogenei attraverso interrogazioni in linguaggio naturale. Il sistema proposto segue una pipeline RAG modulare composta da parsing dei documenti, suddivisione in chunk, generazione degli embedding, indicizzazione in un database vettoriale, retrieval e generazione delle risposte. Sono state valutate diverse configurazioni al fine di selezionare una soluzione finale che bilanciasse qualità del recupero delle informazioni, efficienza computazionale e generazione di risposte basate sulle evidenze. La valutazione ha preso in considerazione sia le tradizionali metriche di Information Retrieval sia metriche di tipo LLM-as-a-judge, consentendo di analizzare le fasi di retrieval e generazione da prospettive complementari. I risultati mostrano che la qualità di un sistema RAG dipende non solo dal modello generativo, ma anche dai componenti precedenti della pipeline, come il parsing, il chunking, la rappresentazione indicizzata, la strategia di retrieval e la configurazione dell’indice. Il sistema finale privilegia un’elevata copertura delle evidenze, una bassa latenza nel recupero delle informazioni e la fedeltà delle risposte, evitando componenti il cui costo computazionale non è giustificato da miglioramenti qualitativi costanti. Poiché non era disponibile un dataset di riferimento annotato manualmente, parte della valutazione si è basata su risposte di riferimento generate sinteticamente. Per questo motivo, i punteggi riportati devono essere interpretati principalmente come indicatori comparativi tra diverse configurazioni, piuttosto che come misure assolute delle prestazioni del sistema. Nel complesso, la tesi dimostra come un sistema RAG specializzato per un dominio specifico possa essere progettato e valutato attraverso una metodologia sperimentale strutturata, al fine di fornire un question answering affidabile basato sui documenti.
Retrieval-Augmented Generation for football-oriented documents
ZACCHETTI, SIMONE
2025/2026
Abstract
This thesis presents the design, implementation, and evaluation of a Retrieval-Augmented Generation (RAG) system for question answering over football-oriented technical, regula- tory, and documentary sources. The work was developed in an industrial context, where users needed a more flexible way to access large collections of heterogeneous documents through natural language queries. The proposed system follows a modular RAG pipeline composed of document parsing, chunking, embedding generation, vector database indexing, retrieval, and answer gener- ation. Different configurations were evaluated in order to select a final setup that bal- ances retrieval quality, computational efficiency, and grounded answer generation. The evaluation considered both traditional Information Retrieval metrics and LLM-as-a-judge metrics, allowing the retrieval and generation stages to be analysed from complementary perspectives. The results show that the quality of a RAG system depends not only on the generative model, but also on earlier pipeline components such as parsing, chunking, indexed repre- sentation, retrievalstrategy, andindexconfiguration. Thefinalsystemprioritizesevidence coverage, low retrieval latency, and answer faithfulness, while avoiding components whose computational cost is not justified by consistent quality improvements. Since no manually annotated ground-truth dataset was available, part of the evaluation relied on synthetically generated reference answers. For this reason, the reported scores should be interpreted mainly as comparative indicators between configurations rather than as absolute measures of system performance. Overall, the thesis demonstrates how a domain-specific RAG system can be built and evaluated through a structured experimental methodology for reliable document-based question answering.| File | Dimensione | Formato | |
|---|---|---|---|
|
2025_07_Zacchetti.pdf
non accessibile
Descrizione: testo della tesi
Dimensione
2.17 MB
Formato
Adobe PDF
|
2.17 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260841