Retrieval-Augmented Generation (RAG) systems have emerged as an approach to en- hance the factual reliability and domain adaptation of large language models by integrat- ing external knowledge retrieval with generative capabilities. While recent advances have demonstrated the effectiveness of this paradigm, many implementations remain tightly coupled and lacking production-oriented architectural considerations. This thesis presents the design, implementation, and validation of a modular RAG architecture explicitly structured into three decoupled layers: ingestion, retrieval and generation. The system leverages a vector database implemented using Milvus for scalable semantic indexing and similarity search, combined with a multi-stage retrieval pipeline that incorporates dense vector search and re-ranking mechanisms. The generation layer is built using Google Con- versational Agents, a recent framework born as an evolution of Google Dialogflow CX that is capable of extracting intents, invocating external tools and handling several sub-tasks in an efficient way. The research investigates how engineering choices at each stage of the RAG pipeline affect recall, latency, and answer faithfulness in a domain-specific conver- sational setting. In the absence of publicly available benchmarks, a structured synthetic evaluation dataset was constructed and validated to assess retrieval effectiveness and response quality. Experimental results demonstrate that the selected indexing configura- tion achieves a favorable recall–latency trade-off, while re-ranking and structured context formatting significantly improve precision and reduce hallucination frequency. Beyond performance evaluation, the proposed architecture integrates governance-oriented design principles, including modular API isolation and retrieval traceability, thereby addressing production-readiness considerations often overlooked in experimental RAG implementa- tions. The findings confirm that the effectiveness of a RAG system depends not solely on the underlying language model, but critically on the architectural and engineering deci- sions governing indexing, retrieval, and orchestration. This work contributes a structured, empirically validated, and governance-aware framework for building scalable and reliable domain-specific conversational AI systems.
I sistemi di Retrieval-Augmented Generation (RAG) si sono affermati come un approccio efficace per migliorare l’affidabilità fattuale e l’adattamento a domini specifici dei mod- elli linguistici di grandi dimensioni, integrando meccanismi di recupero di conoscenza esterna con capacità generative. Sebbene i recenti sviluppi abbiano dimostrato l’efficacia di questo paradigma, molte implementazioni risultano ancora fortemente accoppiate e prive di adeguate considerazioni architetturali orientate alla produzione. La presente tesi propone la progettazione, l’implementazione e la validazione di un’architettura RAG modulare, esplicitamente strutturata in tre livelli disaccoppiati: ingestione, retrieval e generazione. Il sistema utilizza un database vettoriale implementato tramite Milvus per l’indicizzazione semantica scalabile e la ricerca per similarità, integrato con una pipeline di retrieval multi-stage che combina ricerca vettoriale densa e meccanismi di riordina- mento (re-ranking). Il livello di generazione è realizzato mediante Google Conversational Agents, framework evoluzione di Google Dialogflow CX, in grado di estrarre intenti, in- vocare strumenti esterni e gestire sotto-attività in modo efficiente tramite un componente di orchestrazione dedicato. La ricerca analizza come le scelte ingegneristiche adottate in ciascuna fase della pipeline RAG influenzino metriche quali recall, latenza e fedeltà delle risposte in un contesto conversazionale domain-specific. In assenza di benchmark pubblicamente disponibili, è stato costruito e validato un dataset sintetico strutturato per valutare l’efficacia del retrieval e la qualità delle risposte generate. I risultati sperimentali mostrano che la configurazione di indicizzazione selezionata garantisce un compromesso favorevole tra recall e latenza, mentre l’introduzione del re-ranking e di una strutturazione esplicita del contesto migliora significativamente la precisione e riduce la frequenza di allu- cinazioni. Oltre alla valutazione delle prestazioni, l’architettura proposta integra principi di progettazione orientati alla governance, tra cui l’isolamento modulare delle API e la tracciabilità del processo di retrieval, affrontando così aspetti di production-readiness spesso trascurati nelle implementazioni RAG di natura sperimentale. I risultati confer- mano che l’efficacia di un sistema RAG non dipende esclusivamente dall’LLM sottostante, ma in misura determinante dalle decisioni architetturali e ingegneristiche relative a indi- cizzazione, retrieval e orchestrazione. Il lavoro contribuisce pertanto con un framework strutturato, validato empiricamente e orientato alla governance per la realizzazione di sistemi conversazionali scalabili, affidabili e specifici per dominio.
Knowledge access in enterprise environments: a RAG-based conversational architecture
GENTILI, FILIPPO
2024/2025
Abstract
Retrieval-Augmented Generation (RAG) systems have emerged as an approach to en- hance the factual reliability and domain adaptation of large language models by integrat- ing external knowledge retrieval with generative capabilities. While recent advances have demonstrated the effectiveness of this paradigm, many implementations remain tightly coupled and lacking production-oriented architectural considerations. This thesis presents the design, implementation, and validation of a modular RAG architecture explicitly structured into three decoupled layers: ingestion, retrieval and generation. The system leverages a vector database implemented using Milvus for scalable semantic indexing and similarity search, combined with a multi-stage retrieval pipeline that incorporates dense vector search and re-ranking mechanisms. The generation layer is built using Google Con- versational Agents, a recent framework born as an evolution of Google Dialogflow CX that is capable of extracting intents, invocating external tools and handling several sub-tasks in an efficient way. The research investigates how engineering choices at each stage of the RAG pipeline affect recall, latency, and answer faithfulness in a domain-specific conver- sational setting. In the absence of publicly available benchmarks, a structured synthetic evaluation dataset was constructed and validated to assess retrieval effectiveness and response quality. Experimental results demonstrate that the selected indexing configura- tion achieves a favorable recall–latency trade-off, while re-ranking and structured context formatting significantly improve precision and reduce hallucination frequency. Beyond performance evaluation, the proposed architecture integrates governance-oriented design principles, including modular API isolation and retrieval traceability, thereby addressing production-readiness considerations often overlooked in experimental RAG implementa- tions. The findings confirm that the effectiveness of a RAG system depends not solely on the underlying language model, but critically on the architectural and engineering deci- sions governing indexing, retrieval, and orchestration. This work contributes a structured, empirically validated, and governance-aware framework for building scalable and reliable domain-specific conversational AI systems.| File | Dimensione | Formato | |
|---|---|---|---|
|
Executive_Summary___Filippo_Gentili.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Executive summary
Dimensione
877.22 kB
Formato
Adobe PDF
|
877.22 kB | Adobe PDF | Visualizza/Apri |
|
Thesis_Filippo_Gentili.pdf
accessibile in internet solo dagli utenti autorizzati
Descrizione: Thesis
Dimensione
2.92 MB
Formato
Adobe PDF
|
2.92 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253475