Within large institutions, the identification of academic experts and their collaboration networks is a complex task, as relevant information is distributed across heterogeneous data sources with different structures and interfaces. Recent advances in Artificial Intelligence and Large Language Models (LLMs) can facilitate the access to such a fragmented body of knowledge through natural language interfaces, to gather a unified view of researchers and their expertise. Indeed, it is necessary to properly exploit such tools, while handling their limitations, such as outdated information, hallucination, and lack of access to domain-specific data. This thesis proposes FindMyExpert, an intelligent conversational system designed to support faculty at Politecnico di Milano in exploring collaboration networks and identifying researchers with relevant expertise. The system provides a unified natural language interface over multiple institutional sources, including publication repositories, thesis archives, and international agreement records. A data integration pipeline approach collects information into a centralized document database. A tool-based retrieval framework enables structured access to data by exposing each source through dedicated tools. An LLM selects the appropriate tools according to user requests and generating responses grounded in retrieved information. The system supports multi-turn conversations, allowing users to iteratively refine requests. The proposed system is evaluated through experiments analyzing the impact of model size, quantization, reasoning mode, serving framework, tool provisioning strategy, and orchestration framework on tool invocation accuracy, response quality, and latency. Results show that tool-calling performance improves with larger models, that reasoning limited to tool selection provides a favorable accuracy-efficiency trade-off, and that the serving framework affects both latency and output reliability. Overall, the system demonstrates effective tool-based academic information retrieval and provides a foundation in real-world institutional environments.
L’operazione di identificare chi, all’interno di grandi istituzioni accademiche, ha collaborazioni con altre istituzioni e quali aree di ricerca siano coperte è complessa, poiché le informazioni rilevanti sono distribuite su fonti dati eterogenee, caratterizzate da strutture e interfacce differenti e soggette a rapida evoluzione. I recenti progressi nell’ambito dell’Intelligenza Artificiale e dei Large Language Model (LLM) possono facilitare l’accesso a un insieme frammentato di conoscenze tramite interfacce in linguaggio naturale. Per poterli utilizzare è però necessario gestirne i limiti, tra cui informazioni obsolete, allucinazioni. Questa tesi propone FindMyExpert, un sistema conversazionale intelligente a supporto dell’esplorazione delle reti di collaborazione e dell’individuazione di persone con le competenze desiderate. Il sistema offre un’interfaccia unificata in linguaggio naturale per interrogare diverse fonti istituzionali, tra cui archivi di pubblicazioni, tesi e accordi internazionali. Le informazioni raccolte sono integrate in una base dati documentale centralizzata ed il framework è costituito da un insieme di strumenti ciascuno dei quali dedicati ad una fonte dati specifica. Un LLM processa la richiesta dell’utente e seleziona gli strumenti più appropriati, generando una risposta fondata sulle informazioni recuperate. Il sistema supporta conversazioni multi-turno, consentendo di affinare iterativamente le richieste e risposte. Abbiamo valutato il sistema attraverso una campagna sperimentale che analizza numerosi aspetti, tra cui l’impatto della dimensione del modello, la modalità di ragionamento, la strategia di provisioning degli strumenti, valutandone gli effetti su accuratezza dell’invocazione, qualità delle risposte e latenza. I risultati mostrano che le prestazioni migliorano con modelli più grandi, che un ragionamento limitato alla selezione degli strumenti offre un buon compromesso tra accuratezza ed efficienza. Nel complesso, il sistema dimostra la fattibilità della soluzione, la qualità dell’architettura proposta e costituisce un buon prototipo per sviluppi futuri.
FindMyExpert: A Conversational Agent for Academic Collaboration Discovery
CARBAJAL SERRANO, CESAR ADRIAN
2025/2026
Abstract
Within large institutions, the identification of academic experts and their collaboration networks is a complex task, as relevant information is distributed across heterogeneous data sources with different structures and interfaces. Recent advances in Artificial Intelligence and Large Language Models (LLMs) can facilitate the access to such a fragmented body of knowledge through natural language interfaces, to gather a unified view of researchers and their expertise. Indeed, it is necessary to properly exploit such tools, while handling their limitations, such as outdated information, hallucination, and lack of access to domain-specific data. This thesis proposes FindMyExpert, an intelligent conversational system designed to support faculty at Politecnico di Milano in exploring collaboration networks and identifying researchers with relevant expertise. The system provides a unified natural language interface over multiple institutional sources, including publication repositories, thesis archives, and international agreement records. A data integration pipeline approach collects information into a centralized document database. A tool-based retrieval framework enables structured access to data by exposing each source through dedicated tools. An LLM selects the appropriate tools according to user requests and generating responses grounded in retrieved information. The system supports multi-turn conversations, allowing users to iteratively refine requests. The proposed system is evaluated through experiments analyzing the impact of model size, quantization, reasoning mode, serving framework, tool provisioning strategy, and orchestration framework on tool invocation accuracy, response quality, and latency. Results show that tool-calling performance improves with larger models, that reasoning limited to tool selection provides a favorable accuracy-efficiency trade-off, and that the serving framework affects both latency and output reliability. Overall, the system demonstrates effective tool-based academic information retrieval and provides a foundation in real-world institutional environments.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Carbajal_Tesi.pdf
accessibile in internet per tutti
Descrizione: Thesis document
Dimensione
1.7 MB
Formato
Adobe PDF
|
1.7 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Carbajal_Executive Summary.pdf
accessibile in internet per tutti
Descrizione: Executive summary
Dimensione
451.33 kB
Formato
Adobe PDF
|
451.33 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260712