Medical imaging is a cornerstone of modern healthcare, generating vast repositories of quantitative data encoded in the DICOM (Digital Imaging and Communications in Medicine) standard. While DICOM ensures successful image transmission across heterogeneous systems, its metadata — crucial for secondary research and Artificial Intelligence (AI) development — often remains trapped in local silos. These metadata are plagued by vendorspecific dialects, polysemy, and non-standard abbreviations, creating a profound semantic gap, and traditional mapping techniques, ranging from deterministic lexical lookups to statistical NLP and dense vector embeddings, consistently fail to resolve this ambiguity, due to their lack of contextual clinical reasoning. To overcome these barriers, this thesis presents a fully containerized, scalable software pipeline that transforms raw DICOM metadata into a FAIR (Findable, Accessible, Interoperable, Reusable)-compliant, semantically enriched knowledge graph. Adopting an "enrichment-first" strategy, the system preserves the native hierarchical structure of the imaging data within a graph database (Neo4j) while dynamically linking it to a multiontological backbone (SNOMED CT, NCIt, LOINC, RadLex). The core innovation of this work is the integration of a local large language model (Llama 3) to execute context-aware semantic mapping. By analyzing the full context of DICOM attributes, the generative AI successfully disambiguates complex technical terms and normalizes proprietary dialects. Experimental results demonstrate that the proposed neuro-symbolic pipeline significantly outperforms baseline deterministic and embedding-based methods, achieving a near-total coverage (89.5%) and high confidence precision (93%) for the required clinical metadata. Ultimately, this approach bridges the semantic gap, enabling cross-center interoperability and allowing researchers to perform advanced, hierarchical clinical queries that were previously impossible, thereby unlocking the latent value of medical imaging data lakes.
L’imaging medico rappresenta un pilastro fondamentale della sanità moderna, generando vasti archivi di dati quantitativi codificati secondo lo standard DICOM (Digital Imaging and Communications in Medicine). Sebbene il formato DICOM garantisca la corretta trasmissione delle immagini tra sistemi eterogenei, i suoi metadati—cruciali per la ricerca clinica e lo sviluppo di modelli di Intelligenza Artificiale (IA)—rimangono spesso intrappolati in silos proprietari. Tali metadati sono caratterizzati da dialetti specifici dei produttori, polisemia e abbreviazioni non standardizzate, creando un profondo divario semantico. Le tecniche di mapping tradizionali, dalle ricerche lessicali deterministiche ai modelli statistici e agli embedding vettoriali, si rivelano insufficienti per risolvere questa ambiguità, mancando di una reale capacità di ragionamento clinico contestuale. Per superare questi limiti, questa tesi presenta una pipeline software scalabile e containerizzata in grado di trasformare i metadati DICOM grezzi in un Knowledge Graph semanticamente arricchito e conforme ai principi FAIR (Rintracciabile, Accessibile, Interoperabile, Riutilizzabile). Adottando una strategia "Enrichment-First", il sistema preserva la struttura gerarchica nativa dei dati all’interno di un database a grafo (Neo4j), collegandola dinamicamente a un’infrastruttura multi-ontologica (SNOMED CT, NCIt, LOINC, RadLex). L’innovazione centrale del lavoro consiste nell’integrazione di un Large Language Model (Llama 3.3) [40] locale per l’esecuzione di un mapping semantico sensibile al contesto. Analizzando il contesto completo degli attributi DICOM, l’IA generativa è in grado di disambiguare termini tecnici complessi e normalizzare i dialetti proprietari. I risultati sperimentali dimostrano che questa pipeline neuro-simbolica supera nettamente i metodi deterministici e basati su embedding, ottenendo una copertura quasi totale dei metadati clinici richiesti. In conclusione, questo approccio colma il divario semantico, garantendo l’interoperabilità tra centri clinici diversi e abilitando interrogazioni gerarchiche avanzate, sbloccando così il potenziale inespresso dei data lake di imaging medico.
Semantically enriched knowledge graph for medical imaging: an ontology-based approach using LLM mappings
MARTINIS, FEDERICO
2025/2026
Abstract
Medical imaging is a cornerstone of modern healthcare, generating vast repositories of quantitative data encoded in the DICOM (Digital Imaging and Communications in Medicine) standard. While DICOM ensures successful image transmission across heterogeneous systems, its metadata — crucial for secondary research and Artificial Intelligence (AI) development — often remains trapped in local silos. These metadata are plagued by vendorspecific dialects, polysemy, and non-standard abbreviations, creating a profound semantic gap, and traditional mapping techniques, ranging from deterministic lexical lookups to statistical NLP and dense vector embeddings, consistently fail to resolve this ambiguity, due to their lack of contextual clinical reasoning. To overcome these barriers, this thesis presents a fully containerized, scalable software pipeline that transforms raw DICOM metadata into a FAIR (Findable, Accessible, Interoperable, Reusable)-compliant, semantically enriched knowledge graph. Adopting an "enrichment-first" strategy, the system preserves the native hierarchical structure of the imaging data within a graph database (Neo4j) while dynamically linking it to a multiontological backbone (SNOMED CT, NCIt, LOINC, RadLex). The core innovation of this work is the integration of a local large language model (Llama 3) to execute context-aware semantic mapping. By analyzing the full context of DICOM attributes, the generative AI successfully disambiguates complex technical terms and normalizes proprietary dialects. Experimental results demonstrate that the proposed neuro-symbolic pipeline significantly outperforms baseline deterministic and embedding-based methods, achieving a near-total coverage (89.5%) and high confidence precision (93%) for the required clinical metadata. Ultimately, this approach bridges the semantic gap, enabling cross-center interoperability and allowing researchers to perform advanced, hierarchical clinical queries that were previously impossible, thereby unlocking the latent value of medical imaging data lakes.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Martinis_ES.pdf
solo utenti autorizzati a partire dal 02/03/2027
Descrizione: Executive Summary
Dimensione
1.1 MB
Formato
Adobe PDF
|
1.1 MB | Adobe PDF | Visualizza/Apri |
|
2026_03_Martinis_Thesis.pdf
solo utenti autorizzati a partire dal 02/03/2027
Descrizione: Thesis
Dimensione
2.75 MB
Formato
Adobe PDF
|
2.75 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/252819