Large Language Models (LLMs) are computational models based on artificial intelligence designed to understand, interpret, and generate natural language text, and are today employed in a growing number of scientific sectors. At the same time, cancer represents one of the leading public health problems worldwide, making the need to identify new therapeutic strategies. This work pursues a dual objective: to assess the ability of LLMs to support the scientific research pipeline and, as a case study, to conduct a systematic review on the use of non-thermal, non-ionizing electromagnetic fields as a selective monotherapy for tumor cells in vitro. The research was structured into four consecutive phases: literature search, data extraction, data analysis, and hypothesis generation. Each of them was conducted by the researcher in two independent modalities, one depending only on conventional tools and the other with the support of the LLMs. The two approaches were compared at the end of each phase using evaluation metrics specifically identified for the intention. Three state of the art general-purpose LLMs were tested: Claude Opus 4.7, GPT-5.5 Thinking, and Gemini 3.1 Pro. In specific phases, additional tools were also evaluated: SciSpace, Asta AI, NotebookLM, and the Analysis agent of Edison Scientific. Claude Opus 4.7 proved to be the best model. In the literature search phase, its combination with Asta AI – Find papers tool and SciSpace retrieved a considerably more relevant studies than the manual approach. In data extraction, it achieved results comparable to the human Gold Standard. In data analysis, its report was ranked as the best by Gemini 3.1 Pro acting as LLM--as--Judge, surpassing the researcher's report. In the hypothesis generation phase, the human--generated hypothesis outperformed those produced by the LLMs. Regarding the scientific results, the electromagnetic fields analyzed attested to be capable of exercising cytotoxic and antiproliferative effects on tumor cells in vitro. Antitumor efficacy and selectivity were found to be parameter dependent rather than field dependent. However, the scarcity of available studies and the heterogeneity of the analyzed endpoints and tested parameters made it difficult to draw definitive conclusions on the mechanisms of action. Nevertheless, a mechanistic hypothesis integrating the available evidence into a coherent framework was proposed. In the end, LLMs emerged as effective tools for supporting the researcher throughout the entire pipeline, substantially reducing the execution time, while EMFs show a real therapeutic potential, but a deeper understanding of the mechanisms of action is needed to support their future clinical application.
I Large Language Models (LLMs) sono modelli computazionali basati sull’intelligenza artificiale, progettati per comprendere, interpretare e generare testo in linguaggio naturale, e sono ad oggi utilizzati in un numero sempre maggiore di settori industriali e scientifici. Allo stesso tempo, il cancro rappresenta uno dei principali problemi di salute pubblica a livello globale, responsabile di quasi una morte su sei nel mondo; per questo motivo la necessità di individuare nuove strategie terapeutiche è più che mai urgente. Questo lavoro ha un duplice obiettivo: valutare la capacità degli LLM di supportare le diverse fasi della ricerca scientifica e al tempo stesso, come caso studio, condurre una revisione sistematica della letteratura esistente sull’utilizzo dei campi elettromagnetici non termici e non ionizzanti come monoterapia selettiva per le cellule tumorali in vitro. Il processo di ricerca è stato suddiviso in quattro fasi consecutive: ricerca della letteratura scientifica, estrazione dei dati dagli articoli selezionati, analisi dei dati estratti e generazione di un’ipotesi di ricerca. Ciascuna fase è stata condotta in due modalità indipendenti: nella prima, il ricercatore ha lavorato in modo autonomo, utilizzando gli strumenti tradizionali della ricerca, senza cioè senza ricorrere ad alcuno strumento di intelligenza artificiale; nella seconda, le stesse fasi sono state ripetute con il supporto degli LLMs. Al termine di ciascuna fase, i due approcci sono stati confrontati tra di loro e valutati tramite diverse metriche. Gli strumenti testati includono tre LLM general–purpose allo stato dell’arte, Claude Opus 4.7, GPT-5.5 Thinking eGemini 3.1 Pro, affiancati, nelle fasi specifiche, daSciSpace, AstaAI, NotebookLM e dall’agente Analysis Agent di Edison Scientific. Claude Opus 4.7 si è rivelato il modello migliore in tutte le fasi. Nella fase di ricerca bibliografica, la combinazione di questo LLM, Asta AI – Find Papers e SciSpace ha identificato un numero di studi rilevanti nettamente superiore rispetto alla ricerca manuale. Nella fase di estrazione dati, ha ottenuto risultati paragonabili al Gold Standard umano. Nella fase di analisi, il report da esso generato è stato valutato da Gemini 3.1 Pro, in veste di LLM–as–Judge, come il migliore tra tutti, incluso quello del ricercatore. Nella fase di generazione dell’ipotesi, l’ipotesi del ricercatore ha invece ottenuto una valutazione superiore rispetto a quelle generate dagli LLM. Per quanto riguarda i risultati scientifici, i campi elettromagnetici analizzati si sono dimostrati effettivamente in grado di esercitare effetti citotossici e antiproliferativi sulle cellule tumorali in vitro. Tuttavia, l’efficacia antitumorale e la selettività risultano essere parametro dipendenti e non campo dipendenti: esistono specifiche combinazioni di frequenza, intensità, forma d’onda e durata di esposizione che producono effetti significativi su determinate linee cellulari, mentre parametri diversi possono produrre effetti opposti o nulli. La scarsità degli studi disponibili e l’eterogeneità degli endpoint analizzati e dei parametri testati hanno reso difficile trarre conclusioni definitive sui meccanismi d’azione. È stata tuttavia proposta un’ipotesi meccanicistica che integra le evidenze disponibili in uno schema il più possibile coerente. In conclusione, gli LLM si sono dimostrati in grado di affiancare in modo efficace il ricercatore lungo l’intera pipeline della ricerca scientifica, riducendo in modo sostanziale i tempi di esecuzione di quest’ultima. I campi EMF mostrano un potenziale terapeutico reale, ma è necessaria una comprensione più approfondita dei meccanismi d’azione per supportare una loro futura applicazione clinica.
Evaluating LLMs across the scientific research pipeline: a case study on electromagnetic fields as in vitro cancer monotherapy
DENTI, SILVIA
2025/2026
Abstract
Large Language Models (LLMs) are computational models based on artificial intelligence designed to understand, interpret, and generate natural language text, and are today employed in a growing number of scientific sectors. At the same time, cancer represents one of the leading public health problems worldwide, making the need to identify new therapeutic strategies. This work pursues a dual objective: to assess the ability of LLMs to support the scientific research pipeline and, as a case study, to conduct a systematic review on the use of non-thermal, non-ionizing electromagnetic fields as a selective monotherapy for tumor cells in vitro. The research was structured into four consecutive phases: literature search, data extraction, data analysis, and hypothesis generation. Each of them was conducted by the researcher in two independent modalities, one depending only on conventional tools and the other with the support of the LLMs. The two approaches were compared at the end of each phase using evaluation metrics specifically identified for the intention. Three state of the art general-purpose LLMs were tested: Claude Opus 4.7, GPT-5.5 Thinking, and Gemini 3.1 Pro. In specific phases, additional tools were also evaluated: SciSpace, Asta AI, NotebookLM, and the Analysis agent of Edison Scientific. Claude Opus 4.7 proved to be the best model. In the literature search phase, its combination with Asta AI – Find papers tool and SciSpace retrieved a considerably more relevant studies than the manual approach. In data extraction, it achieved results comparable to the human Gold Standard. In data analysis, its report was ranked as the best by Gemini 3.1 Pro acting as LLM--as--Judge, surpassing the researcher's report. In the hypothesis generation phase, the human--generated hypothesis outperformed those produced by the LLMs. Regarding the scientific results, the electromagnetic fields analyzed attested to be capable of exercising cytotoxic and antiproliferative effects on tumor cells in vitro. Antitumor efficacy and selectivity were found to be parameter dependent rather than field dependent. However, the scarcity of available studies and the heterogeneity of the analyzed endpoints and tested parameters made it difficult to draw definitive conclusions on the mechanisms of action. Nevertheless, a mechanistic hypothesis integrating the available evidence into a coherent framework was proposed. In the end, LLMs emerged as effective tools for supporting the researcher throughout the entire pipeline, substantially reducing the execution time, while EMFs show a real therapeutic potential, but a deeper understanding of the mechanisms of action is needed to support their future clinical application.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Denti_Executive_Summary.pdf
accessibile in internet per tutti a partire dal 02/07/2029
Descrizione: Executive summary
Dimensione
2.55 MB
Formato
Adobe PDF
|
2.55 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Denti_Thesis.pdf
accessibile in internet per tutti a partire dal 02/07/2029
Descrizione: Thesis
Dimensione
6.99 MB
Formato
Adobe PDF
|
6.99 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/259862