The increasing cost and duration of the drug discovery process have accelerated the adoption of computational approaches capable of reducing the experimental burden associated with the identification of novel therapeutic compounds. In this context, High-Throughput Virtual Screening (HTVS) has emerged as a key strategy for prioritizing promising candidates by simulating protein-ligand interactions at scale. This approach relies on molecular docking engines to generate and evaluate a large number of potential binding conformations. The generation (or "sampling") phase, in particular, needs to be fast and efficient, while ensuring that the sampled conformations are of sufficient quality to be accurately evaluated and classified. This requirement has led to the widespread adoption of heuristic algorithms that use simplified Scoring Functions (SFs) as a feedback loop, which open up a compromise between speed and predictive accuracy. This thesis investigates the feasibility of incorporating well-known, high-quality SFs into the molecular sampling phase of a production-grade HTVS platform, specifically the LiGen (Ligand Generator) suite. To overcome the resulting computational demands and ensure scalability on HighPerformance Computing (HPC) infrastructures, we applied optimization strategies and hardware acceleration, exploring also the Approximate Computing (AC) paradigm. The developed solutions where then subjected to performance evaluations on CPU and GPU platforms via hardware profiling and qualitatively validated using standardized benchmarks, such as the CASF framework. Based on the final results, the most promising solution is the application of AC techniques to a Knowledge-Based SF, ensuring excellent performance and scalability and no degradation in predictive quality. The final validation on LiGen produced a promising but double-edged result, highlighting the desired increase in docking quality at the cost of a significant decrease in throughput.
L’aumento di costo e durata del processo di scoperta di farmaci ha accelerato l’adozione di approcci computazionali in grado di ridurre l’onere sperimentale associato all’identificazione di nuovi composti terapeutici. In questo contesto, l’High-Throughput Virtual Screening (HTVS), è emerso come una strategia chiave per priorizzare candidati promettenti simulando interazioni proteina-ligando su larga scala. Questo approccio si basa su algoritmi di docking molecolare per generare e valutare un gran numero di potenziali conformazioni di legame. La fase di generazione (o di "campionamento"), in particolare, deve essere rapida ed efficiente, garantendo al contempo che le conformazioni campionate siano di qualità sufficiente per essere valutate e classificate con precisione. Questo requisito ha portato all’adozione diffusa di algoritmi euristici che utilizzano Scoring Functions (SFs) semplificate come ciclo di feedback, le quali aprono ad un compromesso tra velocità e accuratezza predittiva. Questa tesi indaga la fattibilità di incorporare SFs note e di alta qualità nella fase di campionamento molecolare di una piattaforma HTVS industriale, nello specifico la suite LiGen (Ligand Generator). Per superare le conseguenti esigenze computazionali e garantire scalabilità su infrastrutture di High-Performance Computing (HPC), abbiamo applicato strategie di ottimizzazione e accelerazione hardware, esplorando anche il paradigma dell’Approximate Computing (AC). Le soluzioni sviluppate sono state quindi sottoposte a valutazioni delle prestazioni su piattaforme CPU e GPU tramite profilazione hardware e validate qualitativamente tramite benchmark standardizzati, come il framework CASF. Basandosi sui risultati finali, la soluzione più promettente è data dall’applicazione di tecniche AC ad una SF Knowledge-Based, garantendo ottime prestazioni e scalabilità e nessun degrado della qualità predittiva. Le validazione finale su LiGen ha invece prodotto un risultato promettente, ma ambivalente, evidenziando l’incremento sperato della qualità di docking al costo di una significativa riduzione del throughput.
Knowledge-based potentials for High-Throughput molecular docking in large-scale Virtual Screening
Tirri, Enrico
2025/2026
Abstract
The increasing cost and duration of the drug discovery process have accelerated the adoption of computational approaches capable of reducing the experimental burden associated with the identification of novel therapeutic compounds. In this context, High-Throughput Virtual Screening (HTVS) has emerged as a key strategy for prioritizing promising candidates by simulating protein-ligand interactions at scale. This approach relies on molecular docking engines to generate and evaluate a large number of potential binding conformations. The generation (or "sampling") phase, in particular, needs to be fast and efficient, while ensuring that the sampled conformations are of sufficient quality to be accurately evaluated and classified. This requirement has led to the widespread adoption of heuristic algorithms that use simplified Scoring Functions (SFs) as a feedback loop, which open up a compromise between speed and predictive accuracy. This thesis investigates the feasibility of incorporating well-known, high-quality SFs into the molecular sampling phase of a production-grade HTVS platform, specifically the LiGen (Ligand Generator) suite. To overcome the resulting computational demands and ensure scalability on HighPerformance Computing (HPC) infrastructures, we applied optimization strategies and hardware acceleration, exploring also the Approximate Computing (AC) paradigm. The developed solutions where then subjected to performance evaluations on CPU and GPU platforms via hardware profiling and qualitatively validated using standardized benchmarks, such as the CASF framework. Based on the final results, the most promising solution is the application of AC techniques to a Knowledge-Based SF, ensuring excellent performance and scalability and no degradation in predictive quality. The final validation on LiGen produced a promising but double-edged result, highlighting the desired increase in docking quality at the cost of a significant decrease in throughput.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Tirri_Thesis.pdf
non accessibile
Descrizione: Thesis
Dimensione
2.55 MB
Formato
Adobe PDF
|
2.55 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Tirri_Executive_Summary.pdf
non accessibile
Descrizione: Executive Summary
Dimensione
513.25 kB
Formato
Adobe PDF
|
513.25 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261328