Modern antivirus software heavily relies on static analysis, particularly signature-based detection, to efficiently identify malware. However, commercial antivirus engines typically operate as black-box systems, making their internal detection logic and specific signature patterns inaccessible for analysis. This thesis proposes DAMANI (Deep leArning Models for Antivirus mimicking and sigNature extractIon), a framework designed to approximate antivirus behavior and extract representative malicious byte patterns using deep learning techniques. The framework operates in two stages. First, a supervised 2D-CNN is trained to mimic the decision boundaries of a target antivirus engine, achieving high accuracy in replicating both correct classifications and errors. Subsequently, a Convolutional Autoencoder is employed to generate sparse binary masks that isolate the specific byte sequences most responsible for triggering a malicious classification. The framework is evaluated on three distinct antivirus engines: Microsoft Defender, ClamAV, and Avast Premium Security. Experimental results demonstrate that the extracted signatures are both effective, triggering detection when isolated from the original file, and necessary, as their removal significantly reduces the maliciousness score. A comparative analysis against ClamAV’s open-source database confirms that the extracted patterns possess a high degree of similarity to real-world signatures used in production environments. This framework provides deep insights into the behavior of antivirus engines, identifying the representative byte patterns that are essential for their detection logic. Ultimately, by characterizing the static detection patterns of black-box software, this research enables a comparative analysis of the decision logic used by different vendors.

I software antivirus moderni si basano fortemente sull’analisi statica, in particolare sul rilevamento basato su firme, per identificare in modo efficiente i file malevoli. Tuttavia, gli antivirus commerciali operano tipicamente come sistemi black-box, rendendo inaccessibili all’analisi sia la logica interna di rilevamento sia gli specifici pattern di firma utilizzati. Questa tesi propone DAMANI (Deep leArning Models for Antivirus mimicking and sigNature extractIon), un framework progettato per replicare il comportamento degli antivirus ed estrarre pattern rappresentativi di file malevoli mediante tecniche di deep learning. Il framework opera in due fasi: in primo luogo, una CNN bidimensionale supervisionata viene addestrata per imitare i criteri decisionali di un motore antivirus target, ottenendo un’accuratezza elevata nella replica sia delle classificazioni corrette sia degli errori e, successivamente, viene impiegato un autoencoder convoluzionale per generare maschere binarie sparse, che isolano le specifiche sequenze di byte maggiormente responsabili della classificazione malevola. Il framework è stato valutato su tre distinti antivirus: Microsoft Defender, ClamAV e Avast Premium Security. I risultati sperimentali dimostrano che le firme estratte sono sia efficaci, poiché attivano il rilevamento anche quando isolate dal file originale, sia necessarie, in quanto la loro rimozione riduce significativamente il punteggio di malevolenza. Un’analisi comparativa con il database open-source di ClamAV conferma che i pattern estratti presentano un elevato grado di similarità con le firme utilizzate in ambienti reali. Il framework fornisce una comprensione approfondita del comportamento degli antivirus, identificando i pattern di byte rappresentativi che risultano essenziali per la loro logica di rilevamento. Caratterizzando i pattern di rilevamento statico di software black-box, questa ricerca consente un’analisi comparativa della logica decisionale adottata da diversi vendor.

DAMANI: Deep leArning Models for Antivirus mimicking and sigNature extractIon

FILIPPI, NICOLE
2025/2026

Abstract

Modern antivirus software heavily relies on static analysis, particularly signature-based detection, to efficiently identify malware. However, commercial antivirus engines typically operate as black-box systems, making their internal detection logic and specific signature patterns inaccessible for analysis. This thesis proposes DAMANI (Deep leArning Models for Antivirus mimicking and sigNature extractIon), a framework designed to approximate antivirus behavior and extract representative malicious byte patterns using deep learning techniques. The framework operates in two stages. First, a supervised 2D-CNN is trained to mimic the decision boundaries of a target antivirus engine, achieving high accuracy in replicating both correct classifications and errors. Subsequently, a Convolutional Autoencoder is employed to generate sparse binary masks that isolate the specific byte sequences most responsible for triggering a malicious classification. The framework is evaluated on three distinct antivirus engines: Microsoft Defender, ClamAV, and Avast Premium Security. Experimental results demonstrate that the extracted signatures are both effective, triggering detection when isolated from the original file, and necessary, as their removal significantly reduces the maliciousness score. A comparative analysis against ClamAV’s open-source database confirms that the extracted patterns possess a high degree of similarity to real-world signatures used in production environments. This framework provides deep insights into the behavior of antivirus engines, identifying the representative byte patterns that are essential for their detection logic. Ultimately, by characterizing the static detection patterns of black-box software, this research enables a comparative analysis of the decision logic used by different vendors.
ING - Scuola di Ingegneria Industriale e dell'Informazione
26-mar-2026
2025/2026
I software antivirus moderni si basano fortemente sull’analisi statica, in particolare sul rilevamento basato su firme, per identificare in modo efficiente i file malevoli. Tuttavia, gli antivirus commerciali operano tipicamente come sistemi black-box, rendendo inaccessibili all’analisi sia la logica interna di rilevamento sia gli specifici pattern di firma utilizzati. Questa tesi propone DAMANI (Deep leArning Models for Antivirus mimicking and sigNature extractIon), un framework progettato per replicare il comportamento degli antivirus ed estrarre pattern rappresentativi di file malevoli mediante tecniche di deep learning. Il framework opera in due fasi: in primo luogo, una CNN bidimensionale supervisionata viene addestrata per imitare i criteri decisionali di un motore antivirus target, ottenendo un’accuratezza elevata nella replica sia delle classificazioni corrette sia degli errori e, successivamente, viene impiegato un autoencoder convoluzionale per generare maschere binarie sparse, che isolano le specifiche sequenze di byte maggiormente responsabili della classificazione malevola. Il framework è stato valutato su tre distinti antivirus: Microsoft Defender, ClamAV e Avast Premium Security. I risultati sperimentali dimostrano che le firme estratte sono sia efficaci, poiché attivano il rilevamento anche quando isolate dal file originale, sia necessarie, in quanto la loro rimozione riduce significativamente il punteggio di malevolenza. Un’analisi comparativa con il database open-source di ClamAV conferma che i pattern estratti presentano un elevato grado di similarità con le firme utilizzate in ambienti reali. Il framework fornisce una comprensione approfondita del comportamento degli antivirus, identificando i pattern di byte rappresentativi che risultano essenziali per la loro logica di rilevamento. Caratterizzando i pattern di rilevamento statico di software black-box, questa ricerca consente un’analisi comparativa della logica decisionale adottata da diversi vendor.
File allegati
File Dimensione Formato  
2026_03_Filippi_Executive Summary.pdf

accessibile in internet per tutti

Descrizione: Executive Summary
Dimensione 521.31 kB
Formato Adobe PDF
521.31 kB Adobe PDF Visualizza/Apri
2026_03_Filippi_Tesi.pdf

accessibile in internet per tutti

Descrizione: Tesi
Dimensione 2.54 MB
Formato Adobe PDF
2.54 MB Adobe PDF Visualizza/Apri

I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10589/251261