The widespread adoption of GPU sharing in cloud computing environments, research clusters, and multi-user workstations has introduced a significant security concern: when multiple users execute workloads concurrently on the same physical GPU, shared microarchitectural resources become a potential source of information leakage. This thesis presents TEXAS, a contention-based side-channel attack deploying two spy kernels on distinct hardware pipelines of NVIDIA GPUs: the Texture Units pipeline (TEX spy kernel) and the asynchronous global-to-shared memory copy pipeline (Async spy kernel), introduced in NVIDIA's Ampere architecture. Both spy kernels operate entirely from unprivileged user-level code, requiring no elevated permissions and no modifications to the victim's execution environment. TEXAS incorporates a dual-stream classification framework that fuses the timing traces of both spy kernels into a unified BiLSTM-based classifier. By combining signals from two independent pipelines, the framework achieves higher discriminative power than either spy kernel used in isolation. The classifier is structured hierarchically: it first identifies the general workload category — among CNNs, Transformers, RNNs, diffusion models, and cryptocurrency mining — and then refines the classification to the model family and, where possible, the specific model. For CNNs, the framework additionally attempts to infer architectural properties through regression on the collected latency traces. TEXAS is evaluated on three NVIDIA GPUs — RTX 3050, RTX 4070, and RTX 5070 — using 30 models across five categories. The dual-stream classifier consistently achieves weighted F1 scores above 0.90 at the category level across all tested configurations, demonstrating generalization across GPU architectures when retrained on data from the target device. At the model level, F1 scores range from 0.63 for structurally similar models to 0.96 for architecturally distinct ones, establishing that fusing signals from multiple independent pipelines provides a more robust fingerprinting primitive than any single-source approach.
La crescente diffusione della condivisione delle GPU in ambienti di cloud computing, cluster di ricerca e workstation multi-utente ha introdotto una rilevante problematica di sicurezza: quando più utenti eseguono workload in modo concorrente sulla stessa GPU fisica, le risorse microarchitetturali condivise possono diventare una fonte di fuga di informazioni. Questa tesi presenta TEXAS, un attacco side-channel basato sulla contesa che impiega due spy kernel su pipeline hardware distinte delle GPU NVIDIA: la pipeline delle Texture Units (TEX spy kernel) e la pipeline di copia asincrona dalla memoria globale alla memoria condivisa (Async spy kernel), introdotta nell'architettura Ampere di NVIDIA. Entrambi gli spy kernel operano con privilegi utente standard, senza richiedere permessi elevati né modifiche all'ambiente di esecuzione della vittima. TEXAS incorpora un framework di classificazione dual-stream che combina le tracce temporali dei due spy kernel in un classificatore unificato basato su BiLSTM. Unendo i segnali di due pipeline indipendenti, il framework ottiene un potere discriminativo superiore rispetto a ciascun spy kernel usato singolarmente. Il classificatore è organizzato gerarchicamente: identifica prima la categoria del workload — tra CNN, Transformer, RNN, modelli di diffusione e mining di criptovalute — e affina poi la classificazione fino alla famiglia di modelli e, quando possibile, al modello specifico. Per le CNN, tenta inoltre di inferire proprietà architetturali tramite regressione sulle tracce di latenza raccolte. TEXAS è valutato su tre GPU NVIDIA — RTX 3050, RTX 4070 e RTX 5070 — su 30 modelli in cinque categorie. Il classificatore dual-stream raggiunge F1 score pesati superiori a 0.90 a livello di categoria su tutte le configurazioni testate, dimostrando la generalizzazione tra architetture diverse quando riaddestrato sul dispositivo target. A livello di modello, gli F1 score variano da 0.63 per modelli strutturalmente simili a 0.96 per modelli architetturalmente distinti, confermando che combinare segnali da più pipeline indipendenti produce una primitiva di fingerprinting più robusta rispetto a qualsiasi approccio a sorgente singola.
TEXAS: probing workloads in shared GPU environments through TEXture and ASync pipeline side channels
MOTTA, SAMUELE
2025/2026
Abstract
The widespread adoption of GPU sharing in cloud computing environments, research clusters, and multi-user workstations has introduced a significant security concern: when multiple users execute workloads concurrently on the same physical GPU, shared microarchitectural resources become a potential source of information leakage. This thesis presents TEXAS, a contention-based side-channel attack deploying two spy kernels on distinct hardware pipelines of NVIDIA GPUs: the Texture Units pipeline (TEX spy kernel) and the asynchronous global-to-shared memory copy pipeline (Async spy kernel), introduced in NVIDIA's Ampere architecture. Both spy kernels operate entirely from unprivileged user-level code, requiring no elevated permissions and no modifications to the victim's execution environment. TEXAS incorporates a dual-stream classification framework that fuses the timing traces of both spy kernels into a unified BiLSTM-based classifier. By combining signals from two independent pipelines, the framework achieves higher discriminative power than either spy kernel used in isolation. The classifier is structured hierarchically: it first identifies the general workload category — among CNNs, Transformers, RNNs, diffusion models, and cryptocurrency mining — and then refines the classification to the model family and, where possible, the specific model. For CNNs, the framework additionally attempts to infer architectural properties through regression on the collected latency traces. TEXAS is evaluated on three NVIDIA GPUs — RTX 3050, RTX 4070, and RTX 5070 — using 30 models across five categories. The dual-stream classifier consistently achieves weighted F1 scores above 0.90 at the category level across all tested configurations, demonstrating generalization across GPU architectures when retrained on data from the target device. At the model level, F1 scores range from 0.63 for structurally similar models to 0.96 for architecturally distinct ones, establishing that fusing signals from multiple independent pipelines provides a more robust fingerprinting primitive than any single-source approach.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Motta.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
14.43 MB
Formato
Adobe PDF
|
14.43 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Motta_Executive Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
2.16 MB
Formato
Adobe PDF
|
2.16 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260421