Deep neural networks (DNNs) demand immense computational resources, with off-chip DRAM accesses constituting the primary energy bottleneck during inference. While standard layer-by-layer execution requires redundant DRAM transfers for intermediate activations, layer fusion mitigates this by keeping intermediate data on-chip through interleaved scheduling. However, optimizing fused-layer execution introduces cascading inter-layer dependencies and drastically inflates the mapping search space, making the systematic exploration exceptionally challenging. This thesis formally characterizes the complex design space of multi-layer fused DNN execution on spatial accelerators. To systematically navigate this space, an analytical modeling framework is developed. By introducing dimension aliasing and decoupling, the methodology unifies multi-layer computations into a single loop nest while strictly preserving independent, per-layer mapping freedom. Furthermore, an exact producer- consumer constraint-based pruning strategy rigorously eliminates infeasible mappings. The analytical cost model is extended to compute per-layer memory operations, latency, and energy, and is structurally validated against the DepFiN hardware accelerator. Extensive exploration using DepFiN-like and Eyeriss-like architectures was conducted across activation-dominant and weight-dominant workloads. The findings reveal that while full fusion successfully reduces total DRAM traffic by 25% to 95% and significantly lowers latency for activation-dominant workloads (up to 42–48% of savings), it consistently incurs a severe energy penalty for weight-dominant networks due to exponentially enlarged on-chip memory requirements. Conversely, partial pairwise fusion emerges as a highly practical compromise, capturing nearly all latency benefits while substantially mitigating the on-chip memory overhead. Ultimately, this work demonstrates that layer fusion is primarily a workload-dependent latency optimization rather than a universal energy solution, providing designers with a scalable, formalized methodology to systematically navigate these complex architectural trade-offs.
Le reti neurali profonde (DNN) richiedono immense risorse computazionali, con gli accessi alla DRAM off-chip che costituiscono il principale collo di bottiglia energetico durante l'inferenza. Mentre l'esecuzione standard livello per livello (layer-by-layer) richiede trasferimenti DRAM ridondanti per le attivazioni intermedie, la fusione dei layer (layer fusion) mitiga questo problema mantenendo i dati intermedi on-chip attraverso uno scheduling interleaved. Tuttavia, l'ottimizzazione dell'esecuzione a layer fusi introduce dipendenze a cascata tra i layer e ingrandisce drasticamente lo spazio di ricerca del mapping, rendendone l'esplorazione sistematica eccezionalmente complessa. Questa tesi caratterizza formalmente il complesso spazio di progetto dell'esecuzione di DNN a più layer fusi su acceleratori spaziali. Per navigare sistematicamente in questo spazio, viene sviluppato un framework di modellazione analitica. Attraverso l'introduzione dell'aliasing e del disaccoppiamento delle dimensioni, la metodologia unifica i calcoli multi-layer in un singolo loop nest, preservando rigorosamente la libertà di mapping indipendente per ogni layer. Inoltre, un'esatta strategia di pruning basata sui vincoli produttore-consumatore elimina rigorosamente i mapping irrealizzabili. Il modello analitico di costo è stato esteso per calcolare le operazioni di memoria, la latenza e l'energia per singolo layer, ed è validato strutturalmente rispetto all'acceleratore hardware DepFiN. È stata condotta un'ampia esplorazione utilizzando architetture simili a DepFiN ed Eyeriss su workload dominati dalle attivazioni e dominati dai pesi. I risultati rivelano che, sebbene la fusione completa riduca con successo il traffico DRAM totale dal 25% al 95% e diminuisca significativamente la latenza per i workload dominati dalle attivazioni (con risparmi fino al 42-48%), essa comporta costantemente una grave penalizzazione energetica per le reti dominate dai pesi, a causa dell'aumento esponenziale dei requisiti di memoria on-chip. Al contrario, la fusione parziale a coppie emerge come un compromesso altamente pratico, catturando quasi tutti i benefici in termini di latenza e mitigando sostanzialmente l'overhead di memoria on-chip. In definitiva, questo lavoro dimostra che la fusione dei layer è principalmente un'ottimizzazione della latenza dipendente dal workload, piuttosto che una soluzione energetica universale, fornendo ai progettisti una metodologia formalizzata e scalabile per navigare sistematicamente questi complessi trade-off architetturali.
A Framework for exploring fused layers in DNN dataflows for spatial accelerators
Curro', Matteo Salvatore
2024/2025
Abstract
Deep neural networks (DNNs) demand immense computational resources, with off-chip DRAM accesses constituting the primary energy bottleneck during inference. While standard layer-by-layer execution requires redundant DRAM transfers for intermediate activations, layer fusion mitigates this by keeping intermediate data on-chip through interleaved scheduling. However, optimizing fused-layer execution introduces cascading inter-layer dependencies and drastically inflates the mapping search space, making the systematic exploration exceptionally challenging. This thesis formally characterizes the complex design space of multi-layer fused DNN execution on spatial accelerators. To systematically navigate this space, an analytical modeling framework is developed. By introducing dimension aliasing and decoupling, the methodology unifies multi-layer computations into a single loop nest while strictly preserving independent, per-layer mapping freedom. Furthermore, an exact producer- consumer constraint-based pruning strategy rigorously eliminates infeasible mappings. The analytical cost model is extended to compute per-layer memory operations, latency, and energy, and is structurally validated against the DepFiN hardware accelerator. Extensive exploration using DepFiN-like and Eyeriss-like architectures was conducted across activation-dominant and weight-dominant workloads. The findings reveal that while full fusion successfully reduces total DRAM traffic by 25% to 95% and significantly lowers latency for activation-dominant workloads (up to 42–48% of savings), it consistently incurs a severe energy penalty for weight-dominant networks due to exponentially enlarged on-chip memory requirements. Conversely, partial pairwise fusion emerges as a highly practical compromise, capturing nearly all latency benefits while substantially mitigating the on-chip memory overhead. Ultimately, this work demonstrates that layer fusion is primarily a workload-dependent latency optimization rather than a universal energy solution, providing designers with a scalable, formalized methodology to systematically navigate these complex architectural trade-offs.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_03_Curro_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
733.13 kB
Formato
Adobe PDF
|
733.13 kB | Adobe PDF | Visualizza/Apri |
|
2026_03_Curro_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
5.83 MB
Formato
Adobe PDF
|
5.83 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253748