The end of Dennard scaling has pushed computing toward heterogeneous Systems-on-Chip (SoCs) that pair general-purpose cores with domain-specific accelerators. But integrating multiple accelerators does not eliminate complexity -- it relocates it: orchestrating these devices and moving data between them becomes the hard problem, and today this orchestration is entirely CPU-mediated, with the host sitting in the middle of every transfer and every device launch. The reason is largely structural: software stacks are typically built around a single accelerator working on its own, discarding anything that does not originate from that device rather than coordinating with it -- so each component remains an offloading device invoked on request, rather than becoming a main actor capable of driving work by itself. As accelerator integration approaches saturation, removing the CPU from this loop -- letting devices communicate peer-to-peer -- becomes the main remaining lever for efficiency. An illustrative example is the AMD Strix Point (Ryzen AI 300) SoC, whose integrated GPU and NPU share the same physical memory yet are, by default, driven through separate stacks that move data via host-mediated copies and launch NPU work only under CPU orchestration -- overheads the shared memory should make avoidable, but which the undocumented XRT/XDNA stack prevents from being exploited directly. We address this along two orthogonal axes. First, for data movement, we implement and evaluate copy-less approaches against the default copy-based baseline. Second, for kernel launching, we reverse-engineer the XRT/XDNA stack and NPU firmware to expose its command-submission protocol, then use this knowledge, together with a custom kernel module, to dispatch NPU kernels directly from the GPU -- achieving peer-to-peer communication between the two devices through the open-source dma-buf Linux API and bypassing the host entirely. Copy-less data movement consistently beats the baseline, scaling up to a ~2x speedup; combined with our direct NPU dispatch, it reaches up to ~11x on compute-intensive workloads. These results show that meaningful efficiency gains remain accessible in an already-fixed heterogeneous SoC by rethinking communication rather than computation.
Abstract in lingua italiana} La fine del Dennard scaling ha spinto la computazione verso i System-on-Chip (SoC) eterogenei, che accoppiano core general-purpose con domain-specific accelerators. Ma integrare più acceleratori non elimina la complessità: la rilocalizza. Orchestrare questi dispositivi e spostare i dati tra di essi diventa il problema difficile, e oggi questa orchestrazione è interamente mediata dalla CPU, con l'host che siede in mezzo a ogni trasferimento e a ogni lancio di dispositivo. La ragione è in gran parte strutturale: gli stack software sono tipicamente costruiti attorno a un singolo acceleratore che lavora per conto proprio, scartando tutto ciò che non ha origine da quel dispositivo anziché coordinarsi con esso — così ogni componente resta un dispositivo di offloading invocato su richiesta, invece di diventare un attore principale capace di guidare il lavoro autonomamente. Man mano che l'integrazione degli acceleratori si avvicina alla saturazione, rimuovere la CPU da questo loop — lasciando che i dispositivi comunichino peer-to-peer — diventa la principale leva rimasta per l'efficienza. Un esempio illustrativo è il SoC AMD Strix Point (Ryzen AI 300), la cui GPU e NPU integrate condividono la stessa memoria fisica ma sono, di default, pilotate attraverso stack separati che spostano i dati tramite copie mediate dall'host e lanciano il lavoro sulla NPU solo sotto orchestrazione della CPU — overhead che la memoria condivisa dovrebbe rendere evitabili, ma che lo stack XRT/XDNA, non documentato, impedisce di sfruttare direttamente. Affrontiamo il problema lungo due assi ortogonali. Primo, per il movimento dei dati, implementiamo e valutiamo approcci copy-less rispetto al baseline basato su copie di default. Secondo, per il lancio dei kernel, facciamo reverse-engineering dello stack XRT/XDNA e del firmware della NPU per esporne il protocollo di command-submission, e poi usiamo questa conoscenza, insieme a un modulo kernel custom, per fare dispatch dei kernel della NPU direttamente dalla GPU — realizzando una comunicazione peer-to-peer tra i due dispositivi attraverso l'API Linux open-source dma-buf e bypassando completamente l'host. Il movimento dei dati copy-less batte costantemente il baseline, scalando fino a uno speedup di circa 2x; combinato con il nostro dispatch diretto della NPU, raggiunge fino a circa 11x su workload compute-intensive. Questi risultati mostrano che guadagni di efficienza significativi restano accessibili in un SoC eterogeneo già fissato ripensando la comunicazione anziché la computazione.
Towards full peer-to-peer GPU-NPU communication on edge AI SOCs
LAURENZI, MARCO
2025/2026
Abstract
The end of Dennard scaling has pushed computing toward heterogeneous Systems-on-Chip (SoCs) that pair general-purpose cores with domain-specific accelerators. But integrating multiple accelerators does not eliminate complexity -- it relocates it: orchestrating these devices and moving data between them becomes the hard problem, and today this orchestration is entirely CPU-mediated, with the host sitting in the middle of every transfer and every device launch. The reason is largely structural: software stacks are typically built around a single accelerator working on its own, discarding anything that does not originate from that device rather than coordinating with it -- so each component remains an offloading device invoked on request, rather than becoming a main actor capable of driving work by itself. As accelerator integration approaches saturation, removing the CPU from this loop -- letting devices communicate peer-to-peer -- becomes the main remaining lever for efficiency. An illustrative example is the AMD Strix Point (Ryzen AI 300) SoC, whose integrated GPU and NPU share the same physical memory yet are, by default, driven through separate stacks that move data via host-mediated copies and launch NPU work only under CPU orchestration -- overheads the shared memory should make avoidable, but which the undocumented XRT/XDNA stack prevents from being exploited directly. We address this along two orthogonal axes. First, for data movement, we implement and evaluate copy-less approaches against the default copy-based baseline. Second, for kernel launching, we reverse-engineer the XRT/XDNA stack and NPU firmware to expose its command-submission protocol, then use this knowledge, together with a custom kernel module, to dispatch NPU kernels directly from the GPU -- achieving peer-to-peer communication between the two devices through the open-source dma-buf Linux API and bypassing the host entirely. Copy-less data movement consistently beats the baseline, scaling up to a ~2x speedup; combined with our direct NPU dispatch, it reaches up to ~11x on compute-intensive workloads. These results show that meaningful efficiency gains remain accessible in an already-fixed heterogeneous SoC by rethinking communication rather than computation.| File | Dimensione | Formato | |
|---|---|---|---|
|
marco_polimi_executive_summary.pdf
accessibile in internet per tutti
Dimensione
459.79 kB
Formato
Adobe PDF
|
459.79 kB | Adobe PDF | Visualizza/Apri |
|
marco_polimi_master_thesis.pdf
accessibile in internet per tutti a partire dal 02/07/2029
Dimensione
1.75 MB
Formato
Adobe PDF
|
1.75 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261475