Logarithms and exponentials are non-algebraic transcendental functions that cannot be expressed as solutions of polynomial equations. Their efficient computation is essential across applications ranging from machine learning to image analysis, signal processing, and scientific computing. Despite their importance, such floating-point primitives are often absent from the vector datapath of vector CPUs and accelerators, requiring them to fall back to scalar pipelines. Several proprietary and open-source libraries address this issue by implementing vectorized approximations of such operators that, leveraging the available vector- ized primitives, avoid the serialization of otherwise parallel workloads. While effective, existing solutions are largely tied to Central Processing Unit (CPU) vector instruction sets, making them unsuitable for Domain-Specific Accelerators (DSAs) with distinct vector datapaths, such as AMD AI Engine (AIE) and AIE-ML, where support remains limited. Moreover, CPU-based libraries primarily target float32 precision, while many accelerators natively support 16-bit formats like bfloat16. Motivated by this gap, we present DAIE, a vectorized math library for AIEs. DAIE provides fully vectorized implementations of transcendental functions in float32 and bfloat16 precision. DAIE’s kernels rely only on vectorized floating- point addition and multiplication, basic bitwise operations such as AND and OR, and vector selection. We evaluate DAIE across four architectural targets: AIE, where it provides the only available fully vectorized implementation; AIE-ML, where it outperforms existing table-based solutions; and CPUs with vector extensions, including x86 Intel Advanced Vector Extensions 512 (AVX-512) and Arm NEON. Our evaluation includes three studies. First, we measure the accuracy of reciprocal, base-2 logarithm, and exponential through a full sweep of floating-point inputs against a ground-truth implementation. Across all implementations, DAIE achieves an error of at most 4 Units in the Last Place (ULP). Second, we evaluate performance using two common transcendental kernels: SoftMax and entropy. DAIE outperforms all existing solutions on AIE-ML and achieves competitive results on CPUs, obtaining comparable accuracy and throughput to SLEEF and Google Highway. Finally, we integrate our entropy kernel into a state-of-the-art FPGA image mutual-information accelerator, showing that accelerating a single transcendental primitive can improve a larger heterogeneous system. The resulting design achieves 49.5% higher operating frequency, 1.15× speedup, and 1.91× energy reduction over Field Programmable Gate Array (FPGA)-only implementations, highlighting the system-level impact of efficient vectorized math support on AMD AI Engines.
I logaritmi e gli esponenziali sono funzioni trascendenti non algebriche che non possono essere espresse come soluzioni di equazioni polinomiali. Il loro calcolo efficiente è essenziale in applicazioni che spaziano dal machine learning all’analisi di immagini, all’elaborazione dei segnali e al calcolo scientifico. Nonostante la loro importanza, tali primitive floating-point sono spesso assenti dal datapath vettoriale delle CPU e degli acceleratori, richiedendo quindi il ripiego su pipeline scalari. Diverse librerie proprietarie e open-source affrontano questo problema implementando approssimazioni vettorializzate di tali operatori che, sfruttando le primitive vettorializzate disponibili, evitano la serializzazione di carichi di lavoro altrimenti paralleli. Sebbene efficaci, le soluzioni esistenti sono in larga parte legate ai set di istruzioni vettoriali delle CPU, il che le rende inadatte ad acceleratori domain-specific (DSA) con datapath vettoriali distinti, come AIE e AIE-ML, dove il supporto rimane limitato. Inoltre, le librerie basate su CPU mirano principalmente alla precisione float32, mentre molti acceleratori supportano nativamente formati a 16 bit come bfloat16. Motivati da questa lacuna, presentiamo DAIE, una libreria matematica vettorializzata per AIEs. DAIE fornisce implementazioni completamente vettorializzate di funzioni trascendenti in precisione float32 e bfloat16. I kernel di DAIE si basano esclusivamente su addizione e moltiplicazione floating-point vettorializzate, operazioni bitwise di base come AND e OR, e selezione vettoriale. Valutiamo DAIE su quattro target architetturali: AIE, dove fornisce l’unica implementazione completamente vettorializzata disponibile; AIE-ML, dove supera le soluzioni esistenti basate su tabelle; e CPU con estensioni vettoriali, incluse x86 AVX-512 e Arm NEON. La nostra valutazione include tre studi. In primo luogo, misuriamo l’accuratezza del reciproco, del logaritmo in base 2 e dell’esponenziale mediante una scansione completa degli input floating-point rispetto a un’implementazione di riferimento. Su tutte le implementazioni, DAIE ottiene un errore al massimo di 4 ULP. In secondo luogo, valutiamo le prestazioni usando due kernel trascendenti comuni: softmax ed entropia. DAIE supera tutte le soluzioni esistenti su AIE-ML e ottiene risultati competitivi su CPU, raggiungendo accuratezza e throughput comparabili a SLEEF e Google Highway. Infine, integriamo il nostro kernel di entropia in un acceleratore FPGA allo stato dell’arte per la mutual information tra immagini, mostrando che l’accelerazione di una singola primitiva trascendente può migliorare un sistema eterogeneo più ampio. Il progetto risultante raggiunge una frequenza operativa superiore del 49.5%, uno speedup di 1.15× e un miglioramento dell’efficienza energetica di 1.91× rispetto a implementazioni solo FPGA, evidenziando l’impatto a livello di sistema del supporto matematico vettorializzato efficiente sugli AMD AI Engines.
DAIE: a vectorized math library for AMD AI engines
Brunetta, Giacomo
2025/2026
Abstract
Logarithms and exponentials are non-algebraic transcendental functions that cannot be expressed as solutions of polynomial equations. Their efficient computation is essential across applications ranging from machine learning to image analysis, signal processing, and scientific computing. Despite their importance, such floating-point primitives are often absent from the vector datapath of vector CPUs and accelerators, requiring them to fall back to scalar pipelines. Several proprietary and open-source libraries address this issue by implementing vectorized approximations of such operators that, leveraging the available vector- ized primitives, avoid the serialization of otherwise parallel workloads. While effective, existing solutions are largely tied to Central Processing Unit (CPU) vector instruction sets, making them unsuitable for Domain-Specific Accelerators (DSAs) with distinct vector datapaths, such as AMD AI Engine (AIE) and AIE-ML, where support remains limited. Moreover, CPU-based libraries primarily target float32 precision, while many accelerators natively support 16-bit formats like bfloat16. Motivated by this gap, we present DAIE, a vectorized math library for AIEs. DAIE provides fully vectorized implementations of transcendental functions in float32 and bfloat16 precision. DAIE’s kernels rely only on vectorized floating- point addition and multiplication, basic bitwise operations such as AND and OR, and vector selection. We evaluate DAIE across four architectural targets: AIE, where it provides the only available fully vectorized implementation; AIE-ML, where it outperforms existing table-based solutions; and CPUs with vector extensions, including x86 Intel Advanced Vector Extensions 512 (AVX-512) and Arm NEON. Our evaluation includes three studies. First, we measure the accuracy of reciprocal, base-2 logarithm, and exponential through a full sweep of floating-point inputs against a ground-truth implementation. Across all implementations, DAIE achieves an error of at most 4 Units in the Last Place (ULP). Second, we evaluate performance using two common transcendental kernels: SoftMax and entropy. DAIE outperforms all existing solutions on AIE-ML and achieves competitive results on CPUs, obtaining comparable accuracy and throughput to SLEEF and Google Highway. Finally, we integrate our entropy kernel into a state-of-the-art FPGA image mutual-information accelerator, showing that accelerating a single transcendental primitive can improve a larger heterogeneous system. The resulting design achieves 49.5% higher operating frequency, 1.15× speedup, and 1.91× energy reduction over Field Programmable Gate Array (FPGA)-only implementations, highlighting the system-level impact of efficient vectorized math support on AMD AI Engines.| File | Dimensione | Formato | |
|---|---|---|---|
|
ExecutiveSummary_GiacomoBrunetta (1).pdf
accessibile in internet per tutti a partire dal 02/07/2029
Descrizione: Executive Summary
Dimensione
874.49 kB
Formato
Adobe PDF
|
874.49 kB | Adobe PDF | Visualizza/Apri |
|
Thesis_GiacomoBrunetta (2).pdf
accessibile in internet per tutti a partire dal 02/07/2029
Descrizione: Tesi
Dimensione
1.6 MB
Formato
Adobe PDF
|
1.6 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261521