The rise of high-velocity data streams in domains such as network monitoring, fraud detection, and the Internet of Things has exposed the fundamental limitations of centralized streaming machine learning frameworks. Existing Streaming Machine Learning tools, most notably MOA and River, are constrained by single-threaded execution or JVM-based overhead, making them inadequate for the computational demands of large ensemble methods at scale. This thesis presents a distributed streaming ensemble learning system built on Renoir, a high-performance dataflow runtime written in Rust. We design and implement three state-of-the-art ensemble algorithms — Adaptive Random Forest (ARF), Streaming Random Patches (SRP), and Aggregated Mondrian Forest (AMF) — within a unified, extensible architecture that decouples the streaming pipeline from the underlying learning algorithms. The system exploits the embarrassingly parallel nature of ensemble methods by distributing individual base learners across worker nodes, using key-based stream routing to coordinate training and prediction aggregation without a centralized bottleneck. Experimental evaluation across thirteen real-world and synthetic datasets demonstrates that the proposed implementation preserves the predictive accuracy of the original algorithms while delivering decisive throughput advantages at every scale. On a single core, Renoir already outperforms MOA by up to 35.73× and Onelearn by up to 18.24×, driven by Rust's zero-cost abstractions and cache-efficient memory layout. Scaling to 32 cores further widens this gap: our distributed ensembles sustain throughputs that remain an order of magnitude above the centralized competition, confirming that Renoir's dataflow model translates additional hardware directly into processing speed. This behavior is validated across varying hyperparameter configurations, drift conditions, and a real networked deployment, addressing the full range of practical streaming scenarios.
L'ascesa dei flussi di dati ad alta velocità nei domini quali il monitoraggio di rete, il rilevamento delle frodi e l'Internet of Things ha messo in luce i limiti fondamentali dei framework centralizzati di streaming machine learning. Gli strumenti esistenti di Streaming Machine Learning, in particolare MOA e River, sono limitati dall'esecuzione a thread singolo o dall'overhead basato su JVM, risultando inadeguati per le richieste computazionali dei metodi ensemble su larga scala. Questa tesi presenta un sistema distribuito di streaming ensemble learning basato su Renoir, un runtime di dataflow ad alte prestazioni scritto in Rust. Progettiamo e implementiamo tre algoritmi di ensemble allo stato dell'arte — Adaptive Random Forest (ARF), Streaming Random Patches (SRP), and Aggregated Mondrian Forest (AMF) — all'interno di un'architettura unificata ed estensibile che disaccoppia la pipeline di streaming dagli algoritmi di apprendimento sottostanti. Il sistema sfrutta la natura intrinsecamente parallela dei metodi ensemble distribuendo i singoli base learner tra i nodi worker, utilizzando un routing del flusso basato su chiavi per coordinare l'addestramento e l'aggregazione delle predizioni senza un collo di bottiglia centralizzato. La valutazione sperimentale su tredici dataset reali e sintetici dimostra che la soluzione proposta preserva l'accuratezza predittiva degli algoritmi originali, offrendo al contempo decisivi vantaggi in termini di throughput a ogni scala. Su un singolo core, Renoir supera già MOA fino a 35,73× e Onelearn fino a 18,24×, grazie alle astrazioni a costo zero di Rust e a una gestione di memoria efficiente. La scalabilità a 32 core amplia ulteriormente questo divario: i nostri ensemble distribuiti sostengono throughput che rimangono di un ordine di grandezza superiori rispetto alla concorrenza centralizzata, confermando che il modello dataflow di Renoir traduce l'hardware aggiuntivo direttamente in velocità di elaborazione. Questo comportamento è validato attraverso diverse configurazioni di iperparametri, condizioni di drift e una reale implementazione in rete, coprendo l'intera gamma dei flussi di scenari pratici.
Distributed streaming ensemble methods in Renoir
SULFARO, ANTONIO;Shi, Zhi Qiang
2025/2026
Abstract
The rise of high-velocity data streams in domains such as network monitoring, fraud detection, and the Internet of Things has exposed the fundamental limitations of centralized streaming machine learning frameworks. Existing Streaming Machine Learning tools, most notably MOA and River, are constrained by single-threaded execution or JVM-based overhead, making them inadequate for the computational demands of large ensemble methods at scale. This thesis presents a distributed streaming ensemble learning system built on Renoir, a high-performance dataflow runtime written in Rust. We design and implement three state-of-the-art ensemble algorithms — Adaptive Random Forest (ARF), Streaming Random Patches (SRP), and Aggregated Mondrian Forest (AMF) — within a unified, extensible architecture that decouples the streaming pipeline from the underlying learning algorithms. The system exploits the embarrassingly parallel nature of ensemble methods by distributing individual base learners across worker nodes, using key-based stream routing to coordinate training and prediction aggregation without a centralized bottleneck. Experimental evaluation across thirteen real-world and synthetic datasets demonstrates that the proposed implementation preserves the predictive accuracy of the original algorithms while delivering decisive throughput advantages at every scale. On a single core, Renoir already outperforms MOA by up to 35.73× and Onelearn by up to 18.24×, driven by Rust's zero-cost abstractions and cache-efficient memory layout. Scaling to 32 cores further widens this gap: our distributed ensembles sustain throughputs that remain an order of magnitude above the centralized competition, confirming that Renoir's dataflow model translates additional hardware directly into processing speed. This behavior is validated across varying hyperparameter configurations, drift conditions, and a real networked deployment, addressing the full range of practical streaming scenarios.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Sulfaro_Shi_Executive_Summary.pdf
accessibile in internet per tutti
Descrizione: Executive summary
Dimensione
774.41 kB
Formato
Adobe PDF
|
774.41 kB | Adobe PDF | Visualizza/Apri |
|
2026_07_Sulfaro_Shi_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
3.55 MB
Formato
Adobe PDF
|
3.55 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260348