The design of neural architectures is commonly treated as a process separate from parameter optimization, typically relying on human expertise or automated search. Still, manual design requires domain knowledge that does not transfer across tasks, while automated search methods often remain computationally prohibitive. To overcome these limitations, we introduce the Self-Partitioning Adaptive Widening Network (SPAWN), a novel framework that jointly learns both its architecture and its parameters by dynamically increasing its representational capacity during training. SPAWN starts from a single shallow predictor and incrementally builds a soft-gated mixture of experts by partitioning the input space where additional complexity is needed. In contrast to prior approaches that expand networks by adding neurons or layers, SPAWN increases capacity by recursively subdividing the input space, yielding a differentiable tree of affine experts trained end-to-end. This expansion process is governed by data-driven criteria that determine when to grow, where to split, and how to preserve the model's outputs, enabling seamless optimization without retraining from scratch. Empirical evaluations on standard regression and classification benchmarks demonstrate that SPAWN automatically discovers compact architectures that match or surpass the predictive performance of substantially larger, fixed-capacity models, while using markedly fewer parameters.
La progettazione delle architetture neurali è comunemente trattata come un processo separato dall'ottimizzazione dei parametri, affidandosi tipicamente all'esperienza umana o a metodi di ricerca automatica. La progettazione manuale, tuttavia, richiede competenze fortemente legate al dominio applicativo e difficilmente trasferibili, mentre i metodi di ricerca automatica comportano spesso costi computazionali proibitivi. Per superare queste limitazioni, introduciamo la Self-Partitioning Adaptive Widening Network (SPAWN), un framework che apprende congiuntamente architettura e parametri aumentando dinamicamente la propria capacità rappresentativa durante l'addestramento. SPAWN parte da un singolo predittore elementare e costruisce incrementalmente una miscela di esperti con gating differenziabile, partizionando lo spazio degli input là dove è necessaria maggiore complessità. A differenza degli approcci che espandono la rete aggiungendo neuroni o strati, SPAWN accresce la capacità suddividendo ricorsivamente lo spazio degli input e producendo un albero differenziabile di esperti affini addestrato end-to-end. Criteri basati sui dati governano quando crescere, dove suddividere e come preservare le uscite del modello, consentendo un'ottimizzazione continua senza necessità di ri-addestramento. Le valutazioni empiriche su benchmark standard di regressione e classificazione dimostrano che SPAWN individua automaticamente architetture compatte in grado di eguagliare o superare le prestazioni predittive di modelli a capacità fissa significativamente più grandi, impiegando un numero notevolmente inferiore di parametri.
Adaptive self-growing neural networks through differentiable input-space partitioning
Lolli, Federico
2025/2026
Abstract
The design of neural architectures is commonly treated as a process separate from parameter optimization, typically relying on human expertise or automated search. Still, manual design requires domain knowledge that does not transfer across tasks, while automated search methods often remain computationally prohibitive. To overcome these limitations, we introduce the Self-Partitioning Adaptive Widening Network (SPAWN), a novel framework that jointly learns both its architecture and its parameters by dynamically increasing its representational capacity during training. SPAWN starts from a single shallow predictor and incrementally builds a soft-gated mixture of experts by partitioning the input space where additional complexity is needed. In contrast to prior approaches that expand networks by adding neurons or layers, SPAWN increases capacity by recursively subdividing the input space, yielding a differentiable tree of affine experts trained end-to-end. This expansion process is governed by data-driven criteria that determine when to grow, where to split, and how to preserve the model's outputs, enabling seamless optimization without retraining from scratch. Empirical evaluations on standard regression and classification benchmarks demonstrate that SPAWN automatically discovers compact architectures that match or surpass the predictive performance of substantially larger, fixed-capacity models, while using markedly fewer parameters.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Lolli_Tesi.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
11.87 MB
Formato
Adobe PDF
|
11.87 MB | Adobe PDF | Visualizza/Apri |
|
2026_07_Lolli_Executive Summary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
480.6 kB
Formato
Adobe PDF
|
480.6 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260874