This thesis develops and empirically validates a scalable Bayesian nonparametric framework for clustering areal data based on pairwise distances between kernel density estimates. Motivated by the computational intractability of state-of-the-art spatial Bayesian models — which require hours of computation even for moderate geographic scales — we propose a strategy that transforms the inferential problem from tens of thousands of individual observations to a small number of area-specific density representations and their pairwise distances. Building upon the distance-based cohesion-repulsion likelihood of \textcite{Natarajan02042024}, we combine a Normalized Generalized Gamma (NGG) process induced prior for fully data-driven inference on the number of clusters with extensions that incorporate spatial adjacency structure and auxiliary demographic covariates. An efficient MCMC sampling scheme synthesizes the Sequential Allocated Merge-Split sampler, a distance-based adaptation of Locality-Sensitive Sampling for observation selection, a Smart-Dumb-Dumb-Smart move-type selection strategy, periodic Gibbs scans, and shuffle moves. The framework is validated on simulated data and applied to two real datasets: California census income data aggregated over 93 Public Use Microdata Areas, where inference completes in under two seconds compared to ten hours for existing approaches in the field of boundary detection, and Italian municipalities income data comprising 7,903 areal units. Results demonstrate competitive clustering accuracy, interpretable socioeconomic groupings, and substantial computational gains across all model configurations. The entire framework is implemented as a high-performance, modular C++ library with R bindings, publicly available at \href{https://github.com/Filippo-Galli/BNPClust}{https://github.com/Filippo-Galli/BNPClust}.
Questa tesi sviluppa e valida empiricamente un framework bayesiano non parametrico scalabile per il clustering di dati areali, basato sulle distanze a coppie tra stime di densità kernel. Motivati dall’intrattabilità computazionale dei modelli bayesiani spaziali allo stato dell’arte — che richiedono ore di calcolo anche per scale geografiche moderate — proponiamo una strategia che trasforma il problema inferenziale: da decine di migliaia di singole osservazioni individuali a un numero ridotto di rappre- sentazioni di densità specifiche per area e alle loro distanze a coppie. Basandoci sulla verosimiglianza coesivo-repulsiva basata sulle distanze di Natarajan et al. (2024), combiniamo una distribuzione a pri- ori indotta da un processo Normalized Generalized Gamma (NGG) — per un’inferenza sul numero di cluster completamente guidata dai dati (data-driven) — con estensioni che incorporano la struttura di adiacenza spaziale e covariate demografiche ausiliarie. Un efficiente schema di campionamento MCMC sintetizza il campionatore Sequential Allocated Merge-Split, un adattamento basato sulle distanze del Locality-Sensitive Sampling per la selezione delle osservazioni, una strategia Smart-Dumb-Dumb-Smart per la selezione del tipo di mossa, scansioni di Gibbs periodiche e mosse di shuffle. Il framework è val- idato su dati simulati e applicato a due dataset reali: i dati sul reddito del censimento della California aggregati in 93 Public Use Microdata Areas (PUMA), dove l’inferenza viene completata in meno di due secondi rispetto alle dieci ore richieste dagli approcci esistenti nel campo della boundary detection, e i dati sul reddito dei comuni italiani, che comprendono 7.903 unità areali. I risultati dimostrano un’accuratezza di clustering competitiva, raggruppamenti socioeconomici interpretabili e sostanziali vantaggi computazionali in tutte le configurazioni del modello. L’intero framework è implementato come una libreria C++ modulare e ad alte prestazioni con binding in R, disponibile pubblicamente all’indirizzo https://github.com/Filippo-Galli/BNPClust.
Bayesian distance-based clustering of areal data
GALLI, FILIPPO
2024/2025
Abstract
This thesis develops and empirically validates a scalable Bayesian nonparametric framework for clustering areal data based on pairwise distances between kernel density estimates. Motivated by the computational intractability of state-of-the-art spatial Bayesian models — which require hours of computation even for moderate geographic scales — we propose a strategy that transforms the inferential problem from tens of thousands of individual observations to a small number of area-specific density representations and their pairwise distances. Building upon the distance-based cohesion-repulsion likelihood of \textcite{Natarajan02042024}, we combine a Normalized Generalized Gamma (NGG) process induced prior for fully data-driven inference on the number of clusters with extensions that incorporate spatial adjacency structure and auxiliary demographic covariates. An efficient MCMC sampling scheme synthesizes the Sequential Allocated Merge-Split sampler, a distance-based adaptation of Locality-Sensitive Sampling for observation selection, a Smart-Dumb-Dumb-Smart move-type selection strategy, periodic Gibbs scans, and shuffle moves. The framework is validated on simulated data and applied to two real datasets: California census income data aggregated over 93 Public Use Microdata Areas, where inference completes in under two seconds compared to ten hours for existing approaches in the field of boundary detection, and Italian municipalities income data comprising 7,903 areal units. Results demonstrate competitive clustering accuracy, interpretable socioeconomic groupings, and substantial computational gains across all model configurations. The entire framework is implemented as a high-performance, modular C++ library with R bindings, publicly available at \href{https://github.com/Filippo-Galli/BNPClust}{https://github.com/Filippo-Galli/BNPClust}.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_02_Galli_thesis.pdf
accessibile in internet per tutti
Descrizione: Tesi
Dimensione
14.39 MB
Formato
Adobe PDF
|
14.39 MB | Adobe PDF | Visualizza/Apri |
|
2026_02_Galli_executiveSummary.pdf
accessibile in internet per tutti
Descrizione: Executive Summary
Dimensione
460.67 kB
Formato
Adobe PDF
|
460.67 kB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/253003