The choice of locations for urban retail activities often relies on two distinct approaches: expert-driven multi-criteria methods on the one hand, and machine-learning models trained from the observed locations of existing commercial activities on the other. These two approaches are often applied separately and are rarely compared formally with respect to both their outputs and the criteria that determine them. This thesis addresses this issue with reference to the café sector in the Municipality of Milan. The work applies the Analytic Hierarchy Process (AHP) and a Random Forest (RF) classifier as two independent models for assessing commercial suitability over a spatial grid of the city. The original element of the thesis is the use of SHapley Additive exPlanations (SHAP) to compare the variables considered important by the Random Forest classifier with the weights assigned through AHP. In this way, the thesis evaluates whether the data-driven model and the expert-driven model agree not only on the areas considered most suitable, but also on the factors determining such suitability. Both models use open spatial data from OpenStreetMap (OSM), Urban Atlas 2021, ISTAT 2021 census tracts, and night-light imagery acquired by the Visible Infrared Imaging Radiometer Suite (VIIRS). The data were processed through a reproducible pipeline, from data collection and cleaning to the construction of the spatial indicators used in the models and in the WebGIS. The comparison between the maps produced by the two methods shows a high Spearman correlation, ρ = 0.7937: many cells positively evaluated by the AHP model are also positively evaluated by the Random Forest model. However, the permutation test shows that this similarity is not higher than what can be expected when considering the spatial structure of the data. In other words, the two maps may appear similar because they both reflect the spatial distribution of activities and urban functions in Milan, not necessarily because the two models agree on the suitability criteria. This interpretation is confirmed by the comparison between criteria: the correlation between the AHP weights and the variable importance estimated through SHAP is negative and not significant, with ρ = −0.4073 and an empirical two-tailed p-value of 0.2412. The AHP model assigns the highest weight to public transport accessibility, whereas the Random Forest model interpreted through SHAP gives greater importance to local café density and Point of Interest (POI) diversity. Additional checks show that the difference between the two models remains stable when changing the threshold used to define suitable areas and the weight assigned to metro accessibility. The results are finally visualized through a WebGIS, which allows users to inspect, for each cell, both the suitability score and the variables that most influenced the Random Forest prediction.

La scelta della localizzazione di attività commerciali al dettaglio in ambito urbano si basa spesso su due approcci distinti: da un lato, metodi multicriterio fondati sul giudizio esperto; dall’altro, modelli di machine learning addestrati a partire dalla localizzazione osservata delle attività commerciali esistenti. Questi due approcci sono spesso applicati separatamente e raramente vengono confrontati in modo formale rispetto ai risultati prodotti e ai criteri che li determinano. Questa tesi affronta tale problema con riferimento al settore dei café nel Comune di Milano. Il lavoro applica l’Analytic Hierarchy Process (AHP), o processo di analisi gerarchica, e un classificatore Random Forest (RF) come due modelli indipendenti di valutazione dell’idoneità commerciale su una griglia spaziale della città. L’elemento originale della tesi è l’uso di SHapley Additive exPlanations (SHAP) per confrontare le variabili considerate importanti dal classificatore Random Forest con i pesi assegnati tramite AHP. In questo modo, la tesi valuta se il modello basato sui dati e quello basato sul giudizio esperto concordano non solo sulle aree considerate più idonee, ma anche sui fattori che determinano tale idoneità. Entrambi i modelli utilizzano dati spaziali aperti provenienti da OpenStreetMap (OSM), Urban Atlas 2021, sezioni di censimento ISTAT 2021 e immagini di luminosità notturna acquisite dal sensore Visible Infrared Imaging Radiometer Suite (VIIRS). I dati sono stati elaborati attraverso una pipeline riproducibile, dalla raccolta e pulizia delle fonti alla costruzione degli indicatori spaziali utilizzati nei modelli e nel WebGIS. Il confronto tra le mappe prodotte dai due metodi mostra una correlazione di Spearman elevata, pari a ρ = 0.7937: molte celle valutate positivamente dal modello AHP risultano positive anche per il modello Random Forest. Tuttavia, il test di permutazione mostra che questa somiglianza non è superiore a quella che ci si può attendere considerando la struttura spaziale dei dati. In altri termini, le due mappe possono apparire simili perché riflettono entrambe la distribuzione spaziale delle attività e delle funzioni urbane a Milano, non necessariamente perché i due modelli concordino sui criteri di idoneità. Questa interpretazione è confermata dal confronto tra i criteri: la correlazione tra i pesi definiti con AHP e l’importanza delle variabili stimata tramite SHAP è negativa e non significativa, con ρ = −0.4073 e p-value empirico bilaterale pari a 0.2412. Il modello AHP assegna il peso maggiore all’accessibilità al trasporto pubblico, mentre il modello Random Forest interpretato tramite SHAP attribuisce maggiore importanza alla densità locale di café e alla diversità dei Point of Interest (POI). Alcune verifiche aggiuntive mostrano che la differenza tra i due modelli rimane anche modificando la soglia usata per definire le aree idonee e il peso attribuito alla vicinanza alla metropolitana. I risultati sono infine visualizzati in un WebGIS, che consente di consultare per ogni cella sia il punteggio di idoneità sia le variabili che hanno maggiormente influenzato la previsione del modello Random Forest.

Explainable spatial AI for urban retail site selection: an AHP-SHAP audit with WebGIS delivery in Milan

Avila Santos, Miguel Angel;SARWARY, HAFIZULLAH
2025/2026

Abstract

The choice of locations for urban retail activities often relies on two distinct approaches: expert-driven multi-criteria methods on the one hand, and machine-learning models trained from the observed locations of existing commercial activities on the other. These two approaches are often applied separately and are rarely compared formally with respect to both their outputs and the criteria that determine them. This thesis addresses this issue with reference to the café sector in the Municipality of Milan. The work applies the Analytic Hierarchy Process (AHP) and a Random Forest (RF) classifier as two independent models for assessing commercial suitability over a spatial grid of the city. The original element of the thesis is the use of SHapley Additive exPlanations (SHAP) to compare the variables considered important by the Random Forest classifier with the weights assigned through AHP. In this way, the thesis evaluates whether the data-driven model and the expert-driven model agree not only on the areas considered most suitable, but also on the factors determining such suitability. Both models use open spatial data from OpenStreetMap (OSM), Urban Atlas 2021, ISTAT 2021 census tracts, and night-light imagery acquired by the Visible Infrared Imaging Radiometer Suite (VIIRS). The data were processed through a reproducible pipeline, from data collection and cleaning to the construction of the spatial indicators used in the models and in the WebGIS. The comparison between the maps produced by the two methods shows a high Spearman correlation, ρ = 0.7937: many cells positively evaluated by the AHP model are also positively evaluated by the Random Forest model. However, the permutation test shows that this similarity is not higher than what can be expected when considering the spatial structure of the data. In other words, the two maps may appear similar because they both reflect the spatial distribution of activities and urban functions in Milan, not necessarily because the two models agree on the suitability criteria. This interpretation is confirmed by the comparison between criteria: the correlation between the AHP weights and the variable importance estimated through SHAP is negative and not significant, with ρ = −0.4073 and an empirical two-tailed p-value of 0.2412. The AHP model assigns the highest weight to public transport accessibility, whereas the Random Forest model interpreted through SHAP gives greater importance to local café density and Point of Interest (POI) diversity. Additional checks show that the difference between the two models remains stable when changing the threshold used to define suitable areas and the weight assigned to metro accessibility. The results are finally visualized through a WebGIS, which allows users to inspect, for each cell, both the suitability score and the variables that most influenced the Random Forest prediction.
ING - Scuola di Ingegneria Industriale e dell'Informazione
22-lug-2026
2025/2026
La scelta della localizzazione di attività commerciali al dettaglio in ambito urbano si basa spesso su due approcci distinti: da un lato, metodi multicriterio fondati sul giudizio esperto; dall’altro, modelli di machine learning addestrati a partire dalla localizzazione osservata delle attività commerciali esistenti. Questi due approcci sono spesso applicati separatamente e raramente vengono confrontati in modo formale rispetto ai risultati prodotti e ai criteri che li determinano. Questa tesi affronta tale problema con riferimento al settore dei café nel Comune di Milano. Il lavoro applica l’Analytic Hierarchy Process (AHP), o processo di analisi gerarchica, e un classificatore Random Forest (RF) come due modelli indipendenti di valutazione dell’idoneità commerciale su una griglia spaziale della città. L’elemento originale della tesi è l’uso di SHapley Additive exPlanations (SHAP) per confrontare le variabili considerate importanti dal classificatore Random Forest con i pesi assegnati tramite AHP. In questo modo, la tesi valuta se il modello basato sui dati e quello basato sul giudizio esperto concordano non solo sulle aree considerate più idonee, ma anche sui fattori che determinano tale idoneità. Entrambi i modelli utilizzano dati spaziali aperti provenienti da OpenStreetMap (OSM), Urban Atlas 2021, sezioni di censimento ISTAT 2021 e immagini di luminosità notturna acquisite dal sensore Visible Infrared Imaging Radiometer Suite (VIIRS). I dati sono stati elaborati attraverso una pipeline riproducibile, dalla raccolta e pulizia delle fonti alla costruzione degli indicatori spaziali utilizzati nei modelli e nel WebGIS. Il confronto tra le mappe prodotte dai due metodi mostra una correlazione di Spearman elevata, pari a ρ = 0.7937: molte celle valutate positivamente dal modello AHP risultano positive anche per il modello Random Forest. Tuttavia, il test di permutazione mostra che questa somiglianza non è superiore a quella che ci si può attendere considerando la struttura spaziale dei dati. In altri termini, le due mappe possono apparire simili perché riflettono entrambe la distribuzione spaziale delle attività e delle funzioni urbane a Milano, non necessariamente perché i due modelli concordino sui criteri di idoneità. Questa interpretazione è confermata dal confronto tra i criteri: la correlazione tra i pesi definiti con AHP e l’importanza delle variabili stimata tramite SHAP è negativa e non significativa, con ρ = −0.4073 e p-value empirico bilaterale pari a 0.2412. Il modello AHP assegna il peso maggiore all’accessibilità al trasporto pubblico, mentre il modello Random Forest interpretato tramite SHAP attribuisce maggiore importanza alla densità locale di café e alla diversità dei Point of Interest (POI). Alcune verifiche aggiuntive mostrano che la differenza tra i due modelli rimane anche modificando la soglia usata per definire le aree idonee e il peso attribuito alla vicinanza alla metropolitana. I risultati sono infine visualizzati in un WebGIS, che consente di consultare per ogni cella sia il punteggio di idoneità sia le variabili che hanno maggiormente influenzato la previsione del modello Random Forest.
File allegati
File Dimensione Formato  
2026_07_AvilaSantos_Sarwary_Executive_Summary_02.pdf

accessibile in internet per tutti

Dimensione 596.57 kB
Formato Adobe PDF
596.57 kB Adobe PDF Visualizza/Apri
2026_07_AvilaSantos_Sarwary_Thesis_01.pdf

accessibile in internet per tutti

Dimensione 9.83 MB
Formato Adobe PDF
9.83 MB Adobe PDF Visualizza/Apri

I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10589/258137