Individual tree crown detection from very high resolution (VHR) imagery is a core problem in remote sensing for urban and rural vegetation monitoring, yet published studies rarely compare representative methods under identical experimental conditions. The primary goal of this thesis is to systematically compare five representative detection paradigms. This thesis presents a controlled benchmark of five crown detection paradigms applied to a single multispectral scene (JILIN-1 KF01C, approximately 0.5 m resolution, four spectral bands including near-infrared) acquired over Reggio Calabria, Italy, in June 2024, and evaluated against 800 manually delineated reference crown polygons distributed equally between urban and rural landscape strata. The five methods span the full methodological spectrum: local maxima with watershed segmentation, object-based image analysis (OBIA) with Random Forest, a pixel-based Random Forest classifier, DeepForest (a pretrained RetinaNet detector), and U-Net (a semantic segmentation model with a ResNet-34 encoder). All methods share a common preprocessing pipeline, a fixed 512 × 512 px tile grid with spatial block partitioning, and a shared evaluation using Hungarian assignment IoU matching at thresholds of 0.25 and 0.50 on a fixed 302-crown test set. The threshold of 0.25 serves as a lenient criterion that captures coarse crown detections and tolerates annotation boundary uncertainty, while 0.50 provides a stricter standard reflecting more precise crown delineation. As a benchmark comparison, no single method achieved satisfactory detection performance, and results are reported here as reference values rather than evidence of methodological superiority. Across the five methods, DeepForest produced the highest F1-score (F1 at 0.25 = 0.313; F1 at 0.50 = 0.297), the highest recall (R at 0.25 = 0.503), and the highest matched mean IoU (mean IoU at 0.25 = 0.666). U-Net achieved the highest precision (P at 0.25 = 0.307) but the lowest recall among the supervised methods (R at 0.25 = 0.235). The three handcrafted methods (watershed, OBIA with Random Forest, and pixel-based Random Forest) clustered in a narrow F1 band (F1 at 0.25 of 0.103–0.127). The rural stratum was consistently harder across all methods: DeepForest led in both strata but with a pronounced urban advantage (F1 at 0.25 = 0.430 urban versus 0.248 rural). The main failure modes were the spectral similarity between target crowns and surrounding Mediterranean scrub and olive vegetation, and building cast shadows in the urban core. A recurring finding across supervised methods was a train-test distribution shift that required per-stratum threshold recalibration. The benchmark demonstrates that detection accuracy and boundary fidelity are partly independent properties that must be reported jointly, and that a single model and threshold do not transfer optimally across heterogeneous urban and rural Mediterranean landscapes. The reproducible end-to-end pipeline provides a reference for method selection in precision forestry and urban vegetation monitoring under VHR satellite imagery.
Il rilevamento delle chiome arboree individuali da immagini a very high resolution (VHR) rappresenta un problema centrale nel telerilevamento per il monitoraggio della vegetazione urbana e rurale; tuttavia, gli studi pubblicati raramente confrontano metodi rappresentativi in condizioni sperimentali identiche. L’obiettivo principale di questa tesi è confrontare sistematicamente cinque paradigmi rappresentativi di rilevamento. La tesi presenta un benchmark controllato di cinque paradigmi di rilevamento delle chiome applicati a una singola scena multispettrale JILIN-1 KF01C, con risoluzione di circa 0,5 m e quattro bande spettrali, incluso il vicino infrarosso, acquisita su Reggio Calabria, Italia, nel giugno 2024, e valutati rispetto a 800 poligoni di riferimento di chiome delineati manualmente e distribuiti in parti uguali tra strati paesaggistici urbani e rurali. I cinque metodi coprono l’intero spettro metodologico: massimi locali con segmentazione watershed, analisi di immagine basata su oggetti (OBIA) con Random Forest, classificatore Random Forest a livello di pixel, DeepForest (un rilevatore RetinaNet pre-addestrato) e U-Net (un modello di segmentazione semantica con encoder ResNet-34). Tutti i metodi condividono una pipeline di preprocessing comune, una griglia di tile fissa di 512 × 512 px con partizionamento spaziale a blocchi, e una valutazione condivisa basata su matching IoU mediante assegnazione ungherese alle soglie di 0,25 e 0,50 su un test set fisso di 302 chiome. La soglia di 0,25 funge da criterio permissivo, in grado di catturare rilevamenti approssimativi delle chiome e di tollerare l’incertezza nei confini delle annotazioni, mentre la soglia di 0,50 costituisce uno standard più rigoroso, associato a una delineazione più precisa delle chiome. Come confronto di benchmark, nessun metodo ha raggiunto prestazioni di rilevamento soddisfacenti, e i risultati sono riportati come valori di riferimento piuttosto che come prova di superiorità metodologica. Tra i cinque metodi, DeepForest ha ottenuto il valore più alto di F1-score (F1 a 0,25 = 0,313; F1 a 0,50 = 0,297), il recall più elevato (R a 0,25 = 0,503) e la più alta IoU media sui match (IoU media a 0,25 = 0,666). U-Net ha raggiunto la precision più alta (P a 0,25 = 0,307), ma il recall più basso tra i metodi supervisionati (R a 0,25 = 0,235). I tre metodi artigianali (watershed, OBIA con Random Forest e Random Forest a livello di pixel) si sono concentrati in una fascia ristretta di F1 a 0,25 compresa tra 0,103 e 0,127. Lo strato rurale è risultato costantemente più difficile per tutti i metodi: DeepForest ha mantenuto le migliori prestazioni in entrambi gli strati, ma con un marcato vantaggio urbano (F1 a 0,25 = 0,430 in ambito urbano contro 0,248 in ambito rurale). Le principali modalità di errore sono state la somiglianza spettrale tra le chiome target e la vegetazione circostante, composta da macchia mediterranea e uliveti, e le ombre proiettate dagli edifici nel nucleo urbano. Un risultato ricorrente nei metodi supervisionati è stato uno shift di distribuzione tra train e test, che ha richiesto una ricalibrazione delle soglie per ciascuno strato. Il benchmark dimostra che l’accuratezza del rilevamento e la fedeltà dei confini sono proprietà parzialmente indipendenti che devono essere riportate congiuntamente, e che un singolo modello e una singola soglia non si trasferiscono in modo ottimale tra paesaggi mediterranei urbani e rurali eterogenei. La pipeline end-to-end riproducibile fornisce un riferimento per la selezione dei metodi nella selvicoltura di precisione e nel monitoraggio della vegetazione urbana tramite immagini satellitari VHR.
Assessing tree crown detection methods for urban and rural monitoring using VHR remote sensing imagery
SARMIENTO OSPINA, NATALY ALEJANDRA
2025/2026
Abstract
Individual tree crown detection from very high resolution (VHR) imagery is a core problem in remote sensing for urban and rural vegetation monitoring, yet published studies rarely compare representative methods under identical experimental conditions. The primary goal of this thesis is to systematically compare five representative detection paradigms. This thesis presents a controlled benchmark of five crown detection paradigms applied to a single multispectral scene (JILIN-1 KF01C, approximately 0.5 m resolution, four spectral bands including near-infrared) acquired over Reggio Calabria, Italy, in June 2024, and evaluated against 800 manually delineated reference crown polygons distributed equally between urban and rural landscape strata. The five methods span the full methodological spectrum: local maxima with watershed segmentation, object-based image analysis (OBIA) with Random Forest, a pixel-based Random Forest classifier, DeepForest (a pretrained RetinaNet detector), and U-Net (a semantic segmentation model with a ResNet-34 encoder). All methods share a common preprocessing pipeline, a fixed 512 × 512 px tile grid with spatial block partitioning, and a shared evaluation using Hungarian assignment IoU matching at thresholds of 0.25 and 0.50 on a fixed 302-crown test set. The threshold of 0.25 serves as a lenient criterion that captures coarse crown detections and tolerates annotation boundary uncertainty, while 0.50 provides a stricter standard reflecting more precise crown delineation. As a benchmark comparison, no single method achieved satisfactory detection performance, and results are reported here as reference values rather than evidence of methodological superiority. Across the five methods, DeepForest produced the highest F1-score (F1 at 0.25 = 0.313; F1 at 0.50 = 0.297), the highest recall (R at 0.25 = 0.503), and the highest matched mean IoU (mean IoU at 0.25 = 0.666). U-Net achieved the highest precision (P at 0.25 = 0.307) but the lowest recall among the supervised methods (R at 0.25 = 0.235). The three handcrafted methods (watershed, OBIA with Random Forest, and pixel-based Random Forest) clustered in a narrow F1 band (F1 at 0.25 of 0.103–0.127). The rural stratum was consistently harder across all methods: DeepForest led in both strata but with a pronounced urban advantage (F1 at 0.25 = 0.430 urban versus 0.248 rural). The main failure modes were the spectral similarity between target crowns and surrounding Mediterranean scrub and olive vegetation, and building cast shadows in the urban core. A recurring finding across supervised methods was a train-test distribution shift that required per-stratum threshold recalibration. The benchmark demonstrates that detection accuracy and boundary fidelity are partly independent properties that must be reported jointly, and that a single model and threshold do not transfer optimally across heterogeneous urban and rural Mediterranean landscapes. The reproducible end-to-end pipeline provides a reference for method selection in precision forestry and urban vegetation monitoring under VHR satellite imagery.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Sarmiento.pdf
accessibile in internet per tutti
Descrizione: Documento thesis
Dimensione
39.07 MB
Formato
Adobe PDF
|
39.07 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/261482