3D object detection is a key component in autonomous driving. Among the sensors available in autonomous vehicles, LiDAR is the most suitable for 3D object detection. However, LiDAR-based 3D object detectors are subject to domain shifts in sensor characteristics, scene geography, and object-dimension statistics. Thus, detectors trained on one dataset degrade sharply when deployed on another. Annotating every new target domain is costly and impractical, motivating unsupervised domain adaptation, in which a detector is adapted using a labelled source domain together with only unlabelled target data. This thesis proposes an unsupervised domain adaptation method that closes the domain gap in LiDAR-based 3D object detection through coordinated intervention at the input, feature, and output levels, rather than at any single stage of the pipeline. At the core is the Mean-Teacher self-training scheme, in which a teacher produces filtered and corrected pseudo-labels for a student, and the teacher's model is updated as an exponential moving average of the student's weights. This is complemented at the input level by Random Object Scaling, which mitigates object-size mismatch between domains, and Hard Instance Sampling, which augments target scenes with difficult source objects. Also, at the feature level, Domain-Specific Normalization keeps source and target statistics separate. The method is evaluated on the challenging nuScenes-to-KITTI adaptation scenario using two representative detectors, CenterPoint and PointPillars. For CenterPoint, it recovers 63.67% of the BEV and 57.92% of the 3D average-precision gap, under the Strict IoU threshold, relative to the source-only baseline. For PointPillars, the corresponding figures are 40.55% and 35.57%. Ablation studies confirm that each component contributes to the gain, supporting the claim that effective adaptation requires a multi-level approach.
Il rilevamento di oggetti in 3D (3D object detection) è una componente fondamentale nella guida autonoma. Tra i sensori disponibili nei veicoli autonomi, il LiDAR è il più adatto a questo compito. Tuttavia, i rilevatori di oggetti 3D basati su LiDAR sono soggetti a domain shift nelle caratteristiche del sensore, nella geografia della scena e nelle statistiche delle dimensioni degli oggetti. Di conseguenza, i rilevatori addestrati su un dataset subiscono un netto degrado delle prestazioni quando vengono impiegati su un altro. Annotare ogni nuovo dominio target è costoso e poco pratico, il che motiva l'adozione dell'unsupervised domain adaptation (adattamento di dominio non supervisionato), in cui un rilevatore viene adattato utilizzando un dominio sorgente etichettato insieme a soli dati target non etichettati. Questa tesi propone un metodo di unsupervised domain adaptation che colma il divario tra domini (domain gap) nel rilevamento di oggetti 3D basato su LiDAR mediante un intervento coordinato a livello di input, di feature e di output, anziché in una singola fase della pipeline. Al centro vi è lo schema di self-training Mean-Teacher, in cui un teacher produce pseudo-etichette filtrate e corrette per uno student ed è aggiornato come media mobile esponenziale (exponential moving average) dei pesi dello student. Questo è completato, a livello di input, dal Random Object Scaling, che attenua la discrepanza nelle dimensioni degli oggetti tra i domini, e dall'Hard Instance Sampling, che arricchisce le scene target con oggetti sorgente difficili. Inoltre, a livello di feature, la Domain-Specific Normalization mantiene separate le statistiche del dominio sorgente e di quello target. Il metodo è valutato sul complesso scenario di adattamento da nuScenes a KITTI utilizzando due rilevatori rappresentativi, CenterPoint e PointPillars. Per CenterPoint, recupera il 63.67% del divario di average precision in BEV e il 57.92% in 3D, con la soglia IoU Strict, rispetto alla baseline source-only. Per PointPillars, i valori corrispondenti sono 40.55% e 35.57%. Gli studi di ablazione confermano che ogni componente contribuisce al miglioramento, a sostegno della tesi secondo cui un adattamento efficace richiede un approccio multi-livello.
Domain adaptation in LiDAR-based 3D object detection
Shahinpour, Erfan
2025/2026
Abstract
3D object detection is a key component in autonomous driving. Among the sensors available in autonomous vehicles, LiDAR is the most suitable for 3D object detection. However, LiDAR-based 3D object detectors are subject to domain shifts in sensor characteristics, scene geography, and object-dimension statistics. Thus, detectors trained on one dataset degrade sharply when deployed on another. Annotating every new target domain is costly and impractical, motivating unsupervised domain adaptation, in which a detector is adapted using a labelled source domain together with only unlabelled target data. This thesis proposes an unsupervised domain adaptation method that closes the domain gap in LiDAR-based 3D object detection through coordinated intervention at the input, feature, and output levels, rather than at any single stage of the pipeline. At the core is the Mean-Teacher self-training scheme, in which a teacher produces filtered and corrected pseudo-labels for a student, and the teacher's model is updated as an exponential moving average of the student's weights. This is complemented at the input level by Random Object Scaling, which mitigates object-size mismatch between domains, and Hard Instance Sampling, which augments target scenes with difficult source objects. Also, at the feature level, Domain-Specific Normalization keeps source and target statistics separate. The method is evaluated on the challenging nuScenes-to-KITTI adaptation scenario using two representative detectors, CenterPoint and PointPillars. For CenterPoint, it recovers 63.67% of the BEV and 57.92% of the 3D average-precision gap, under the Strict IoU threshold, relative to the source-only baseline. For PointPillars, the corresponding figures are 40.55% and 35.57%. Ablation studies confirm that each component contributes to the gain, supporting the claim that effective adaptation requires a multi-level approach.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_07_Shahinpour_Executive Summary_02.pdf
accessibile in internet per tutti
Descrizione: executive summary
Dimensione
665.04 kB
Formato
Adobe PDF
|
665.04 kB | Adobe PDF | Visualizza/Apri |
|
2026_07_Shahinpour_01.pdf
accessibile in internet per tutti
Descrizione: thesis text
Dimensione
6.56 MB
Formato
Adobe PDF
|
6.56 MB | Adobe PDF | Visualizza/Apri |
I documenti in POLITesi sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/10589/260845